New Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore.
Scientific discovery begins where known problems end: a system has to frame hypotheses, design experiments, and learn from the results. Evaluating this is hard. Genuinely new answers cannot be checked quickly, and in familiar domains a model can simply recall what it has seen.
Addressing this challenge, researchers from Tencent Hy, Fudan University, and Tsinghua University built verifiable Alien Worlds. Their rules are executable, so every answer is checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks.
🔹 Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems)
🔹 A flawed manual, four rounds of self-designed probes, and closed-book tests after every round
🔹 Every answer graded by an interpreter or a proof checker, with no LLM judge
What we found across 10 frontier AI systems:
1️⃣ Getting feedback is more effective than thinking alone. No AlienCode run starts above 15.7%; after four rounds the best reaches 89.0%, while the same turns without feedback stay at 0.5–11.0%.
2️⃣ Designing the experiments matters. Replaying a system's own best probes gives it exactly the same evidence, yet in AlienCode 9 of 10 systems do worse than when they chose the probes themselves.
3️⃣ Knowing a rule is not using it. Even when every required rule is stated correctly, tasks are solved only 73.4% of the time.
4️⃣ One score hides a lot. The same system under the same budget ended anywhere from 5.7% to 79.0%, and rankings barely transfer between the two worlds.
CL-bench asked whether models can learn from context. ExplorationBench asks whether they can discover the rules themselves.
📄 Paper: https://t.co/eamwrsJyxQ
🌐 Website & leaderboard: https://t.co/gxSeoCfc2s
📝 Blog: https://t.co/jyrwH7ZnQH
💻 Code (coming soon): https://t.co/i7VlgErkcl
The Rise and Potential of LLM Based Agents
This is probably the most comprehensive overview of LLM based agents.
From how to construct these agents to how to harness them for good.
A great read for the weekend.
https://t.co/zH0wCMGEok
LLMs Can Align Themselves without Finetuning?
This paper discovers "that by integrating self-evaluation and rewind mechanisms, unaligned LLMs can directly produce responses consistent with human preferences via self-boosting"
Does seem to have benefits in terms of generating safer outputs. We have seen the importance of high-quality data to align these models. However, this work approaches alignment without any extra data or any training for that matter.
The authors claim that their approach can "improve the harmlessness rate of LLaMA 30B over vanilla inference from 82% to 97%, while maintaining the helpfulness rate."
While it doesn't require any weight updates or fine-tuning, the disadvantage of this method is the longer inference time (a 4-fold increase on the LLaMA 30B model and the HH dataset).
Interestingly, the author suggests potentially using this approach to generate better-quality data that can subsequently be used to finetune models.
The inference speed will only improve in the coming months, so it will be interesting to see how these slower methods to interact with LLMs, including self-consistency and self-refinement, will benefit and reemerge as effective ways to tune prompts.
Also, not all approaches are built the same. Some will be helpful for things like data augmentation/generation, some for eliciting reasoning, and some for fast inference.
There is a lot to explore in this space.
https://t.co/uyMfKRVdER
@omarsar0 Thanks for recommending our work! In the paper, we also talk about opportunities and challenges of Agent Society.
Welcome everyone to follow and discuss!😃😃😃
🔗github:https://t.co/K2oJZS2uew
The luminous clouds of Jupiter! ☁️ Taken by our @NASAJuno mission on its 20th close pass of the planet, this view reveals the highest clouds in bright white casting shadows on layers of clouds below. Image processed by citizen scientist Kevin M. Gill: https://t.co/SJcHiq8QR3