Banger paper from Google Cloud AI Research.
If you follow autonomous research agents, this one is worth your time.
ScientistTwo takes a problem from a human expert and runs the full discovery cycle without further intervention.
It establishes state-of-the-art baselines, generates seed ideas aimed at resolving stated limitations in existing work, screens each idea on a data subset before committing to full experiments, runs its own ablation studies to attribute the gain, and revises the idea from that attribution.
Manuscript drafting includes a simulated peer-review and rebuttal engine, plus a meta-review pass that feeds back into the work.
They benchmark it against papers accepted at ICLR, ICML and NeurIPS. The reported solutions outperform the human state-of-the-art models on those problems, and the generated papers score higher average ratings than the human-authored ones under automated AI reviewers.
Paper: https://t.co/3i9Zgsahj4
this paper is f*cking brilliant
a Google paper replaces append-only conversation history with SKILL.state to scale long-horizon agent skills
the result: explicit state updates eliminate context-poisoning failures while cutting cumulative token usage by 16x at 100 execution turns
the crazy part is how mutable state abstractions prevent context window collapse
instead of appending endless observation logs, SKILL.state feeds the model only the skill specification, latest observation, and a structured state object
most agent runtimes continuously stuff past conversation logs into an ever-growing prompt context
this architecture converts long-running skill execution into a constant prompt footprint
read the complete paper + article below
bookmark it for future reference
They were building in stealth for 2 years, I was building in stealth for 2 hours���
Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe.
⚡️Demo below on a M4 MacBook⚡️
every LLM has the ability to efficiently batch inference every key of a JSON at the same time and generate probabilities from a set of possible categories. No new training required, but it’s easy to optimize if you need!
On hugging face now!
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
@thsottiaux “Selected model is at capacity. Please try a different model” I am a Pro20x user. This issue has been bothering me for several days, and I haven't been able to work. If this continues, I will cancel my Codex Pro subscription.
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz