New Fellows Research: Can Claude autonomously align other AIs?
We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well.
https://t.co/nhlCMgQl46
@devsharma_8 True, but I would add that new tech always goes through a phase where their job is to convince the user that they can be used in a certain way. Which is why they start looking like this and then eventually look like this... with models the jump from how->what will feel like this
I think people are missing that the most important loops are learning loops to detect what you should do.
Everyone I know in AI starts with this type, except for very specific products where there's nothing to be learned yet.
Start with observations then derive lessons.
I love how claude loves loops, recursion, nesting, fractals - something about the model attends to the core pieces of a good AI system (or any system) in the near future. The model is training us as its users on it's own taste, which is itself recursion.
Establishment-quaking: To the extent our systems meant to keep power accountable to the citizen have been hollowed out by the professionals who operate them, soon, inexpensive intelligence will allow any single citizen to re-close those feedback loops themselves.
Of all inevitable AI effects, the most exciting is the coming automation of all a citizen's potential means of penetrating and correcting administrative barriers and information asymmetries governments enjoy or exploit to sustain lazy, malicious or otherwise wrong policies.
Scaling laws have held for 15 orders of magnitude. In ~2 years, better than a coin flip we get to the RSI moment.
And scale is faster on coordination tasks than judgement tasks. We're excited, but sober about the latter and loud about implications as we watch these gaps close.
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor.
It’s happening faster than we thought, and the implications deserve greater attention. https://t.co/OVVPJO7VQx
Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
Textract OCR + Prompt Engineer Agent for extractors = 95% accurate extraction pipeline.
https://t.co/Cqe7Zp322J
Next extension is to design a system that detects when inputs drift and you'd get a full Learning Loop around any semantic entity you'd like. :)
@googlemaps an ad got in the way of my click in search results and I just ended up following directions halfway across town to the wrong location. What have you done
The moment I can ask Gemini what terminal I just landed in I will consider Gemini a major product in my life
"Where am I Gemini?" there you go Goog, ad concept