AI systems often use more direct experience than a human could get in a lifetime. Humans require less experience, because they "transfer" past experience to new tasks. Recent work I led found an equation to characterize transfer in a simple setting.
https://t.co/2OezsOi8lV
👇
Today, AI can generate tons of code—but how do we know if it's good?
That's why we built Sculptor: the first coding agent environment.
Sculptor helps you catch issues, write tests, and improve your code—all while you work in your favorite editor.
Introducing Claude for Enterprise.
Now your entire organization can collaborate securely with Claude—with no training on chats or files.
Comes with:
📚 Expanded 500K context window
🧑💻 Native GitHub integration
🔐 Enterprise-grade security features
https://t.co/CUiwAkKpoh
A big part of my job these days is to think about what technical work Anthropic needs to do to make things go well with the development of very powerful AI.
I digested my thinking on this, plus some of the Anthropic zeitgeist around it, into this piece:
https://t.co/dXAwiUNI6I
I'm excited to join @AnthropicAI to continue the superalignment mission!
My new team will work on scalable oversight, weak-to-strong generalization, and automated alignment research.
If you're interested in joining, my dms are open.
New Anthropic research paper: Scaling Monosemanticity.
The first ever detailed look inside a leading large language model.
Read the blog post here: https://t.co/6RYwxt6nWI
Today, we're announcing Claude 3, our next generation of AI models.
The three state-of-the-art models—Claude 3 Opus, Claude 3 Sonnet, and Claude 3 Haiku—set new industry benchmarks across reasoning, math, coding, multilingual understanding, and vision.
New Anthropic Paper: Sleeper Agents.
We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through.
https://t.co/mIl4aStR1F
Our new model Claude 2.1 offers an industry-leading 200K token context window, a 2x decrease in hallucination rates, system prompts, tool use, and updated pricing.
Claude 2.1 is available over API in our Console, and is powering our https://t.co/uLbS2JNczH chat experience.
Today, we’re publishing our Responsible Scaling Policy (RSP) – a series of technical and organizational protocols to help us manage the risks of developing increasingly capable AI systems.
We should expect a jump that in some sense feels like the jump between gpt3 and claude/gpt4 in the next ~2 years, based on smooth underlying exponentials in effective compute.
There is lots of meaning in trying to make that go well and meaning is top of the hierarchy of needs.
How quickly is A.I. advancing? And should you be working in the field? Checkout my recent conversation on these topics with @Hernandez_Danny:
https://t.co/6SyZzRKur6
Introducing Claude 2! Our latest model has improved performance in coding, math and reasoning. It can produce longer responses, and is available in a new public-facing beta website at https://t.co/uLbS2JNczH in the US and UK.
We develop a method to test global opinions represented in language models. We find the opinions represented by the models are most similar to those of the participants in USA, Canada, and some European countries. We also show the responses are steerable in separate experiments.
Introducing 100K Context Windows! We’ve expanded Claude’s context window to 100,000 tokens of text, corresponding to around 75K words. Submit hundreds of pages of materials for Claude to digest and analyze. Conversations with Claude can go on for hours or days.
Proud to partner with @awscloud to give people an easy way to access Claude in their cloud environments!
We believe this work will drive immense value for businesses looking to build generative AI applications with AWS tools and capabilities.
Stay tuned for more.
Today we are releasing the new Claude App for @SlackHQ, in beta. Now every company in the world has the chance to have a “virtual teammate” who can help make work more fun and productive.
Only model I'm aware of comparable to ChatGPT.
Have some notable customers. Some prefer its personality, clarity, summarization, creativity, etc. I Makes sense for any business using ChatGPT to compare it to Claude.
After working for the past few moths with key partners like @NotionHQ, @Quora, and @DuckDuckGo, we’ve been able to carefully test out our systems in the wild. We are now opening up access to Claude, our AI assistant, to power businesses at scale.
Writing AI evals is the first AI researcher task I've seen where models feel clearly better than me.
This automation will enable much better measurement and understanding of LM's broadly.
Reduces eval iteration time from ~week to ~hour.
>10x more leverage from a researcher.
It’s hard work to make evaluations for language models (LMs). We’ve developed an automated way to generate evaluations with LMs, significantly reducing the effort involved. We test LMs using >150 LM-written evaluations, uncovering novel LM behaviors.
https://t.co/1olqJSvhDA