New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude.
Memory on Claude Managed Agents is now in public beta.
Your agents can now learn from every session, using an intelligence-optimized memory layer that balances performance with flexibility.
Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software.
It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans.
https://t.co/NQ7IfEtYk7
@DoorDash_Help A Dasher faked delivery by using the SAME photo for 2 different orders. $122.30 order never arrived. Refund denied without a single explanation, no investigation, nothing. Not the first time. @DoorDash keeps the money, the customer keeps the loss. #doordash
nanochat now trains GPT-2 capability model in just 2 hours on a single 8XH100 node (down from ~3 hours 1 month ago). Getting a lot closer to ~interactive! A bunch of tuning and features (fp8) went in but the biggest difference was a switch of the dataset from FineWeb-edu to NVIDIA ClimbMix (nice work NVIDIA!). I had tried Olmo, FineWeb, DCLM which all led to regressions, ClimbMix worked really well out of the box (to the point that I am slightly suspicious about about goodharting, though reading the paper it seems ~ok).
In other news, after trying a few approaches for how to set things up, I now have AI Agents iterating on nanochat automatically, so I'll just leave this running for a while, go relax a bit and enjoy the feeling of post-agi :). Visualized here as an example: 110 changes made over the last ~12 hours, bringing the validation loss so far from 0.862415 down to 0.858039 for a d12 model, at no cost to wall clock time. The agent works on a feature branch, tries out ideas, merges them when they work and iterates. Amusingly, over the last ~2 weeks I almost feel like I've iterated more on the "meta-setup" where I optimize and tune the agent flows even more than the nanochat repo directly.
🔷What is the most likely diagnosis in this 50 y/o M w/ PMH of sickle cell disease initially presenting w/ pain crisis who subsequently developed acute encephalopathy, hypoxic respiratory failure and multiorgan dysfunction? He also developed progressive anemia, thrombocytopenia and hyperbilirubinemia.
#radres #futureradres #medicine #neurology #neurosurgery #ENT #MRI #FOAMed @Radiopaedia@RSNA #Ophthalmology
⚡️ Multiple sclerosis is a potential cause of secondary trigeminal neuralgia. Patients w/ MS have a 20-fold increased risk of developing trigeminal neuralgia ⚡️
#Neurology#Neurosurgery#radres#futureradres#Medicine#FOAMed#MRI#Radiology#MS@TheASNR
*(Used the T1 here since others were degraded but normally FLAIR shows the plaques the best & CISS or FIESTA show atrophy the best; also, ignore the partially imaged left vest. schwannoma)*
Google Research introduced Learn Your Way, an AI-powered experiment that reimagines textbooks into personalized, multimodal learning experiences.
Built with LearnLM and integrated into Gemini 2.5 Pro, it adapts content to students’ grade level and interests, then generates multiple representations like narrated slides, quizzes, audio lessons, and mind maps.
In a study, students using Learn Your Way scored 11% higher on retention tests than those with standard digital readers.
Your connected tools are now available in Claude on your mobile device.
Now you can access projects, create new docs, and complete work while on the go.
@bee__computer Hey. I bought the device 10 days ago and no updates or reply of the customer service. Could you please provide me with some information?
"We show that current state-of-the-art #LLMs do not accurately diagnose patients across all pathologies, follow neither diagnostic nor treatment guidelines, and cannot interpret laboratory results, thus posing a serious risk" @NatureMedicine
https://t.co/pP5K3mhFeJ
Read the perspectives of experts from @MICCAI_Society and @RSNA on the clinical, cultural, computational, and regulatory considerations to adopt #AI technology successfully in radiology https://t.co/CsPEVa2O1R #AIME2024
Having trouble remembering what you should look for in vascular dementia on imaging?
Almost everyone worked up for dementia has infarcts. Which ones are important?
Here's what you must remember when the patient can't remember--what's important to look for in vascular dementia
▶️Subcortical infarcts:
Breaks important white matter connections between parts of the brain so they can’t function
▶️Hypoperfusion cortical infarcts:
🔸These infarcts don’t cause dementia themselves—they’re just a sign of the underlying disease.
🔸Indicate chronic neuronal hypoperfusion at a cellular level causing damage, dysfunction & dementia
▶️Hemorrhage:
🔸Sign of both hypertensive & amyloid small vessel disease.
🔸Amyloid angiopathy has a very strong correlation w/dementia
🔸Amyloid causes neurodegeneration & stroke by build up of amyloid proteins in the vessel wall leading to hemorrhage & decreased waste clearance
▶️Strategic infarcts:
🔸Infarcts located in structures directly related to cognition
🔸Remember w/the mnemonic: One HIT CAUses dementia
H=hippocampus
I=insula
T=thalamus
CAU=CAUdate
So now you know the important signs to look for when you are reading a study for vascular dementia.
Now you will always know what to de-mention when the patient has dementia!
Last week, I spoke about AI and regulations at an event at the U.S. Capitol attended by legislative and business leaders. I’m encouraged by the progress the open source community has made fending off regulations that would have stifled innovation. But opponents of open source are continuing to shift their arguments, with the latest worries centering on open source's impact on national security. I hope we’ll all keep protecting open source!
Based on my conversations with legislators, I’m encouraged by the progress the U.S. federal government has made getting a realistic grasp of AI’s risks. To be clear, guardrails are needed. But they should be applied to AI applications, not to general-purpose AI technology.
Nonetheless, some companies are eager to limit open source, possibly to protect the value of massive investments they’ve made in proprietary models and to deter competitors. It has been fascinating to watch their arguments change over time.
For instance, about 12 months ago, the Center For AI Safety’s “Statement on AI Risk” warned that AI could cause human extinction and stoked fears of AI taking over. This alarmed leaders in Washington. But many people in AI pointed out that this dystopian science-fiction scenario had little basis in reality. About six months later, when I testified at the U.S. Senate’s AI Insight forum, legislators no longer worried much about an AI takeover.
Then the opponents of open source shifted gears. Their leading argument shifted to the risk of AI helping to create bioweapons. Soon afterward, OpenAI and RAND showed that current AI does not significantly increase the ability of malefactors to build bioweapons. This fear of AI-enabled bioweapons has diminished. To be sure, the possibility that bad actors could use bioweapons — with or without AI — remains a topic of great international concern.
The latest argument for blocking open source AI has shifted to national security. AI is useful for both economic competition and warfare, and open source opponents say the U.S. should make sure its adversaries don’t have access to the latest foundation models. While I don’t want authoritarian governments to use AI, particularly to wage unjust wars, the LLM cat is out of the bag, and authoritarian countries will fill the vacuum if democratic nations limit access. When, some day, a child asks an AI system questions about democracy, the role of a free press, or the function of an independent judiciary in preserving the rule of law, I would like the AI to reflect democratic values rather than favor authoritarian leaders’ goals over, say, human rights.
I came away from Washington optimistic about the progress we’ve made. A year ago, legislators seemed to me to spend 80% of their time talking about guardrails for AI and 20% about investing in innovation. I was delighted that the ratio has flipped, and there was far more talk of investing in innovation.
Looking beyond the U.S. federal government, there are many jurisdictions globally. Unfortunately, arguments in favor of regulations that would stifle AI development continue to proliferate. But I’ve learned from my trips to Washington and other nations’ capitals that talking to regulators does have an impact. If you have a chance to talk to a regulator at any level, I hope you’ll do what you can to help governments better understand AI.
[Original text (with links): https://t.co/tw2iT0ooLT ]
I've really enjoyed using @crewAIInc 's tools to build multiagent AI systems -- in addition to being productive, it's also fun to use! It was great hanging out with its creator @joaomdmoura to chat about best practices for building agentic workflows.
Multi-agent collaboration has emerged as a key AI agentic design pattern. Given a complex task like writing software, a multi-agent approach would break down the task into subtasks to be executed by different roles -- such as a software engineer, product manager, designer, QA (quality assurance) engineer, and so on -- and have different agents accomplish different subtasks.
Different agents might be built by prompting one LLM (or, if you prefer, different LLMs) to carry out different tasks. For example, to build a software engineer agent, we might prompt the LLM: "You are an expert in writing clear, efficient code. Write code to perform the task …".
It might seem counterintuitive that, although we are making multiple calls to the same LLM, we apply the programming abstraction of using multiple agents. I'd like to offer a few reasons:
- It works! Many teams are getting good results with this method, and there's nothing like results! Further, ablation studies (for example, in the AutoGen paper cited below) show that multiple agents give superior performance to a single agent.
- Even though some LLMs today can accept very long input contexts (for instance, Gemini 1.5 Pro accepts 1 million tokens), their ability to truly understand long, complex inputs is mixed. An agentic workflow in which the LLM is prompted to focus on one thing at a time can give better performance. By telling it when it should play software engineer, we can also specify what is important in that subtask: For example, the prompt above emphasized clear, efficient code as opposed to, say, scalable and highly secure code. By decomposing the overall task into subtasks, we can optimize the subtasks better.
- Perhaps most important, the multi-agent design pattern gives us, as developers, a framework for breaking down complex tasks into subtasks. When writing code to run on a single CPU, we often break our program up into different processes or threads. This is a useful abstraction that lets us decompose a task -- like implementing a web browser -- into subtasks that are easier to code. I find thinking through multi-agents roles to be a useful abstraction.
In many companies, managers routinely decide what roles to hire, and then how to split complex projects -- like writing a large piece of software or preparing a research report -- into smaller tasks to assign to employees with different specialties. Using multiple agents is analogous. Each agent implements its own workflow, has its own memory (itself a rapidly evolving area in agentic technologies -- how can an agent remember enough of its past interactions to perform better on upcoming ones?), and may ask other agents for help. Agents themselves can also engage in Planning and Tool Use. This results in a cacophony of LLM calls and message passing between agents that can result in very complex workflows.
While managing people is hard, it's a sufficiently familiar idea that it gives us a mental framework for how to "hire" and assign tasks to our AI agents. Fortunately, the damage from mismanaging an AI agent is much lower than that from mismanaging humans!
Emerging frameworks like AutoGen, Crew AI, and LangGraph, provide rich ways to build multi-agent solutions to problems. If you're interested in playing with a fun multi-agent system, also check out ChatDev, an open source implementation of a set of agents that run a virtual software company. I encourage you to check out their github repo and perhaps even clone the repo and run the system yourself. While it may not always produce what you want, you might be amazed at how well it does!
Like the design pattern of Planning, I find the output quality of multi-agent collaboration hard to predict. The more mature patterns of Reflection and Tool use are more reliable. I hope you enjoy playing with these agentic design patterns and that they produce amazing results for you!
If you're interested in learning more, I recommend:
- Communicative Agents for Software Development, Qian et al. (2023) (the ChatDev paper)
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Wu et al. (2023)
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework, Hong et al. (2023)
[Original text: https://t.co/4gTbcQfikx ]