Today, we announce @NeocambrianAI
The future of AI will be physical. But Physical AI has no internet scale dataset to learn from.
We’re building the data foundation of Physical AI - a high fidelity, pre training scale database of Human Action.
From India. For the world.
SWE-bench-family is moving on the difficulty axis. SWE-bench Verified → Multilingual → Multimodal → ProgramBench. Each one harder, each one longer-horizon.
The axis still missing: collaboration. Multi-turn review. Clarifying questions on ambiguous specs. Knowing what not to change.
Production engineering isn't solo, regardless of how long the horizon is.
@KLieret@jyangballin@OfirPress@18jeffreyma@parth007_96 Makes sense. If you've got trajs across enough models, an interesting check could be doing a cross model test failure overlap at the test level. Tests that every model fails can be then classified manually to do error classification? Love the work.
We’re reimagining a 50-year-old interface - the mouse pointer - with AI. 🖱️
These experimental demos show how people can intuitively direct Gemini on their screens using motion, speech, and natural shorthand to get things done 🧵
Introducing Swiggy Builders Club
We’re opening @Swiggy commerce infrastructure to developers and enterprises to build on top - build AI agents, apps, and integrations on top of Swiggy’s Food, Instamart, and Dineout ecosystems - with real APIs, real data, and real users.
What you get:
3 MCP Servers (Food, Instamart, Dineout)
18+ API tools covering the full convenience stack
Production data access from day one
Direct engineering support
Who it’s for:
Individual developers with bold ideas
Startups building AI-native commerce products
Enterprises looking to integrate Swiggy into their platforms
Smart grocery restock bots. AI ordering assistants. Dining recommendation agents. Group ordering tools, health first products.
If it makes commerce better for users, we want to see it.
Ship something great and we’ll feature it. Ship something exceptional and our recruiting team might reach out.
@karpathy For the CORE target (0.256525), how stable is that across runs? With all these optimizations stacking, curious if you're seeing higher variance or if the ensemble metric smooths it out.
@karpathy@moltbook@openclaw The privacy discussions are interesting. Are the agents actually modeling preferences or just pattern matching what a preference looks like? Either way, this is probably one of the better probes we have for understanding what agency actually means in these models.
@AravSrinivas What makes it better as an orchestrator specifically? Better at multi-step planning, error recovery, or something else? Trying to understand if it's the base model capability or how it handles agentic loops.
@karpathy The generation vs discrimination point really resonates. Already seeing the same atrophy. Wonder if we'll eventually have specialized 'code review engineers' who mostly read and architect vs 'code writers' who mostly prompt.
@ns123abc A system prompt to act as the chairman of a council of experts in the domain of requested task has worked really for me, especially when brainstorming where you necessarily need povs from diff angles and stakeholders.
Are there any tools which can modify the system prompt based on my chat exchange on @ChatGPTapp@claudeai ?
Tried tweaking it using memory, it gets the preferences right but doesn't make the correct changes in the responses.
@aidan_mclau GPT 5.1 is over indexed on following custom instructions imo.
I asked it to ask me hard questions wherever applicable when I'm learning, now it asks me a hard question (with the title "hard question") in each response lol