I've done a complete U-turn on my opinion on open source AI. We should be careful (and probably disallow release) of models on par with current open source Astra/Opus level models.
AI is now reverse engineering games from binaries and remaking them. Reverse engineering games is a very hard problem. If this is possible then reverse engineering banking software, ID systems (Aadhaar), flights, etc is possible too. A game ships its binary to every player. Bank and Aadhaar backends are protected by servers, but it's unlikely the security teams at these are smarter than advanced AI that can do this to games.
We are protected rn because a Claude will refuse to tick "I'm not a robot" on websites. When OSS models don't respect that, we have a looming cybersecurity problem. Apologies I didn't see this earlier.
Knowing “exactly” what you want and that has always been the bottleneck.
If i know what i want, i can do what i want. I cannot however, want what i want. Only my heart can do that - want what it wants.
Some countries invest in sovereign AIs to balance bigger corps. Says a lot about an individual’s ability to disrupt and compete is such an environment.
Coempt Eduteck only qualified for CBSE On-Screen Marking because tender terms were changed to benefit it: @sidhant_sarthak, who revealed the changes, tells @kthaparoffice for The Wire.
Watch the full interview on The Wire's Youtube Channel at 01.30pm today.
seriously, working with AI is MISERABLE for one and only one reason: having to re-explain the same thing
"oh yeah this new session obviously doesn't know what proper case trees are, so let me explain it for the 5000th time in my life"
I'm tired
AGENTS.md doesn't solve this because it is impossible to fit the entire domain knowledge without nuking the context - it would be 1m+ tokens worth
RAGs don't solve this, the agent won't search unknown unknowns
SKILLs don't solve this unless I keep like a collection of 1750 skills with specific cuts of domain knowledge for each possible subset of my domain that I might need in a given chat, but that's a lot of manual work
recursive LLMs or whatever don't solve this for the same reason, you can't dump a domain book and expect the AGENT will magically guess that it is supposed to search for a specific bit knowledge. unknown unknowns
fine tuning doesn't solve this (OSS models suck and OpenAI / Anthropic gave up on user fine tuning)
I honestly think a good product around fine tuning on your domain would be a major hit and an underdog lab should take this opportunity
👋 We kept MRCR in the system card for scientific honesty, but we've actually been phasing it out slowly.
Two reasons: (1) it's built around stacking distractors to trick the model, which isn't how people actually use long context, and (2) we care more about applied long-context capability than needle-retrieval. Graphwalks is a better signal for applied reasoning over long context, and internally we've seen this model do really well on long-context code.
MRCR wasn't included in the Mythos Preview system card for these reasons, but Graphwalks was - that will be the case for future models too.
Known surface vs. novel surface might be the most underrated variable in agent evals. The Python shortcut problem is already documented in OSWorld. @EpochAIResearch found ~45% of tasks can be solved by scripting around the GUI using terminal, openpyxl instead of LibreOffice etc;
OSWorld is about computer use, but many tasks require little use of graphical user interfaces.
About 15% can be solved with only the terminal and a further 30% can rely heavily on Python scripts.
We even found cases of models downloading packages to manipulate spreadsheets.
🚨 Shocking: Frontier LLMs score 85-95% on standard coding benchmarks. We gave them equivalent problems in languages they couldn't have memorized. They collapsed to 0-11%.
Presenting EsoLang-Bench.
Accepted to the Logical Reasoning and ICBINB workshops at ICLR 2026 🧵
Browser and computer use feel like obvious candidates. Models score impressively on known interfaces. Given an unfamiliar UI - one they couldn’t have memorized and I suspect the collapse looks a lot like this.
Some countries invest in sovereign AIs to balance bigger corps. Says a lot about an individual’s ability to disrupt and compete is such an environment.
There’s two ways to look at AI and job displacement.
One is that the AI is taking away jobs and making big corps bigger without needing you.
The other is that AI now gives everyone the ability to compete with the big corp. Money is no longer a limiting factor in making things. You can sit in your bedroom and make a new world class IP or app. In fact having big teams and approval processes in place slows things down with magical tools like this.
AI is the great equalizer.
UPI is not an open interoperable protocol, it is a closed messaging protocol that happened to have some public documentation at launch.
Standards need public access and a vendors neutral standards body. NPCI fits neither.
Also, XMPP exists?
Nature doesn’t run on laws of physics, instead laws of physics run on our brain to approximate how nature runs.
How could we ever know what the “fundamental” layer of reality is, and if it even exists.
All we can do is observe and make models of what we observe.
I used em dashes (a lot) before ChatGPT made them the official punctuation of soulless AI prose.
(can't believe I've had to gentrify my own writing style)