I made Codex Pet talk.
Talking Pets is a small local add-on that reads Codex Pet bubbles / assistant replies aloud with local TTS.
No Codex patching.
No signed app modification.
Just local logs + local voice.
Demo:
[attach video]
Repo:
https://t.co/eMlb4KxllD
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
Too many people have been psyopped into seeing AI only through the lens of fear and dystopia
Almost anyone I ask gives me the same answer. I really haven't met many people outside the tech community who are optimistic that AI is going to create a prosperous future
For some reason, that entire idea feels dead
It's like people have abandoned the possibility of an abundant, prosperous and really good future
that's incredibly sad, because that future is still there for us to build
After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework.
This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development happens safely and securely.
We're working hard to make Astra broadly available, and get its advanced cyber capabilities into the hands of defenders.
https://t.co/9lMIvdeMMJ
Lost my phone at the office and spent 30 minutes turning the place over. Find My was disabled by MDM.
Out of ideas, I asked Claude how I could find it. It suggested tracking the Bluetooth signal strength, then wrote me a meter in about a minute.
I walked around watching the number climb. Found it.
Apparently you can just make the tool you need now.
Code: https://t.co/fmnISzHfZ2
MIT and Harvard argue LLMs are nowhere near doing real scientific discovery.
They published a paper called “Evaluating Large Language Models in Scientific Discovery.”
Every week, tech labs claim an LLM has made a breakthrough in biology, physics, or chemistry.
But this proves they are faking it.
For years, AI benchmarks have tested models using static, multiple-choice science trivia. Models ace these tests, leading everyone to believe AI is right on the verge of autonomous scientific discovery.
Researchers built a new evaluation framework called SDE to test what happens when you take LLMs out of the multiple-choice quiz and put them into real, open-ended research projects.
They tested frontier models across biology, chemistry, materials science, and physics.
The results are sobering.
When forced to handle the actual loop of discovery—proposing a testable hypothesis, designing simulations, running experiments, and interpreting ambiguous results iteratively, current LLMs fall apart.
There is a massive, glaring performance gap between passing standard science benchmarks and doing real science.
Why do they fail? Because real science requires iterative reasoning, handling imperfect evidence, and adapting to unexpected observations.
LLMs are built to predict the next token based on existing internet data. They can regurgitate a textbook explanation of photosynthesis or quantum mechanics instantly.
But when placed inside an uncharted loop where the textbook doesn't have the answer yet, they hit a wall.
Worse still, the researchers discovered diminishing returns. Simply scaling up model sizes and adding raw compute isn't fixing the gap. Top-tier models from different providers share the exact same blind spots.
We are miles away from general scientific superintelligence.
The tech industry is selling a narrative that AI is about to automate labs, run clinical trials, and invent materials on autopilot.
But right now, AI isn't doing science.
It's just remembering it.
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
we've cut prices on luna by 80%, making it by far the most price-efficient model in its class.
a lot of our research is about how to create incredibly efficient models for any given level of intelligence.
excited to see what you all do with intelligence too cheap to meter!