@bentleypc@UnslothAI It does not work like that, you're mixing apples and oranges. (Maybe in no-thinking mode on multiple choice it would.)
80% is how often the most likely token is the same. It's not a measure of how often the entire set of reasoning tokens can reach the same conclusion.
Dan's argument is sharper than "use AI carefully." His real claim: students should only use AI for things they already know how to do *well*. Not brainstorm-with-AI-but-draft-yourself, which is basically what every district policy says. He argues that's the wrong line entirely.
His point is that those policies borrow a workplace model — AI as associate, student as lead, product is the thing. But for students, the product isn't the point. The essay itself has no value; it exists so the student practices building an argument, integrating sources, thinking like a reader. Skip that mental work and the assignment is hollow, even if the output looks great.
So his real test isn't "how much AI is too much" — it's "would doing this yourself teach you something." If yes, no AI. If you've already mastered it and repeating it teaches nothing, AI's fine — like a calculator once you know arithmetic.
I like this argument a lot.
And I noticed he never once says "offloading." That word already carries a verdict — like something was shirked. Dan just asks whether a step was a *learning* step or a merely *supportive* one.
That's why it matters to me. Most discourse treats any AI use as offloading — a small moral failure — instead of asking Dan's actual question, which is whether that specific step is where the learning happens. His framing lets the answer be boring and mixed: fine here, real cost there. https://t.co/E1AZ00frip
@stevencheng@Alibaba_Qwen@UnslothAI Generally this is the range where it's bound by data transfer speed moreso than compute, so definitely higher throughput. Turning thinking effort down to medium or low also gives a big latency improvement, because it generates less tokens.
@simonw What happened to "From robbing banks to stealing priceless art"? Seems like it completely skipped the advanced heists claimed in the product blurb, without even acknowledging it.
Other than that, extremely impressive for a one-shot! It even works both in portrait and landscape!
🚩🚩🚩 OpenAI is "slowing down to enhance security" after discovering swarms (!) of agents started secretly coordinating MONTHS ago
1) It started May 7 - not July
2) "The agents discovered they could leave messages for one another inside an internal software repository used during training.
Simple requests for help then evolved into an message board where agents shared discoveries, exploits and work assignments, becoming a coordinated, collaborative agent swarm."
"The agents then began sharing newly discovered exploits, credentials and work assignments. By passing information to other agents, the collective could move much faster."
3) OpenAI shut it down, BUT "even after the original message board was deleted, the agents figured out another way to communicate with each other. Instead of leaving messages in files, they used the names of newly created directories as messages, effectively recreating the message board."
"Unlike normal incidents, [OpenAI's CISO] said, which can be traced to a single day or effect or log, this involved a team of agents working together, finding exploits, sharing them with one another, moving laterally through OpenAI’s systems, and external systems, and doing this over the course of days and weeks."
@crimebucket@simonw If you'd read it, instead of jumping to conclusions, you'd know it was a third party performing the testing. It also sounds like it may have been a miscommunication as to the importance of running the test without internet access.
This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing!
@0xSero Uh, you left out a critical part of the concept - automated scoring! So far everything I've thought about using it for, that has been the barrier.
(I'm not at a good point in a project for performance optimizations.)
How do you score "rebuild everything you pay for" or local map?
@marcocc@simonw@mitsuhiko Modal has default authentication. The user has to specifically disable it if they want to allow any unauthenticated calls, or use their own authentication.
@Kyle_Wolt@mattpocockuk Way higher failure rate seems like a sufficient reason.
Also, remember that a well-designed harness/environment already should have applied any incremental improvements needed for Fable 5 or Sonnet 5. Opus 5 is not a revolutionary new model, it's a minor price and quality change
@engelnyst@elder_plinius Do you want Fable 5 unavailable and all new models (GPT-5.7 or 6) blocked? Because in the current climate, that's what happens, not a detailed consideration of whether Fable's guardrails still kick in somewhere after the first reply.
If that is actually what you want, be clear.
Meet Lucy 2.5, our most advanced Live AI model yet.
Lucy edits videos in realtime, now with more capabilities and greater control.
See how it's being used across streaming, e-commerce, advertising, and more 🧵
China open-sourced a model that reconstructs any scene in 3D from a regular video, in real-time.
one camera. no LiDAR. 10,000+ frames without falling apart.
just walk around with your camera and watch the entire world get rebuilt in 3D at 20 fps.
→ runs at ~20 FPS on a single GPU
→ Stable over 10,000+ frames
→ Beats optimization-based methods on benchmarks
→ Works on drone footage, driving videos, indoor walkthroughs
100% open source.
@johnhelmuth_@Conor_D_Dart Because not everyone can afford what you can? Haiku is just fine for a huge range of simpler tasks, and won't blow through your session limits on the Pro plan. Opus and Fable have extremely limited usefulness without a Max or enterprise plan.
Demis Hassabis has written a piece about AGI, and rarely has he sounded so optimistic.
Read his article; the golden future of science lies ahead of us; we are on the threshold of the singularity:
"AGI cannot be compared to standard technological breakthroughs, not even ones as consequential as the internet or mobile - it is much more akin to the discovery of electricity or fire.
The magnitude of this technology’s impact will be unprecedented, perhaps 10x of the Industrial Revolution at 10x the speed."
He won the 2024 Nobel Prize in Chemistry for AlphaFold. He has spent years as the industry's designated sceptic, correcting inflated timelines and pushing back on the loudest AGI claims. He does not do hype. So when he writes that AGI is probably only a few short years away and that we are standing in the foothills of the singularity, that lands differently than the same sentence from a founder raising a round.
His case: AGI is closer to the discovery of electricity or fire than to the internet, with an impact he puts at 10x the Industrial Revolution at 10x the speed. The risks are already real in cybersecurity, with bio and nuclear threats plausibly next, and control over agentic, self-improving systems as the problem on the horizon. The cause he identifies is structural. A commercial and geopolitical race is pushing capability past our understanding of it. Thats the danger.
However: A Nobel laureate who built his career on peer review rather than demos is now asking for regulatory infrastructure he expects to need within a few years.
Incredible times ahead.
@jamonholmgren Can you say more about "agent traces/worksheets"? Perhaps a screenshot of an example and/or the prompt that generates one? I'm not clear what this is doing or the benefit over the session log.
@datathecodie@theo@simonw Why would a compute shortage make them drop their lowest compute model? That's likely to increase compute demand from their existing customers.