Unlocking AI in Web3 | AI Content | Love to play with Grok Imagine | Outer space is a vast | Cosmic Discoveries | The Moon | The Mars | DYOR always | "INDIAN"
Most of the bill for long-running coding agents is not reasoning.
It is the harness: replaying context, dumping giant observations, and spending a full model turn on work that should have been one tool call.
SoL-Pi (NVIDIA / NTU / MIT) treats that harness as the thing to search over.
A research loop inspects traces, proposes harness changes, and keeps only the ones that clear both a capability gate and an efficiency gate.
Four mechanisms survive: Action Fusion, ObservationPack, Online Context Compact, and an Evidence-Preserving Reducer.On the 51-task EdgeBench suite, SoL-Pi stays comparable to Pi on GPT-5.6 Sol and Opus 5 while cutting recorded token traffic by 44.7–49% and API cost by about one third. Estimated savings: $8.75–$13.50/hr vs native Codex and Claude Code harnesses, $4.36–$5.71/hr vs Pi.
The useful idea is not make the agent cheaper. It is: scale auto-research until the harness stops wasting tokens on work the agent already finished.
https://t.co/igbW1tILI0
She's a fly, she thinks she's a helicopter, the Air Force disagrees.
Watch a giant fly patrol downtown at dusk canyon runs between skyscrapers, searchlights tracking her, tanks aiming up, and one slow-motion mid-air swat that ends a fighter jet's whole career.
She filed no flight plan.
She regrets nothing.
#3D #ThreeJS #Animation #WebGL
A teacher model that loves cats can generate a dataset of random-looking number sequences. Nothing in the text says “cats.” Fine-tune a student on it anyway, and the preference can still transfer.
That is subliminal learning. It is also a data-poisoning problem, because the trait is not legible in the dataset.
New Stanford work by Nathan Hu, Sanmi Koyejo, and Christopher Potts makes the hidden signal readable.
They treat prompted subliminal learning as a special case of context distillation: in theory, the dataset identifies the teacher’s prompt. Then they recover that prompt with SALVE (Search-Aided Latent Verbalization):
- Optimize a soft prompt so the original model predicts the dataset well
- Ask the same model to verbalize the soft prompt
- Use beam search to keep the wording fluent and faithful
On the standard animal-preference setup, SALVE recovers prompts that name the trait in 18/20 runs. Common text-optimization baselines do not. In some cases it even recovers the trait from data where student fine-tuning fails to pick it up.
They also detect the effect in mixed data, activation-steered teachers, and subsets of real preference data selected with Logit-Linear Selection (sycophancy / misalignment).
The practical point is simple: you may be able to audit a dataset for hidden behaviors before it ever reaches fine-tuning.
Paper: https://t.co/1mncXR5Fyv
Insiders quoted by the New York Post accuse OpenAI and Anthropic of overstating recent model incidents to push Washington toward rules that could protect incumbents.
That’s an allegation not a proven fact.
You're a fly driving through Times Square at midnight in a red convertible 🪰🌃
Rain on the neon. Synthwave on the radio. Wind in your wings. Not a single care in the world.
Built entirely in code. Three.js + Tone.js. No video files.
Which city should she drive through next? 👇
#Threejs #CreativeCoding #WebGL #NYC
OpenAI just backed mandatory outside safety evaluators for frontier AI labs.
The bipartisan FRONTIER Act would let independent verification organizations assess how leading AI models are developed and tested.
Turn your coding agents into research agents
Due to popular demand, we're happy to announce that a beta version of OpenResearch is now available for Windows users.
Can an AI bot go from idea to working product in 3 days?
Grok Bot Galaxy starts today, Sept. 15–17. Watch founders build live, explore role-specific workflows, and see integrations like Chief of Staff in action.
Free livestream: https://t.co/0nOiNZiqvZ
8:30 a.m.–6 p.m. PT
Meet Buzz. She's a fly chilling at a golden-hour beach, sipping beer, eating fries, biting burgers sunglasses on, umbrella overhead, not a care in the world.
Meanwhile, the Indus River the cradle of one of the world's oldest civilizations still flows today, reminding us that great things last.
Built with Indus(by @SarvamAI) Three.js + Tone.js. Procedural ocean waves, seagulls, wind chimes, and a fly who has a better beach day than you.
#Threejs #CreativeCoding #WebGL #Indus
Unconfirmed, but the rumor mill is loud: some Opus 5 traffic especially in Claude Code is already being routed to Opus 5.2.
If that’s real, the early read is a clear step up from Opus 5. Faster. Cleaner. Less of the stalling. Still not Fable 5.
What I actually want from the next drop is simpler than another benchmark chart:
- do the work without being asked twice
- stop padding every answer
- use fewer tokens to get to a better result
Anthropic has heard this for weeks. A release looks close. I’m hoping they treated the criticism as a product problem, not noise.
Elon’s entire engineering philosophy in one sentence:
“Physics is the law, and everything else is a recommendation”
"Physics is a harsh judge, and there's no fooling physics. If something's wrong, the rocket's going to explode. It's not going to get to orbit. So it's not like, "Elon, you're amazing," meanwhile the rockets are blowing up. It's hard to say you're really kicking ass here when the rockets are exploding.....that's just not the case
Generally, the rockets need to get to orbit, the satellites need to work, the Starlink connection needs to work, or bad things happen. This is sort of a physics situation, and physics is a harsh judge. I say physics is the law, and everything else is a recommendation. I've seen people break the laws made by humans, but I've not seen anyone break the laws made by physics. And rockets are ruled by physics"
@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.