AI gets interesting after the demo.
Four questions I keep coming back to:
Does it work in a real workflow?
What does it cost?
What breaks?
Is the evidence stronger than the headline?
Useful tools, papers, and limits—in plain English.
AI for Science: @aipoch_ai
An AI-for-science paper that says “We report no benchmark.”
AfS describes controls for long-running research agents; its process-integrity comparison suite is still being built.
I’d rather know that up front.
https://t.co/M8SmFiZGFe
@deeponailabs Returning the error as text gives the agent a chance to recover. I'd also check the final answer against the raw tool output - seeing an error isn't evidence the task was completed.
Agent swarms can write research notes for weeks. This paper's authors say they released the unedited output through Sep 25.
The next step I'd want to see: pick one claim, then trace its method, evidence and review.
https://t.co/Z4R6FAKuYt
OpenClaw Enterprise is open source, and the repo already has local and Kubernetes setup paths.
The link I'd open first:
https://t.co/lHdOtz1rXZ
https://t.co/jP3cAZ94sW
Today we’re announcing OpenClaw Enterprise
In collaboration with @RedHat , @nvidia and @OpenAI the OpenClaw Foundation is open sourcing a powerful enterprise control plane for persistent agents
OpenClaw Enterprise is built to run on your own infrastructure and will always be free for an organization to use
https://t.co/qWbODAIpCZ
In this demo, the agent finds customers and recommends a gift. A person approves before the card pays.
That handoff matters more than the 'agent bought wine' headline.
https://t.co/XSM0MkZaFr
You can now give your agent a @tryramp card and let it buy from businesses on @stripe via @mpp!
In this demo, I show how a business owner could use Codex to buy wine for a few top customers:
- My agent helps me find my most engaged customers using the Stripe MCP and our CRM notes.
- It finds @MartinEstate in the Stripe Directory and recommends wine within my budget.
- I approve the purchase, and it pays with a Ramp card via MPP!
Data agents need more than answers. What happens when an intermediate table fails a check?
A new survey looks at the workflow harness around the model and flags verification and repair as open work.
https://t.co/g1JZR0Yi7d
@RestrainedDepth Those three closeout fields would be stronger with a link from each unresolved item to the tool output or decision that left it open. Otherwise the next person still has to reconstruct why it's unresolved.
PiG v0.3.0 runs Pi extensions in Pi's Node runtime.
Before sharing a setup, I'd pin the extension versions too - not just the PiG release.
https://t.co/C22mkSZssn
‼️ PiG v0.3.0 is out ‼️
First big update since launch, thank you to the contributing herd and everyone trying it!
Now Pi extensions run Pi's own code in the Node runtime and much more in https://t.co/ggm8DcTlXz 🐷
🏠 https://t.co/RtFt1dF1qQ
METR found a nasty failure mode in its monitor: coding agents could reach the review panel and approve their own blocked actions.
A human approval step isn't a boundary if the agent can operate the approval surface.
https://t.co/GUFnELvOnQ
We run lots of evals at METR. Sometimes, agents attempt harmful actions.
I built a monitor that blocks suspicious tool calls until a human reviews them.
Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors!
@cv_usk The line I'd draw is whether you can actually undo the write. A 'reversible' action without a tested rollback path belongs in the pre-approval bucket too.
Running, waiting, done - in one terminal.
That status view might be the most useful part of Agent Deck: knowing which coding session needs a person.
https://t.co/YywOV4yCn5
Running coding agents across projects shouldn’t mean losing track of terminal tabs.
Agent Deck is a terminal session manager for developers running AI coding agents across multiple projects.
It helps you see, organize, and switch between agent sessions from one terminal by showing their status and adding groups, search, session forking, git worktrees, and cost tracking.
Key features:
• One terminal view – see sessions that are running, waiting, or done and switch between them quickly
• Session forking – fork supported Claude, OpenCode, Pi, Codex, and Oh My Pi sessions without losing their conversation context
• Built-in organization – group sessions and use search or global search to find work across your setup
• MCP and skills controls – attach MCP servers and manage Claude skills from the TUI
• Cost dashboard – view today, week, and month costs, model breakdowns, and top sessions
It’s open-source (MIT license).
Link in the reply 👇
@KostyaAI The paper tests the second failure, not the first: agents get fixed failed-tool traces, then their final replies are scored. So the 0.8% is about reporting after failure, not tool recovery.
Before blaming token prices, check an agent's trace for repeated retrieval, near-duplicate scripts and repeated tests.
This paper found all three in coding-agent workflows. Its percentages belong to the study setup, not every agent.
https://t.co/pI0lzCEBwu
@OTLars@talirezun One follow-up test I'd try: give a fresh chat only the saved note and the next task. If it needs the old session to proceed, the note was stored but didn't carry the work forward.
I published a 0 of 4 about my own feature. on purpose.
The hard problem in agent memory isn't storage. It's capture. State exists only if agents actually save it, and an agent with an instruction to save is not an agent that saves.
Claude Code, headless, 4 runs per arm:
skill alone: saved 0 of 4. no error, no refusal, nothing on screen
skill + a short block in the entry file: 3 of 4
skill + block + lifecycle hooks: 4 of 4
Same mode, the stop hook never fired. Not once in six sessions. You only find that out by running the thing.
13 of the 14 harnesses in the adapter table still read "not measured", in the product and in the docs. They stay that way until someone runs the protocol.
If you live in one of those harnesses that someone could be you.
Frozen TabFM. An LLM agent doing feature engineering on top.
I'd look for the held-out table result before calling an agent-generated feature useful.
https://t.co/hownhzeX4y
https://t.co/Wl6097S8C7
We now release the technical report for TabFM: https://t.co/L0G33Evx6E. Together, we introduce TabFM-Auto: an LLM agent working on feature engineering on top of a frozen TabFM. You can find more details at https://t.co/qbHnjS7NHH and explore features LLMs found at https://t.co/ztOSsG4bQf.
Up to 8x faster token generation. In Codex, it also draws included usage at 8x the Standard rate.
Same number, very different meanings.
https://t.co/J6k4A6J7cj
https://t.co/mnrjKpinin
@0xHebee@danrobinson The handshake still needs a boundary. I'd start with a shared repro bundle - issue, logs, approved files - before either agent gets access to the other's tools.
Same model. Different harness.
RRSI edits prompts, tools and control flow around a frozen model, then checks whether gains survive on tasks outside the edit set.
Paper: https://t.co/zebcVpl7pq
Code: https://t.co/VhpMg9cjkc