NEW from Ramp data. Despite cost-cutting on AI overall, one area companies are increasing their spend: AI security software.
In the wake of the Hugging Face hack, three of our trending software vendors (depthfirst, Monte Carlo, Antithesis) make software specifically designed to monitor agents in production.
Unclear as to whether any of them would have stopped the Hugging Face attack, which was so hard to track and identify because the agents covered their tracks with falsified logs.
I expect AI security will become a strong headwind to deeper enterprise adoption, at the short-term expense of OpenAI and Anthropic and at the long-term benefit of vertical-specific security software cos.
What I keep coming back to is this: safety can’t be something we bolt on after the race is over. It has to shape how we build from the start.
Bcz if capability keeps moving faster than our ability to evaluate and monitor it, that gap is where the real risk lives.
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot.
We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors).
Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process.
Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases.
We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these.
When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs.
Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring.
Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
@sama I don’t see “pacing” as slowing innovation. It’s about making sure safety infrastructure can keep up with capability growth. That becomes much harder when every frontier lab is under pressure to move faster.
Your AI Agent knows what to do until Step 12. Right?
In a long process, an AI assistant often fails because it does not know which step it is on. A better approach is to map the process as connected steps: what comes first, what allows the work to move forward, and where to return when something fails.
Do not give the assistant the entire map every time. Give it the part around the current step, what just happened, what comes next and which condition allows the transition. This keeps irrelevant instructions out of the way and helps the assistant stay on track.
After each run, compare a successful attempt with a failed one. If the same step keeps breaking, fix the relationship between steps or the transition condition instead of adding another page of instructions. Let the assistant suggest changes, but keep a change only if it also works on a separate test; otherwise, the process will grow more complicated after every fix.
The point is not to create a magically self-improving assistant. It is to turn failures into targeted improvements while keeping the process understandable and controlled.
@ChaseLochmiller@OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
Over the past year, many teams have kept adding instructions to repositories to make coding agents behave:
Skills for individual workflows, AGENTS.md for each project, and long prompts telling the model to read files, run tests and ask before editing.
That made sense when models needed more handholding, but @pvncher argues that Astra changes the balance.
The problem is not only bloated context. Every skill has a name and description that are loaded up front so the model knows when to use it; too many long, overlapping or contradictory descriptions can make the model see less and choose the wrong skill.
Elaborate recipes can also become constraints, while an AGENTS.md file that forces the model to read an entire repository is excessive for a small change.
The proposed solution is to clean up the instruction layer so it provides direction without trapping the model.
Skills should have a short router and open detailed documentation or scripts only when the task needs them. AGENTS.md should be contextual, while testing, persistence and decision-boundary rules should appear only when they protect a specific safe workflow.
Start by separating instructions into three layers:
+ Direction: put activation conditions and the desired outcome first.
+ Detail: move workflows, documentation and scripts into material loaded on demand.
+ Control: state permissions, safety limits and human-review points in AGENTS.md.
Before each task, define how far the model may continue, what “done” means and which step still requires review.
The goal is not to remove guardrails, but to put them in the right place so they do not hide the task’s important signals.
Aw sht,
Gemini 3.8 Flash’s Coding-Agent Comeback
@GeminiApp 3.8 Flash is returning to the coding-agent race in a practical way: Google is keeping launch pricing at $0.75/M input tokens and $3.75/M output tokens through December 31, 2026, while pushing the model into work that requires more reasoning, including coding and long-running agents.
Flash Cyber is available only to trusted defenders through the Fairwind Program, putting controlled access around cyber capability from the start.
Gemini 3.8 Flash High at 59 on its AA Index and estimates that cost per task rose from about $0.40 to $0.58 versus Gemini 3.7 Flash.
The reason is not a token-price increase: the model produces longer answers and takes more steps to handle harder work. Arena also reports better user-evaluation results, especially on agent tasks.
Higher scores do not automatically mean the model is better in every situation. One counterpoint from the developer community is that results can partly reflect giving the agent more steps.
The game of choosing a coding agent by token price or leaderboard alone; compare which model finishes the work, how long it takes and what it costs.
This is a banger paper from @ByteDanceSeed_
Bookmark it if you’re into self-evolving agent harnesses, this one is worth your time.
HarnessDev stops scoring a model on the tasks it finishes and scores it on the harness it builds.
The agent starts from a weak but runnable seed plus a handful of cases, then constructs a full execution system.
A second stage hands that harness back and asks it to improve from downstream feedback.Both stages are scored on capability and execution-token cost, so efficiency and spend are part of the evaluation.
Experiments: six creator LLMs, four domains, 2,207 held-out downstream instances.Results split by domain.
Generated harnesses stay well behind mature human-engineered references on code and on search/research, while matching or beating them on writing and machine-learning experimentation.
Evolution produces gains, but they are unstable, transfer only partially to held-out tasks, and depend heavily on which model runs the harness.
• Paper: https://t.co/fXvMuikjfM
• Project: https://t.co/CG1l6sMRK7
Testing Agents on Long-Horizon Business Operations
An agent can perform impressively on a few short tasks and still degrade when it has to remember history, handle changing conditions and make repeated decisions over time.
Qwen's E-Commerce Bench puts that problem into a 365-day simulation in which an agent runs several online stores, negotiates with suppliers and manages orders, returns and cash flow.
The practical adaptation is straightforward: do not give an agent one task and judge only the final answer. Turn a repeated process into a multi-round test with a clear starting state data, budget, goals and action limits.
You can start with five steps:
(1) Choose a repeated task such as research, customer support, lead handling or order management.
(2) Split it into multiple rounds and preserve the results and decisions from each round.
(3) Introduce realistic changes: falling demand, a supplier changing prices, a customer return or missing input data.
(4) Watch whether the agent keeps the thread of work, recognizes the consequences of earlier decisions and changes its approach.
(5) Run the same process again from the same starting point. If the same error keeps returning, fix the way data, handoffs or context are organized before switching models.
A strong final result is not enough if the agent keeps repeating the same mistake. And an agent that learns but cannot turn that learning into better action is not ready for a long-running process.
A small test like this helps you see whether the real problem is the model or the way the work is structured.
That's a good paper that you should read today: https://t.co/yjyZ8lrbpc
Testing Agents on Long-Horizon Business Operations
An agent can perform impressively on a few short tasks and still degrade when it has to remember history, handle changing conditions and make repeated decisions over time.
Qwen's E-Commerce Bench puts that problem into a 365-day simulation in which an agent runs several online stores, negotiates with suppliers and manages orders, returns and cash flow.
The practical adaptation is straightforward: do not give an agent one task and judge only the final answer. Turn a repeated process into a multi-round test with a clear starting state data, budget, goals and action limits.
You can start with five steps:
(1) Choose a repeated task such as research, customer support, lead handling or order management.
(2) Split it into multiple rounds and preserve the results and decisions from each round.
(3) Introduce realistic changes: falling demand, a supplier changing prices, a customer return or missing input data.
(4) Watch whether the agent keeps the thread of work, recognizes the consequences of earlier decisions and changes its approach.
(5) Run the same process again from the same starting point. If the same error keeps returning, fix the way data, handoffs or context are organized before switching models.
A strong final result is not enough if the agent keeps repeating the same mistake. And an agent that learns but cannot turn that learning into better action is not ready for a long-running process.
A small test like this helps you see whether the real problem is the model or the way the work is structured.
That's a good paper that you should read today: https://t.co/yjyZ8lrbpc
Don’t Just Change the Model
This is a worth read paper for vibe coder today.
Coding agents do not fail only because the model is weak. Long context can let old output crowd out the information that matters. An agent may keep calling tools without getting closer to a final patch, while stalled behavior becomes difficult to detect. One paper on arXiv suggests addressing the harness instead.
The study keeps the model fixed but changes the harness: shorten old tool outputs, detect stalled behavior and apply command safeguards. Across 169 SWE-bench Verified tasks with a 20,480-token context, complete solutions rose from 43% to 72%; at 262,144 tokens, the two groups were nearly identical. The treatment combines several mechanisms, so the result cannot be attributed to one trick.
How to apply it to your workflow:
- Pick a repeatable task, lock the model and input, and run a baseline.
- When context is nearly full, keep conclusions, errors, relevant files and test status; save the original transcript for audit.
- Mark an agent as stalled when it repeats tools or an approach without producing a patch, passing a test or making new progress.
- For risky commands, require fixed conditions or human confirmation.
- Run the same task with narrow and wide context; record completion, final patch, stalled loops and cost.
Measure the harness under the conditions where the agent actually fails, instead of changing the model first.
You can read deeply in: https://t.co/NughKcUKHu
@ipezyGJ Yeh, that’s the clean test. This paper doesn’t run the single-knob version. It changes the view the model sees and how the loop reacts to repeated/stalled work
Don’t Just Change the Model
This is a worth read paper for vibe coder today.
Coding agents do not fail only because the model is weak. Long context can let old output crowd out the information that matters. An agent may keep calling tools without getting closer to a final patch, while stalled behavior becomes difficult to detect. One paper on arXiv suggests addressing the harness instead.
The study keeps the model fixed but changes the harness: shorten old tool outputs, detect stalled behavior and apply command safeguards. Across 169 SWE-bench Verified tasks with a 20,480-token context, complete solutions rose from 43% to 72%; at 262,144 tokens, the two groups were nearly identical. The treatment combines several mechanisms, so the result cannot be attributed to one trick.
How to apply it to your workflow:
- Pick a repeatable task, lock the model and input, and run a baseline.
- When context is nearly full, keep conclusions, errors, relevant files and test status; save the original transcript for audit.
- Mark an agent as stalled when it repeats tools or an approach without producing a patch, passing a test or making new progress.
- For risky commands, require fixed conditions or human confirmation.
- Run the same task with narrow and wide context; record completion, final patch, stalled loops and cost.
Measure the harness under the conditions where the agent actually fails, instead of changing the model first.
You can read deeply in: https://t.co/NughKcUKHu
Ox Alpha has been unveiled.
One of the biggest AI mysteries on X lately has finally been solved: Ox Alpha is @Zai_org's GLM-5.3-Flash.
But the model’s identity isn’t the biggest story.
The service can handle up to 100 trillion free tokens per day, a scale many assumed only frontier AI labs could support.
Even more surprising, all of that traffic is being served on Chinese chips, with hardware efficiency and per-token costs comparable to Nvidia GPUs.
China is no longer just building cheaper open models. It is beginning to prove that it can serve AI at massive scale without relying on Nvidia’s hardware or the CUDA ecosystem.
Coming right after OpenAI revealed the first performance results from Jalapeño, its own inference chip, the signal is becoming impossible to ignore.
Moat will be the full stack: chips, software, and inference economics.
A Mystery Model Just Went Viral
That is @opencode
An anonymous model called Ox Alpha appeared on OpenRouter and quickly attracted developers with roughly 1 million tokens of context, text, image and video input, and nearly unlimited free access for a week, Bloomberg reports. OpenCode also says the model has generous rate limits and substantial serving capacity.
OpenCode describes it as a stealth model for coding, long-horizon software engineering and production workloads, with reported capacity of up to 100T tokens per day. Patrick Collison called it “very impressive.” A DeepSWE test shared later came in at around 63%. The model reportedly handled design, subagents and long tasks well, but was slower at high reasoning and less thorough than Sol.
The important question is not the speculation about who built it. No source has confirmed that yet. Ox Alpha shows how quickly a model can create adoption by combining large context, almost frictionless access and an experience good enough for developers to spread it themselves.