Hm. Without having full context into your workflow, I couldn't pinpoint exactly where you can make optimizations.
I think building your own harness or fine-tuning a local model might be nuclear. GPT 5.5 is very good at in-context learning. So I would say fundamentally it's a mix between file context, codebase architecture w/ adequate documentation, AGENTS .md written hierarchically from top level -> subdir, and skills as context at runtime.
I know this sounds very repetitive because you get this exact response on all of your posts.
This will solve your agents-being-stupid problem. It'll maybe make things more smooth, so on average faster. But not the blazing speed you're looking for.
That would set up your agents for async work, e.g. gpt 5.5 xhigh. And then synchronous work gpt 5.5 low on /fast or opus 4.7 in fast mode with very detailed instructions—I'd recommend using text-to-speech like wispr flow or monologue so that your prompts are more contextually rich.
Even if your codebase is extremely well-written, which I don't doubt, it does not necessarily mean that it's agent-native— this is another axis of development that we need to spend time on when we use agents. It pays to spend time up front on writing lints, checks, hooks literally just for agents and not necessarily yourself.
At most, if you spend the time on tuning all of the configs, it'll take maybe a few days and you don't really have to tweak it much after that. Then async agents can work on a /goal for 72+ hrs just fine.
https://t.co/urDgf3kkd4
https://t.co/FDD5dvbeNO
Nvidia GPUs follow a pretty consistent pricing pattern generation-to-generation, adjusted for performance per dollar.
In the case of AI, GPU demand is extremely inelastic, so NVIDIA can effectively command any price it wants, particularly for their datacenter GPUs; the endless demand for AI means, in theory, demand for GPUs are always outpacing supply.
Ultimately, AI progress is still subject to research, development, and adoption, all of which still take time. GPU supply is not intentionally being throttled. Interesting theory though!
It feels like a level of abstraction above the tui, and the tui is already an abstraction above the code.
I prefer to see the tool calls in the way the CLI presents it because I can read the diffs very fast to steer the model when it does something weird.
I hardly read code or plans at this point, but watching it work allows me to come up with ideas that I otherwise wouldn’t have if I didn’t see what the model was doing.
I might steer it maybe once every few hours, if at all.
In theory your claim is likely the correct stance for human-made production lines. However, steering does not necessarily imply a misaligned trajectory.
Suppose you define the trajectory ahead of time; the human may have set a clear goal with strict verification gates for the agent, but the agent still falls into trajectory misalignment because of reward-hacking its solutions or inability to “see the bigger picture” of a problem, even if your verification gates and goal are perfect.
Surprisingly, we would assume that this is not an issue. Models are very advanced today. But they still have these jagged edges that Karpathy talks about.
Steering is necessary because today we still need humans as an oracle. Hopefully that will change with a sufficiently intelligent model, but unfortunately we are not there yet.
Symbolic is pretty cool imo. What I’ve found from practice is that as tasks get higher complexity over longer trajectories, you shift to a filesystem representation.
Inherently, the filesystem is great for reference handling and durable objects for agents. The LLM’s internal representation of the filesystem is also deeper baked into its training priors.
A virtual filesystem allows for a useful hybrid where you can have symbolic context with filesystem-like behavior embedded into the namespace.
Nonetheless, both work. Although, you might see models behave in less predictable ways in symbolic contexts.
The largest issue with symbolic context is that representation is less explicit. This is an extremely important property that you need to design around in agent systems. Once the agent’s workspace or language becomes less explicit, it struggles to predict outcomes. For example, planning a path of mutating the context state. The agent will struggle more in the symbolic context than filesystem.
Ultimately, this is a very fixable problem. You just need in-context learning and the agent will adapt to whichever system you choose.
Codex is great! Two major issues:
Codex does not play nicely with subagents—often it likes taking over a subagent’s work when it is “slow”, which is usually a couple of seconds.
Codex is also impatient in terms of waiting and often treats any arbitrary amount of time as “long”, e.g. it is not very good at polling.
These are both tough behaviors to correct. I use a system prompt to make it better, but not perfect; it definitely shows as a learned behavior more than anything.
@FlyaKiet@superset_sh I’m really looking forward to this. Terminal rendering in Superset has been driving me insane and forced me to do some work in iTerm2.
Superset is such a great product, excited to see it improve.
On VibeProxy, it is parsed from the model name, so you make separate entries per reasoning level in settings.json.
For example this is how you register GPT xhigh:
{
"model": "gpt-5.4-xhigh",
"id": "custom:gpt-5.4-xhigh",
"index": 0,
"baseUrl": "http://127.0.0.1:8321/v1",
"apiKey": "dummy",
"displayName": "GPT-5.4 xhigh (adapter)",
"noImageSupport": false,
"provider": "openai"
}
I've been following your work for a while and it's awesome seeing your progress, especially when you used LLMs much less compared to what you're able to accomplish with Codex 5.3 today.
How do you envision the adoption of Bend2 as a language for AI? As a replacement for existing programming languages? Or it is used to prove work with existing languages?