founder @ synthhaven
Building the orchestration layer for AI media. ex-interpretability @ Google Brain PAIR(rip)
OSS type shit + building software factories
Started to add grafana boards and otel/RUM to pretty much every development env now. Makes it so much easier to spot issues and get a good feedback cycle early on the process for your agents.
funny thing, but if you want to find jank your coding agents are doing or potential nonsense, just grep your past sessions for the frequency of the word "honest" in assistant messages and track writes to /tmp
seems dumb, but it's a great proxy metric.
Playing around with the deepseek harness. Got it to run the whole night doing a bunch of different k8s maintenance and debugging stuff.
I'm very impressed at how surprisingly good it is off the box at handling long horizon boring tasks,and their setup also makes it easy to check background processes, calls made and reconstructing all the things that happened in the sessions.
Also very pleased at how effective it is at searching other past sessions and dissecting things and patterns.
@arakharazian Not surprised at all. Fable is fun at first, but the code can quickly become unmaintainable hell. Codex just offers a much better developer experience.
As your codebase becomes enormous and more complex, the cost and frustration of mistakes and jank exponentiate.
I've been really enjoying the deepseek harness for goal focused tasks. The ability to track the exact background tasks and context used is really useful. As well as supporting session search on workspace by default. Great implementer overall
Still using Pi for more complex tasks tho
Don't sleep on https://t.co/Nd6q7i5zA0 for @pidotdev though.
Makes a lot of compaction summarization bullshit go away and I found it lets you run really long sessions of agents that delegate to agents while also preemptively tagging issues that emerge on your process or jank
Deepseek flash really feels a lot like that really special moment when opus 4.5/4.6 + @steipete openclaw gateway planes + @pidotdev innovation really blew up and all of a sudden it became really, really easy to mesh and connect code together and really fast to build.
As in, it's so cheap and so fast, yet so good at interacting with your codebase that it makes a lot of shit possible.
With the huge bonus that it is also incredibly good at building good software, dealing with delegating, and using your existing verification loops and processes
Overall, that's the really interesting trend, as video and media generation moves away from prompt render, and you start having extremely powerful context engineering based capabilities.
Those build on top of one another, enabling even more complex applications.
Minimax H3 is lowkey amazing for complex user interaction design, specially if you use the right game design terms.
Makes sense, since model is omnimodal.
Can also just pull up a tablet pen and draw notes on some prototypes and H3 can animate. Then feed to deepseek and build
@Suhail Hard to trust a company that openly admits that they Gold Experience Requiem you away if you're working on AI/ML etc by "steering" your agent away.
It's complete insanity that shit like this flew, specially since the models' been enshittified and overpriced to death.
Seriously, give @skypilot_org a chance. Extreme versatile and useful tool if you do heavy agentic coding, but also super useful for serving your own models, running some large scale map reduce style model calls, iterate on kernel work etc,etc,etc
So. Versatile.
It's beautiful
@deepseek_ai Really obssesed with this right now. Really want to give a shot at connecting this with @skypilot_org . The plugin system of this harness paired with the "invertible" nature of sessions + the off the box event bus feels like a perfect fit for mass coding tasks a la @_lopopolo
@ltx_io@reactorworld Beast. LTX is the most fun local video model to use, specially around how good it is for in context tasks. LTX really spoils you and makes any gen that takes longer than a few seconds unbearable.
Y'all are awesome!