nice to see a real benchmark instead of vibes. same 1,000 messages, accuracy basically tied, so the choice comes down to cost vs latency. for most support flows 2.8x cheaper wins. worth reading the full breakdown.
i benchmarked Jev vs OpenAI Decisions API on the same 1,000 customer messages.
nearly tied on accuracy, Jev is ~2.8× cheaper, OpenAI is ~1.4× faster.
full breakdown on my blog: https://t.co/KKv7ohYbu5
saving this for the weekend. the harness and context engineering parts are what i'm most curious about. picking up right where last weekend's harness reading left off.
How do always-on, proactive agents like @Bot work?
Watch my lecture at Stanford CS146S:
2:35 How we got here
6:13 What changed in the models
10:06 Inside an always-on agent
18:07 The harness
27:26 Context engineering
34:26 Where this is going
saving this for the weekend. the harness and context engineering parts are what i'm most curious about. picking up right where last weekend's harness reading left off.
How do always-on, proactive agents like @Bot work?
Watch my lecture at Stanford CS146S:
2:35 How we got here
6:13 What changed in the models
10:06 Inside an always-on agent
18:07 The harness
27:26 Context engineering
34:26 Where this is going
the takeaway isn't skip tests. an agent testing its own reading of the spec mostly confirms itself. the check has to come from outside the loop. a spec, a reviewer, a real user flow.
ok it's official... AI-written unit tests and integration tests are empirically proven to be unhelpful. you should tell your agents to stop writing tests by themselves
on the deepswe eval set, banning sonnet 5.5 high from writing any tests actually resulted in slightly higher success rate (non stat-sig), with less time and token spent (stat-sig). this is pretty hard empirical evidence
across the tests that were written by the "tests allowed" baseline arm, 65% of them were unit tests, 35% were integration tests, neither bucket resulted in any improvement compared to not writing any tests at all
i also picked a random sample of 44 tasks subset where i completely disabled executing even existing tests - it also did not affect success rate at all
i spot checked many tests written in the baseline arm, and my intuition is that most tests are simply a repetition of the implementation
agent-written tests do not add any value because both the implementation and the tests were simply the agent's interpretation of our intent. the tests aren't any more accurate than the implementation itself
i suspect we can still extract some value from unit tests and integration tests if we describe them ourselves when we believe we can articulate our intent better through test cases than through requirements. but i have not proven this yet
also worth noting, during deepswe eval the agent wrote almost no e2e tests (only 17, compared to 3000+ unit/integration tests written). so this analysis does not prove nor disprove the value of e2e tests - i will do another eval specifically for that
@mattpocockuk the /pr evidence bit is the part that matters. a screenshot or a short "here's how i checked it" note beats a long description every time. less back and forth, clearer merge risk.
@fatih same thing on the school run. my daughter asks for a song, i'm driving, and siri just won't play it. i end up waiting for a red light, or pulling over with the hazards on, just to get it started.
spending the weekend on harness engineering. this looks like the cleanest starting point i've seen. not about making the model smarter. a closed loop so it keeps context, verifies the work, and doesn't call it done too early.
@housecor the specialization reason is fading, sure. but teams weren't only there because code was slow. someone still has to own the product call, the taste, and what not to ship. one person per slice can work. calling that the end of teams skips the part that was never about syntax.
knowing yourself matters. the other day, talking with my brother, we saw how different we are. he makes a coffee slowly and actually enjoys it. i treat that as wasted time, hit the machine, and get back to the computer. there's no right or wrong. go with what fits you and what you enjoy. same with role models. don't copy them as the correct version. look for what you can actually add to yourself.
@tylergibbs yeah claude still feels ahead on UI taste, that's fair. the annoying part is the sub. a multi-model harness only helps if you can actually bring the model you're already paying for.
video turing test, 48% pass rate, first HIM… the demos keep leveling up and somehow i feel less of that early wow. maybe the bar moved, or maybe i just got used to the pace.
Introducing Griffin, the first model to pass the video Turing test.
48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
It’s the first Human Interaction Model (HIM).
@antirez this lands. AI for shipping speed, hand-written practice for depth. the trap is treating generated code as the training ground instead of using the time it frees for the hard stuff.
günün sonunda ortada bir ürün var. onu üreten şirket de kullanan kişi de kodun veya ürünün nasıl yapıldığını çoğu zaman o kadar önemsemiyor. diploma / alaylı tartışmasına girmeden söylemek gerekirse, asıl farkı AI’ın getirdiği verimlilik yaratıyor. koda hakim olmak ve kod detayları geliştiricinin işi; şirket ve kullanıcı ürünü bekliyor. bu ayrım geliştirici tarafında bence hâlâ net değil.
@resat_dev katılmıyorum. clean code tartışması sert olabilir ama bu üslup eleştiri değil, aşağılama gibi duruyor. tartışmayı fikir üzerinden yürütmek varken kişiyi küçültmek hem meselenin özünü kaçırıyor hem de kimseyi ikna etmiyor.
@SSShken i usually get solid results on medium too. for me the difference is documenting as I go, so the plan doesn't have to live only in the model's head across steps. that keeps longer multi-file work steadier without bumping effort.
opus 5.5 is the best model I've used so far. even on medium effort it works really well. solid on speed and tokens too. past few months I've been bouncing between claude and codex but looks like I'm sticking with claude for a while unless they make another absurd call on their side.
jason fried spent years looking for a writing tool that matched how he thinks, then built it himself. same thing in product: design for everyone and it fits no one. a small tool shaped around a real workflow often beats the generic suite.
For years, maybe a decade, I've wanted to build a writing tool that just did a few things I couldn't find anywhere else.
A system that mimicked how I think when I write. A system that provided clear affordances, features, and places to do the things I wanted to do. A way to play with my writing as I'm writing it. A way to settle into just the right way to say something.
It never made sense to make it a 37signals product because it's not commercially viable, nor would it be worth pulling people off their other work to hack on this. Plus, it would have taken months the old way.
So this weekend I just made it myself, with Claude's help.
It's called Write_On and the short video walks you through it.
Essentially it's "alternative control" at the word, sentence, and paragraph level, plus a way to dim stuff back, and stash stuff near by. Not version control, but alternative control. You'll see what I mean when you watch the video below.
So here it is. Write_On.
@okandavut_ ben de paylaşımından görüp bu hafta sonu araştırmayı planlıyorum. senin yorumlarını da merak ediyorum şu anda çok yeni bir alan olduğu çok fazla farklı yöntem var ama kısa bir göz gezdirdikten sonra bu benim de aklıma yatan en iyi yöntem gibi.