"It's super, super easy for people to mistake fluency, or some kind of cognitive cosplay, for epistemic grounding."
Two lawyers who left the law to build legal AI on context rot, the verification paradox, and why "not found" is the hardest answer to trust.
Link below 👇🏻
"We were implementing a system that had no precedent in the AI industry."
@nitmusai and the engineers behind @GenieAI 3.0 on state machines, agentic evals, and shipping a Cowork-shaped product before Cowork existed.
EP5 of Out of the Bottle is out now 🎧
Link below 👇
I don’t know if you feel the same, but Anthropic’s rate of innovation has currently plateaued. The amount of crazy stuff they did in the last 6-8 months has slowed down, and others have woken up.
It feels like every few months one of the leading labs comes up with a breakthrough, and then everyone rallies behind it and plays catch-up. Then some big, new, different thing drops, and that creates a whole new wave of upgrades.
But I am really enjoying the thriving ecosystem that has built up, one that is extremely innovation-led and blazing towards unlimited progress.
@swyx@candidosales Don’t see direct problem with worktrees.
1. It is an issue with node modules dependency trees
2. It is an issue with agent not maintaining hygiene once the work is done.
I believe from the point Open AI broke their news it has all turned into a PR thing to show their model is superior and lethal.
I am very certain that labs are doing this on purpose to market their models to showcase its power.
And it is working.
@GergelyOrosz same dynamic is playing out in legal. once AI output crosses a quality threshold reviewing every clause creates more cognitive load than it's worth. the question of what replaces review is the right one. for code it's evals. for legal it's scenario generation
@rauchg the shift most teams haven't made yet: stop treating agents as a one time prompt and start building the factory. once you do that compounding kicks in
@emollick the guide needing an update in 48h is the feature not a bug. we're not in a steady state. that chatbot vs agentic systems frame in your diagram is exactly the right distinction to anchor on right now
@__tosh the boilerplate is going to zero. what remains is the design of the tool surface. 9 lines today, and the next version of this is the model designing its own tools
@omarsar0 orchestration > model choice. once routing for cost vs intelligence is dynamic, the orchestration layer becomes the actual IP. models commoditize fast, the routing logic is what compounds
@dair_ai harness choice as a training variable is the insight people keep undervaluing. everyone's optimizing the model while the execution environment shapes outcomes just as much
I support open-source models distilling what commercial companies distilled for free from the entire internet. I published distillation for free in 1991 in Europe - this was copied in the US and in China (https://t.co/mddh8XmfAs)