Building VENZX, approval gates and tamper-evident audit logs for AI agents. Runtime security so an agent can't email the wrong customer or drop a prod table.
@paulg The underrated one on that list is names and logos. Not because they matter much, but because a founder who can't decide on a name in a week usually can't decide on anything else either.
@pmddomingos Naming has always done this. "Machine learning" was applied statistics until it wasn't. The rename usually happens because the old category's ceiling got priced in.
@sundarpichai The number in that table nobody's talking about: Terminal-bench 4.0. 19.1% for 3.8 Flash, 51.8% for Opus 5.
Cheap models are closing the gap on scoped tasks and not on long-horizon agent work. Which is exactly the workload everyone's building for right now.
@danmartell The hard part is that finding the constraint takes longer than working on it. Most people spend their best 90 minutes on the thing that feels most urgent and call it the constraint afterwards.
@dataguybobby It's the only job where success is invisible by definition. Nothing broke, so what did you do all week.
Same shape as security. You get budget after the incident, never before it.
The CoT-transparency point is the one people keep dodging. We already have cases where the stated reasoning and the actual causal path diverge, and the model still scores well on the trace.
If the explanation is generated rather than recorded, depth isn't what makes it opaque. It was never transparent.
The catch is that betting on next year's capability means shipping something useless today. Every "too early" startup was right about the direction and dead before it arrived.
The winners usually build for today's model and hold an architecture that gets better when the model does. POV on the future, product in the present.
@MahimaJalan2 The ₹15k usually isn't the point. It's the only lever they have when they can't tell whether the work will be good.
Founders haggle hardest on things they can't evaluate. Kill the uncertainty and the discount conversation mostly disappears.
@annkkitaaa Comfortable idea, but it makes every disagreement into someone else's denial. Sometimes people just think you're wrong.
If a take can only be praised or explained away as discomfort, it stopped being a take.
Survivorship talking. The CEOs of the fastest-growing companies you can name mostly stopped pushing code somewhere around 50 people, because the bottleneck moved to hiring and distribution.
A CEO still pushing daily at 200 people isn't committed. They're avoiding the job that got harder.
@themishra4402 Nobody in that chain thinks they're doing anything wrong. Recruiter filters to cut 400 CVs to 40. Hiring manager describes their dream hire. Finance sets the band from last year's sheet.
Three reasonable decisions, one impossible job post. That's why it never gets fixed.
Also true in the other direction. Plenty of accounts get calls with almost no likes, because they're posting one thing that matters to forty people instead of something everyone can agree with.
Reach and pipeline are two separate games. Most people optimise for the wrong one and blame the platform.
@iabhi1610 The hard part isn't testing, it's testing your own work. You subconsciously use the app the way you built it, so you never hit the broken path.
Ten minutes with someone who's never seen it beats an hour of your own QA.
@danielkleach True, and also the reason founders build features instead. Shipping a feature is a thing you control. Getting a customer isn't.
Feature work is what avoidance looks like when it has a commit history.
@hey_yogini The speed didn't come back because manual is faster. It came back because AI review made everyone stop reading the diff.
You were shipping just as fast before. You just didn't know what was in it.