Today, we're releasing shadcn/cli v4. It packs a ton of features: shadcn/skills, presets, dry-run, monorepo and more.
If you're using shadcn/ui with coding agents or need better control over the defaults, this is for you.
Here's everything new:
The Human Verification Paradox
The most "advanced" AI agents in production don't trust themselves.
74% rely on humans for final verification.
Here's why human-in-the-loop isn't a bug—it's the business model 🧵
The uncomfortable truth from 306 deployed agents:
❌ Public benchmarks: rarely used
❌ Automated CI/CD: breaks on non-determinism
✅ Human review: 74% of systems
✅ LLM-as-a-judge: 52% (but ALWAYS paired with humans)
Automation's dirty secret? It's supervised.
Why can't we automate evaluation?
Because production tasks are bespoke:
Insurance claim analysis → no public dataset
Internal HR policy lookup → company-specific
Client onboarding flow → unique per business
One team spent 6 months building a 100-example benchmark FROM SCRATCH.
The "Golden Set" pattern we saw everywhere:
Step 1: Deploy agent with 100% human review
Step 2: Domain experts label "golden" outputs
Step 3: Build LLM judge on golden set
Step 4: Judge filters 95%, humans verify 5%
This is the actual production stack. Not sexy, but it works.
Real case study: Site Reliability Agent
What it does:
Analyzes system failures
Proposes fixes
What it DOESN'T do:
Execute fixes automatically
Why? Because when the agent is wrong, the cost is downtime.
Human SRE = final approver. Always.
The verification hierarchy in practice:
Tier 1 (automatic pass):
LLM judge scores >90% confidence
Tier 2 (human spot-check):
5% random sample of Tier 1
Tier 3 (mandatory review):
Judge scores 70-90%
Tier 4 (rejected):
<70% → rerouted to human from start
Interview quote that stuck with me:
"We tried running it fully autonomous for a week. Three clients called to complain. Now we have a nurse review every output. Zero complaints in 6 months."
The ROI isn't eliminating humans.
It's making each human 10x more productive.
Here's the genius part:
Agents aren't REPLACING workers.
They're TRIAGING work.
Medical agent example:
Agent handles 80% of routine pre-auth
Routes 20% complex cases to specialists
Specialists focus ONLY on hard problems
Net effect: 5x throughput, same headcount.
Why "LLM-as-a-judge" alone fails:
❌ Judges inherit model biases
❌ No ground truth for novel tasks
❌ Adversarial inputs fool judges
❌ Drift when base model updates
Every team using judges ALSO uses:
→ Human calibration
→ Regular spot-checks
→ Feedback loops
The evaluation gap researchers miss:
Academia: "We achieved 94% on SWE-bench"
Industry: "Cool, but our task isn't on SWE-bench"
75% of deployed agents evaluate WITHOUT formal benchmarks.
They use:
A/B tests
User satisfaction scores
Business metrics (time saved, errors reduced)
What this means for your agent:
If you can't define "correct" programmatically, you need humans.
But here's the trick:
Don't review EVERY output
Review SAMPLES
Let the agent learn the review criteria
Gradually reduce review rate
Closer:
The future of agents isn't "no humans."
It's "humans in the right place, at the right time."
From the MAP study: 86 deployed systems, 26 domains.
How do YOU verify your agent's outputs? 👇
https://t.co/kfAoTRoKU7
A service that builds a marketplace connecting consumer-grade GPU owners with users who need GPU capacity for workloads such as AI. It brings a large supply of GPUs to the market at pretty low rates.
https://t.co/k8umHjujRj was scheduled to be discontinued on March 5, 2026, following its acquisition by OpenAi. A few thoughts on developing its replacement:
https://t.co/1EUcBUSrRm
How often developers accept code changes made by AI?
Research says:
We would have expected that less experienced developers tend to use and accept agent at higher rates: it seems like the opposite is true!
https://t.co/9Uh1aD0UVg
If you're concerned about AI, the answer is straightforward: just invest in yourself. It really is that simple. This year is a particularly challenging time to be asleep at the wheel when it comes to personal development.
https://t.co/9nt5YpnVP1
@ShitcoinSherpa It is not a smart contract. It is just usual wallet account. Just race who pay more tx fee. For ex I got some wei from address: https://t.co/GAXzZn0ztf