@nateparrott After working with it for some time, I see one improvement that can help a lot - something like @storybookjs for Design system preview and documentation. It can hugely improve how CD previews the components. I just miss the state management and docs.
I’m not sure how other people read the options, but for me, AI-generated options sound a bit “sharp” and “truncated often”. AI gives good explanations, but the flow suffers.
We made a blind taste test to see whether NYT readers prefer human writing or AI writing.
86,000 people have taken it so far, and the results are fascinating. Overall, 54% of quiz-takers prefer AI. A real moment!
https://t.co/Gpbr3TAiiI
🚨BREAKING: Alibaba tested AI coding agents on 100 real codebases, spanning 233 days each.
the agents failed spectacularly.
turns out passing tests once is easy. maintaining code for 8 months without breaking everything is where AI collapses.
SWE-CI is the first benchmark that measures long-term code maintenance instead of one-shot bug fixes.
each task tracks 71 consecutive commits of real evolution.
75% of AI models break previously working code during maintenance.
only Claude Opus 4 stays above 50% zero-regression rate. every other model accumulates technical debt that compounds over iterations.
here's the brutal part:
- HumanEval and SWE-bench measure "does it work right now"
- SWE-CI measures "does it still work after 6 months of changes"
agents optimized for snapshot testing write brittle code that passes tests today but becomes unmaintainable tomorrow.
Alibaba built EvoScore to weight later iterations heavier than early ones. agents that sacrifice code quality for quick wins get punished when consequences compound.
the AI coding narrative just got more honest: most models can write code. almost none can maintain it.
AI didn't remove my job. It replaced the part I was good at with the part nobody wants to do.
New post: The AI Triangle — why money, time, and effort constrain AI far more than the hype suggests.
https://t.co/rxiPXg6Agw
Wrote about what it would take to fix this — faster reporting, standardized data, and why the US is behind Australia and the UK on nonprofit transparency.
https://t.co/lXUZwudGTZ
It feels surreal that a maximalist prediction piece by @Citrini7 is getting serious, heated responses. Since when do we treat long-range economic fan fiction as actionable analysis?
Wrote about this in more detail — from IRS data chaos to why the graveyard of "Palantir for the rest of us" is full of technically sound products. https://t.co/ahK8br2T4H
The biggest data company in the world (Palantir) basically does three things: collect messy data, clean it up, build dashboards. Why hasn't anyone done this for small businesses?
Can LLMs change this? They're genuinely good at normalizing chaotic data into a clean structure. If that cost drops, "Palantir for SMBs" might finally work. Or maybe distribution is still the real moat.
@mikulaja What I'm surprised about is that they didn't face any significant issues with anti-fraud measures during the PPP program in 2020-2021. I believe SBA even issued guidance to banks to stop funding any loans to GD routing numbers.
🚨 Cedar Park first responders are working a large brush fire near 12820 W. Parmer Ln. The nature of this fire is requiring some evacuations at the apartment complex. Please avoid the area if possible. Updates to come as able.
“Product value is the benefit that a customer gets by using a product to satisfy her needs minus associated costs. Complexity is the effort associated with delivering such a product to the customer.” — @productboard https://t.co/s7oEzdoXgU