I’m significantly older than you. I started coding in the late 60s. My current strategy is to not read any of the code written by my agents. That’s the only way I can take advantage of their productivity. What I do instead is to surround the agents with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, test coverage, and a plethora of others. In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.
- You: "My stocks went up 10%."
- Government: "Congratulations! That's taxable."
- You: "But inflation was 8%. I only really gained 2%."
- Government: "We tax the full 10%."
- You: "You tax the inflation you created?"
- Government: "Capital gains are capital gains."
- You: "So you profit from debasing my money twice?"
- Government: "You're not an economist."
Good morning, Asia. While you were sleeping, one of our most-read stories was about Google burning through cash for the first time since going public decades ago as it raised its spending forecast for AI infrastructure. https://t.co/mYimix2GWy
We had a team of agents rebuild SQLite from its 835-page manual.
It created a replica in Rust which passed 100% of a held-out test suite.
Interestingly, cost varied 15x depending on which model mix we used.
grok 4.5 made me give grok build a serious run today
here's my honest first impression (non affiliated neutral view point):
1. grok build is a very good harness
firstmate stretches harness capabilities to their limits, and i've been testing it with claude code, codex, opencode, pi and grok build
so far, grok build and claude code are the only two harnesses that can automatically wake up when the background polling process finishes
codex hard fails on this kind of background polling task, with no escape hatch. opencode and pi can both do it with custom plugins, but not out of the box
grok build also feels really clean, smooth, and transparent. you can see what background tasks are running, what hooks got triggered at what step, context window, token usage etc all out of the box and somehow the UI does not feel cluttered at all
2. grok 4.5 is a very good model
i've been using opus 4.8 as my primary firstmate, and today i did a full switch to grok 4.5. so far i don't think anything degraded, while token throughput is a lot faster, although time-to-first-byte for each response seems long - i wonder if prompt caching is done properly or no
3. the quota that comes with X premium is quite generous, and there's no session level limit
so overall i'm quite pleased by this and plan to switch a lot of my tasks to grok. the landscape just got a lot more interesting...
In Brussels for the first time. My hotel tells me I can't control the A/C because of the environment. They promise they will keep it at 26C for me the whole time.
We are staying for only 3 days, so they will not clean my room or change my towels. I have to take the trash out of the room myself.
This is so bizarre.
I want to find a way to make sense of this, but it's hard not to make the cynical observation that the hotel is benefiting at the guests' expense.
@big_duca Someone has to prompt the Claudes, talk to customers, coordinate with other teams, decide what to build next. Engineering is changing and great engineers are more important than ever.
Do you think we, programmers, may lose our jobs in the era of AI coding agents? No, we won’t. Instead, there will be more jobs for us, because software will become much cheaper and easier to replace. We won’t upgrade software; we will just buy new ones—just like we do with shoes and skirts. The world will need many AI operators, formerly known as programmers. I’ve just published a new blog post about this: https://t.co/YclwT9BUWd
A guy just used @AnthropicAI Claude to turn a $195,000 hospital bill into $33,000.
Not with a lawyer. Not with a hospital admin insider.
With a $20/month Claude Plus subscription.
He uploaded the itemized bill. Claude spotted duplicate procedure codes, illegal “double billing,” and charges that Medicare rules explicitly forbid. Then it helped him write a letter citing every violation.
The hospital dropped their demand by 83%.
This isn’t just a feel-good story. It’s a preview of what AI will really do next: flatten systems built on opacity.
Hospitals, insurance companies, legal firms—all rely on asymmetry. They win because you don’t have access to the same data, code books, or language.
Claude gave one person the same leverage as a compliance department. That’s a revolution.
We thought AI would replace jobs. Turns out, it’s replacing excuses.
I've been in crypto for over 10 years and I’ve Never been hacked. Perfect OpSec record.
Yesterday, my wallet was drained by a malicious @cursor_ai extension for the first time.
If it can happen to me, it can happen to you. Here’s a full breakdown. 🧵👇