Most Claude Code content is demos.
A to-do app. A Twitter clone. "Look what I built in ten minutes."
I built my own operating system with it. 33 tables. An agent runtime that spawns claude -p behind a native MCP server. Slack on socket mode. 291 tests across 62 files. It runs on my laptop and I open it every morning.
It took months, not ten minutes, and most of what I learned came from things breaking in ways nobody has written down.
That's what I'm going to post here.
No team. No CTO. One person and a laptop.
the failure modes are the part i'd read first.
anthropic opened a docs home for people building with claude: engineering deep dives, claude code and api guides, from the teams building it. i'll read them.
but i've built a 33-table system, an agent runtime that spawns claude -p behind an mcp server, and a client's migration of 2,490 products, none of it with a developer's training. every expensive thing i hit was environmental rather than api. a working directory. two env vars the child inherited. a flag i set because it sounded like the careful choice.
not one of those three raised an error.
https://t.co/sRyO7OigQ7 is our new home for developers building with Claude.
You'll find engineering deep dives, Claude Code and API guides, tips from the teams building Claude, and some fun easter eggs.
i did not choose the model my agents are running today.
claude code 2.1.284 made sonnet 5.5 the default sonnet. my runtime spawns claude -p behind an mcp server, and the children come up on whatever the default is.
so the model underneath my own system moved because of a line in someone else's changelog.
three things i went to check in my repo, none of which are in the release notes:
whether i pin a model anywhere or just inherit it
whether the 291 tests still pass on the new default
whether a spawned child ends up on the same model as the session that spawned it
the announcement says 30% faster and up to 30% cheaper. the third one is the one i still cannot answer from the docs.
Sonnet 5.5 is smarter, more efficient, and 30% faster than Sonnet 5. It costs up to 30% less for most work, so your Claude Code usage goes further too.
Use it for well-scoped everyday tasks like fixing bugs and quickly iterating on features.
a personal operating system is really just 8 decisions.
mine is 33 tables and 291 tests across 62 files. these are the eight that mattered.
1. sqlite, one file, no hosted database. it survives everything except two processes writing at once. set busy_timeout or you get SQLITE_BUSY the first time a cron overlaps a request.
2. drizzle for the schema. 33 tables means 33 things that have to survive a migration you run on yourself on a tuesday night.
3. tests before features. 291 of them in 62 files. the point is not coverage. it is that an agent can run them and find out it was wrong before i do.
4. the agent runtime spawns claude -p behind a native mcp server. that one decision is where the next four live.
5. spawn from the project directory, not /tmp. spawning from /tmp breaks mcp server registration. the child comes up with no tools and says nothing about it.
6. do not sanitize the child env. cleaning it kills subscription auth. you get an auth error that looks like a billing problem and is not.
7. do not pass --permission-mode. it breaks tool access. the child runs, answers, and never touches a file.
8. do not run the runtime inside a claude code session. the child inherits CLAUDECODE and CLAUDE_CODE_ENTRYPOINT and comes up degraded. same prompt, worse agent, no error.
four of those eight are not code. they are environment variables and a working directory.
claude code 2.1.281 added mcp server checks to claude plugin validate.
it now tells you which .mcp.json entries would be silently dropped at load.
silently dropped is the whole point. i have spent more hours on mcp servers that came up with no tools and no error than on any bug that actually threw.
spawning from /tmp broke my server registration. no error.
clearing the child env killed subscription auth. an error, but the wrong one.
--permission-mode killed tool access. the agent still answered.
every one of those cost me an afternoon because the failure looked like the model being bad at the task.
a validator that names a dropped entry turns one of those afternoons into one line of output.
if you run agents alone, the expensive bug is never the one that throws. it is the config that loads clean and does half of what you wrote.
claude code 2.1.283 has a long fix list, and three of the entries describe things i would have spent a day blaming on my own code.
my runtime spawns claude -p behind a native mcp server, so all three sit in my path.
1. stdio mcp servers were being left running when a session ended while they were still starting.
orphaned processes. i would have gone straight to my own spawn and teardown code, because that is the first place i look when something is still alive that shouldn't be.
2. a stateless remote mcp server that returned a brief 404, a proxy mid-redeploy for example, stayed unusable for the rest of the session and kept showing as connected.
connected and not working is the worst state a thing can be in. there is nothing to debug because nothing is reporting a problem.
3. progress notifications from a long tool call were discarded once the call moved to the background.
so the call goes quiet and you cannot tell slow from stuck.
separately, /context was not counting mcp server instructions at all. they get their own row now and count toward the total. if you have been wondering where your context window goes before you have typed anything, that row is new information.
claude code started showing spend in dollars instead of tokens.
the changelog example is $271.40 / $500.00 spent this month. it sits in the status line where the token count used to be.
i have read token counts for a year like weather. something that happens to you, in a unit you cannot feel.
a dollar figure lands in half a second. a token count never landed at all.
that matters more than it sounds, because my agent runtime spawns claude -p behind an mcp server. i am not sitting there watching the children run. they start, they finish, and the only honest thing i could say about what they cost was that it felt fine.
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
five things broke when i wired claude code into a runtime of my own. four of them broke without an error.
the thing runs claude -p as a child process behind an mcp server. my os around it is 33 tables and 291 tests in 62 files, and none of those tests caught any of this, because the failures never raised anything to catch.
1. spawning from /tmp
the mcp server registered nothing. no error, no warning, not even an empty tool list i could have asserted on. the child came up with fewer tools than i gave it and answered anyway. i noticed because the answers got worse, not because anything told me.
2. stripping the child's env
this felt like the responsible move. minimal env, don't pass what it doesn't need. it killed my subscription auth and the child started asking for an api key i don't use. the fix was passing more through, not less.
3. --permission-mode
i set it to control what the child could touch. it broke tool access completely. i went through my own tool definitions twice before i thought to take the flag back out.
4. CLAUDECODE and CLAUDE_CODE_ENTRYPOINT
i was building inside a claude code session, so the child inherited both, concluded it was nested, and quietly degraded itself. it worked from a plain terminal and failed in the exact setup i work in every day. after that i stopped trusting anything i tested that way.
5. two processes on one sqlite file
no busy_timeout, so SQLITE_BUSY, in my face, with a name i could search. the loudest failure of the five and the only one that cost me nothing.
claude code 2.1.282 shipped this week. one line in it: fixed more cases of continued or resumed sessions re-sending earlier messages in a changed form.
so if you used --continue or --resume, the history the model got back was not quite the history you had. not truncated. changed. nothing in the terminal said so.
i don't build software the way a dev does. i keep a session open for hours, close the laptop, resume in the morning, and the only test i have on the conversation itself is whether the answers still feel as sharp as they did yesterday.
i have no idea whether this ever hit me. that's the part i keep chewing on.
every time the quality dropped across a resume i had three explanations ready and all three were about me. bad prompt. too much at once. late.
none of them were checkable.
if you build with claude code and you've been carrying that particular suspicion about yourself, some of it was upstream.
the fix is one bullet in a long release.
i assumed my own system needed to look like a product. it cost me the first half of the build.
nacho io is my operating system. next.js 16, sqlite, drizzle, 33 tables, 291 tests in 62 files. here's what i assumed and what it actually turned out to be.
1. i assumed i needed postgres.
it's sqlite. one file. the whole database fits in a backup i can read. the first schema i wrote was for a database i never used.
2. i assumed i'd build a ui for every part of it.
most of it is slack, in socket mode. i talk to my own system in the app i already have open. the screens i designed first are the screens i never opened again.
3. i assumed the agent layer would be the hard part.
it's a runtime that spawns claude -p behind a native mcp server. that part works. the hard part is everything around it deciding what's worth spawning for.
4. i assumed integrations were a later problem.
outlook, notion and apify are what make it mine instead of a demo. without them it's a very well tested empty room.
5. i assumed tests were something devs do to feel safe.
291 tests in 62 files is the only reason i can let an agent change my own system while i'm not watching. it's not confidence. it's permission.
6. i assumed the design work would be the easy part, because i'm a designer.
i still don't know what this thing should look like. i know exactly what it should do.
the 33 tables are the part i'm proudest of and the part nobody will ever see.
I wrote CLAUDE.md like a README for months. That's why it never helped.
Now it has five sections and none of them describe the project.
1. Versions, exact. Not "Astro" but Astro 5.2, content collections, no SSR. A guess at the minor version costs an afternoon.
2. What I already decided. URL shape, slug rules, where images live. Written as rules, not preferences, so nothing gets relitigated on file 40.
3. What will go wrong in production. "Editors paste 400 words into a field sized for 40. Every text field needs an overflow case." This one section has saved me more hours than the other four together.
4. What not to touch. The handful of files that are hand-tuned and stay hand-tuned.
5. How I want to be told. "Show me the plan before you write more than one file."
Section 3 took me longest to learn. It's the only one about the people who will actually use the thing.