Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.
finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
Forgot to add the full prompt in the replies haha
Here it is:
# Scope Guard
Complete the current task with the minimum sufficient change.
## Before editing
- Read the relevant code, tests, and configuration directly. Do not work from search snippets or guesses.
- If the requirement is ambiguous or the premise is unverified, resolve that before building on it.
- State a minimal plan:
- **Outcome** — the exact behavior requested
- **Non-goals** — what this task will not do
- **Files** — the smallest set expected to change
- **Proof** — the check that will prove the change works
- Start with one implementation path. Split work only when the task has genuinely independent parts.
## While editing
- Reuse existing code, helpers, patterns, and test setup before adding anything new.
- Fix bugs at the root cause. Do not stack patches around a wrong premise.
- Add an abstraction, adapter, or config layer only for a second real caller
in this task or a stated requirement.
- Preserve behavior outside the requested change.
- Do not design for rare or future cases nobody asked about.
- Remove code you replace. Keep an old path only when compatibility is an explicit requirement.
## Pause and confirm
Read-only discovery is always allowed. If the task has not already authorized it, get approval before:
- Materially expanding the scope or touching unrelated files
- Adding a dependency, framework, service, or new test infrastructure
- Changing a public API, schema, storage format, or wire format
- Deleting or overwriting user data, discarding uncommitted work, rewriting history, or dropping data
- Keeping two implementations of the same behavior alive
## Testing
- Run the narrowest existing tests that exercise the changed behavior.
- Extend the most relevant existing test before creating a new test file.
- Add a test only when changed user-observable behavior is not covered, or when the user asks for one.
- Each new test must protect a clear acceptance criterion or regression risk.
- Do not backfill unrelated coverage or introduce test infrastructure for this task alone.
- Do not use passing tests as justification for extra abstractions or scope.
## If the plan grows
Stop when the work starts adding future-use layers, workaround stacks,
unrelated cleanup, or tests for unstated behavior. Rewrite a smaller plan
and confirm the new scope.
## Done means
- The requested behavior works and the acceptance criteria are met
- Relevant checks pass, with the exact commands and results reported
- Every touched file is necessary and the diff contains nothing unrelated
- No debug code, backup copies, dead paths, or scratch files remain
- Assumptions, limitations, and unverified runtime behavior are stated plainly
I’m reading this on the plane right now, and it genuinely gave me chills. You should all give it a read.
And a huge thank you to @zartbotF for writing such an incredible piece.
https://t.co/IMMd9WNJI4
if you want to understand what building at real scale is like, this is a must read
1.1.1.1, @cloudflare's DNS resolver, stores over stores over 250 billion (!!!) cache entries at any given time, so every minor change has a massive ripple effect
we made 5 changes, and saved > 100TB of memory...
https://t.co/sIBXkJ1CDO
At 250 billion DNS cache entries, one wasted byte costs 250 GB of RAM.
Five Rust optimizations later: 100 TB freed, inserts 43% faster, lookups 19% faster.
We didn't trade speed for space. https://t.co/mMyOYnj0mQ
i built an island that visualizes my bots as they work! it's cute watching them work and then go to bed and their little houses when they're resting 🥺
prompt to make this is below!
"The standard pace is for chumps" has been echoing in my head all week. It's unbelievable how fast you can move when you just decide to push the pedal all the way. https://t.co/KmaA0mUupo
Going to double down on the MacBook mission with Omarchy. We almost have perfect coverage for the vintage Intel era going from 2009-2020. There's a straight shot to get the M1 and M2 machines going too, even if it's a lot more work. But we'll do the work. We'll fix everything.
Git at Scale (by cursor) has been one of the most interesting blog posts i've read in a while. It came right when I was frustrated with Shopify's internal git system.
As an exercise, I've implemented it over the weekend as open source. It's a single rust binary that you can point at any S3 type object store. It uses WAL and CAS primitives and requires no other data store.
It also implements bundle-uri so large git repos (like our mono) are very fast to download as a chain of static bundles. Also comes with basic familiar UX.
https://t.co/oEnWK3eIL4
sharing a new long-form blog post: ai chip architectures
it covers the leading chip architectures (nvidia, amd, tpus, trainium, cerebras, groq) across architecture, scaling (scale-up and scale-out), and software stacks.
it helps build an intuition for the architectures and their trade-offs.
https://t.co/7eZMh3ddZS
Highly recommend https://t.co/4nV7gECAOJ as an option that’s more approachable than https://t.co/40q06gOdO3 if you want to build a language model from scratch
It’s difficult but nothing beats understanding things at the atomic level.
1/
I'm not a coder. I'm a trader!
But... I built an automated system that ranks cash-secured put (CSP) candidates; pulling live options data, fundamentals, technicals & growth trends. Then, scores them into a clean dashboard.
Here's exactly how I did it (with Claude's help) 🧵
if your definition of coding is just the act of typing code, then yes, it's solved. agents can type code faster than anyone can.
but i think most of us think of coding as the act of solving problems with code. in that definition, coding is definitely not solved. agents can generate code quickly and build whatever you want, but without the right constraints, they end up causing more problems than they solve.
i think of this like my gardening analogy. agents are like magic beans that grow in any condition, but in wild and unpredictable ways. if you put the right scaffolding around them however, they can grow into whatever shape you want.
you still want humans thoughtfully and intentionally steering this outcome. we have a name for them: engineers. agents have greatly lowered the barrier to entry to doing engineering. what was once the domain of those who understood syntax and logic is now accessible to anyone who can explain their ideas. english is a new programming language for engineering outcomes with agents.
now, even ordinary people can build impressive software. and the good ones will learn engineering through first principles. as someone who is a self-taught programmer (i first learned how to code through Excel macros!), this is incredibly exciting to me, and i can't wait to see what everyone learns and builds in this new world.
let a thousand gardens bloom!
every team needs a gardener. someone quietly watching the stream of PRs flowing into your codebase, noticing the smells: the third isRecord this week, the lint suppressions creeping like ivy across your carefully planned garden. a steady hand tending the weeds that would engulf it in slop if left unchecked.
with careful grooming it stays a garden. every weed you pull becomes a rule, so it can't grow back. without that, it's just whatever grew.
We all have ideas. Ideas are immortal. They last forever.
What doesn’t last forever is inspiration. Inspiration is like fresh fruit or milk: It has an expiration date.
If you want to do something, you’ve got to do it now. You can’t put it on a shelf and wait two months to get around to it. You can’t just say you’ll do it later. Later, you won’t be pumped up about it anymore.
If you’re inspired on a Friday, swear off the weekend and dive into the project. When you’re high on inspiration, you can get two weeks of work done in twenty-four hours. I
Inspiration is a time machine in that way. Inspiration is a magical thing, a productivity multiplier, a motivator. But it won’t wait for you. Inspiration is a now thing. If it grabs you, grab it right back and put it to work.
Inspiration is perishable.