"Claude Sonnet 5.5 has roughly matched Claude Fable 5.1 (165)."
Anthropics mid-size model (!) now matches it state of the art and probably the best model in the world within just a few months of release.
I am so excited for Fable 5.5, you have no idea.
The LLM scheduling loop, clearly explained:
(bookmark this)
Whenever a user sends a prompt to an LLM, the resulting inference request does not go straight from the API endpoint to the GPU.
Instead, it first enters the serving engine, where a scheduler decides when and how its tokens will be processed.
Here is the complete lifecycle:
1) The request enters the waiting queue
Each request arrives with a prompt, sampling settings, and a maximum output length.
Prompt lengths vary, and the scheduler cannot know the final generation length in advance. It therefore manages requests one iteration at a time.
2) The scheduler builds the next batch
Before every model step, the scheduler checks two main constraints:
- How many tokens can be processed in this iteration
- How much KV-cache space remains
It then applies a scheduling policy, often prioritizing running decode requests before admitting new work.
3) New requests enter prefill
During prefill, the model processes the prompt tokens and creates the K and V tensors needed by attention.
This stage can process many prompt tokens in parallel. Long prompts may be split into chunks so they do not block active generations for too long.
4) Running requests enter decode
Once prefill finishes, the request moves to decode.
The model now generates the next token using the KV cache created for all previous tokens. Standard autoregressive decoding usually adds one new token per active request during a model step.
5) The GPU executes the selected work
Modern serving engines can place prefill and decode work in the same GPU batch.
For example, request A may be processing its prompt while requests B and C generate their next tokens.
The batch is therefore not a fixed group that runs until every request finishes. Its composition can change after every model step.
6) The engine updates request state
After execution, generated tokens are sampled and appended to their requests.
A finished request releases its KV-cache blocks and leaves the batch. An unfinished request keeps its state and becomes eligible for the next scheduling iteration.
The scheduler can then use the freed token budget and cache space to admit another waiting request.
In the visual below, C finishes after iteration 1, so D enters during iteration 2. By iteration 3, the active batch contains only A and D.
This iteration-level replacement is continuous batching. It keeps the GPU working while requests with different prompt and output lengths progress independently.
If you want to dive deeper, I wrote a detailed article covering the full LLM inference pipeline, including tokenization, prefill, decode, KV caching, and the latency metrics each stage controls.
Read it below.
Send this prompt to Claude.
It finds 40 real people who need what you sell this month, shows you where it found each one, and writes every one of them a pitch.
One rule inside it does most of the work: no prospect without a public link. You get real people, not made-up leads.
Full prompt below:
Claude makes AI agents f...cking illegal
4 GitHub repos that turn Claude from a chatbot into a coding system
01 Claude Code
▸ https://t.co/SkjKYCjDnK
→ terminal → codebase → git → execution
02 Claude Agent SDK
▸ https://t.co/xguka3Giwv
→ build your own agents on the Claude Code harness
03 Agent Skills
▸ https://t.co/92AxiUAo9k
→ reusable instructions + scripts + specialized workflows
04 Claude Code Action
▸ https://t.co/YtEj3OJm3L
→ put Claude directly inside GitHub PRs + issues
the stack:
Claude → understand repo → load skills → use tools → write code → test → open PR → review → iterate
that's the part getting powerful
Claude isn't just generating code anymore
Anthropic is building the infrastructure around it to keep working after the first prompt ⭣
Why Widening Credit Spreads Could Be The Next Major Warning
Corporate credit risk is being repriced higher.
5 year CDX protection is around 61 basis points for investment grade and roughly 350 for high yield. Investors are demanding more compensation to insure against corporate credit losses.
This does not mean a credit crisis has arrived. Current levels remain below earlier stress peaks. What matters is the speed and breadth of the move.
Why This Matters
Credit spreads are not Treasury yields.
A company can benefit from a falling government bond yield and still face a higher borrowing cost if its credit spread widens faster.
The Fed raised rates 25 basis points in September to 3.75% to 4.00%. The bigger issue is refinancing expectations. A borrower expecting cheaper money may suddenly have to refinance at a much higher cost.
At the same time, energy, freight and other input costs are pressuring margins. If cash flow weakens while refinancing becomes more expensive, interest coverage can deteriorate before a company misses a payment.
The Bank Data Are Saying Something Similar
The Dallas Fed September survey showed
• Loan demand fell from 46.7 to 8.2
• Loan pricing jumped from minus 3.2 to 25.0
• Credit standards moved from minus 1.7 to minus 10.4
• The 6 month business outlook fell from 24.2 to minus 14.8
Bankers also expected further tightening in parts of business and commercial real estate lending. Current loan performance was still improving. Credit can tighten before defaults rise.
A borrower can still be paying yesterday’s loan while no longer qualifying for tomorrow’s refinancing.
History Shows The Pattern
In 2007, credit spreads widened before the full economic damage appeared. Fed cuts lowered policy rates, but deteriorating borrower quality and tighter lending standards overwhelmed much of that relief.
In 2015 and 2016, spreads widened sharply around energy and global growth fears without producing a national recession. Widening spreads are a warning, not a guaranteed outcome.
In 2020, Treasury yields collapsed while corporate financing conditions deteriorated dramatically.
Lower Treasury yields do not automatically mean easier financial conditions.
One Important Caveat
CDX indexes roll into new series in September. Changes in maturity and constituents can create discontinuities in a generic chart, so the exact size of this move should not be attributed entirely to worsening fundamentals.
But the broader direction matters because lending conditions are becoming more defensive at the same time.
What concerns me most is the combination
• Credit investors demanding more compensation
• Banks becoming more cautious
• Refinancing costs remaining high
• Operating costs staying elevated
If that persists, the next stage can be weaker investment, reduced hiring, refinancing failures and eventually higher defaults.
The real danger here is if lenders move from asking borrowers to pay more to deciding they no longer want to refinance them at all.
IT’S HAPPENING. 😂
I’ve been stalking this Mac mini like it owes me money.
Hong Kong, China ✅
Anchorage, Alaska ✅
Next stop: Iowa. 🌽
Apple says Tuesday, Oct 6.
My desk is about to get dangerous. 🔥