@scottstts Built it almost 5 years ago in a triple Pi cluster config and enclosure. It fits in a travel bag and is air-gapped. The family can't imagine a day without it, and our smart home runs through it. Good times.
@rezoundous Bro, Google, at a minimum, has gotten us out of these problems for quite some time, lol. It just takes a minute to find a reliable source these days, but it still works! ๐
This is a routing, switching, and policy problem if this is the case.
I'd love to see what "under the hood" looks like because this is crazy and proves no runtime constraint-aware verification and validation testing was done at all.
"Move fast and break things" strikes again.
I just paid $321 for a coding session where Fable 5 refused to do the work.
Here is where the work actually went:
Fable 5: $78
Opus 4.8: $242
75% of the session got routed to Opus because the new classifiers kept flagging routine coding work as cybersecurity risk.
The model I chose did a quarter of the job.
The fallback did the rest.
Anthropic said a small fraction of tasks would fall back.
My receipts say otherwise.
@nalinrajput23 The Google effect: Everything they build is just waiting to be abandoned at any given time because they don't commit to anything worthwhile that isn't tied to ads or Google cloud full send. ๐คฃ
@jun_song Google and its marketing team. That's it. Because there's no way. It's useless at anything that requires attention and functional tool use. It's a glorified prose generator that's serviceable as a summary generator at best when not forced inside their workspace apps. ๐
@morgantepell Google got left behind ages ago and insist on an AI plan just for gemini to be useful in their apps, marginally, while leaving it utterly useless in its actual home app, Gemini. They're a joke and have been since they fumbled the bag with Bard from the start.
@karpathy Been doing this for 5 years. Since GPT3 flawlessly with public receipts.
The issue was never intelligence it was containment, context routing, tool boundaries, memory discipline, handoff logic, security posture, etc and state persistence.
Got told: "You're doing too much." ๐คท๐พ
@elder_plinius There are ways to prevent drift from possible adversarial injections that cause jailbreaks, though. There are always ways to clamp the model's thinking and output so it never comes out, with receipts. I don't know why these labs haven't done it yet.
@jonathan_wilke Because its cheap. Thats literally it. Navigating anything from CLI is trash. Even devs have to jump to an IDE of sorts to review their code if they plan to traverse it meaningfully. Anyone saying otherwise shouldn't be taken seriously lol a well designed GUI + harness is amazing
@ClaudeDevs LiteLLM sidecar deployments should be on every dev's list of things to deploy in their AI-assisted architectures now, or you risk some lab pulling the plug on your brittle sandbox depending on one model.
@Rnb_Wd@RileyRalmuto@OpenAIDevs@OpenAI Stuff like this is why I stick to SIL prompting in app. Not letting these shit architectures touch my local hardware AT ALL. ๐คฃ