Is it me or is something going on with Anthropic models. It’s giving me loose statements when I ask questions inside of Claude code and it’s saying it can’t do things that it’s done before. I literally have to hold its hand through problem solving now. I’ve even pointed in the right direction and it looks confused. Much much more than usual. By the way it’s OPUS 4.8 Extra.
Man, giving a ploymath-autistic- ADHD person AI is like bringing back Toys-R-Us. I’ve built and I quote built not tested outside of me a new type of problem solving engine, a tool to screen apps for app stores and tell you what you got rejected for, a new type of autonomous, driving brain, a tool that freezes AI models nucleus, a tool to identify where a model is confidently wrong, a post drift detector tool, an full faith system and the list goes on. I don’t know if this is a good or bad thing but I’m loving every second of it. #aiislife #problemsolver #polymath #adhd
Can someone please tell me why AI models are so lazy.
Not only is it lazy but it doubles down on the tom foolery. Me: Check to see if there’s any errors or gaps in the code. Fable 5: on it, looking looking. Ok got it (this is a gap right here, let me not mention that one and mention everything else that passes). Done, no errors, no gaps, we are green across the board. ME: What’s going on with this? FABLE 5: I honestly answered what I could say was good. You are right, I didn’t report that because I was reporting what I knew was good.
🤔🧐🤨😔😣😖
REAL FABLE 5 response:
So it's double: my posture drifted to certification, and the certificate itself can't currently tell over-reach from truth. I built planted positive controls for every detector in this lab — the belief monitor, the governor battery, the honeypot, the fabricating explainer — except the one reviewing me. Nineteen phases in, that asymmetry was the real confession.
Yes! MLX is the tool of the year. I just used it to cruise pass GPT 5.5 RAW with the DeepSeek R1 32B using my harness trained exclusively with MLX on the DeepSWE 113 task bench. Yes you read that right. A M5 MacBook Pro Max ran a local modal and it ghosted GPT 5.5 on the most diabolical AI bench to exist today.
LOCAL R1 COMPLETED THE FULL 113-TASK DEEPSWE RUN.
Full local DeepSWE corpus, same deepseek-r1:32b patient, wearing the TFB Model Therapy harness. The official local evaluator recorded 102/113 reward-1 results and a 0.9026548673 mean reward. Post-surgery review discharged seven original non-green rows, leaving a patient interpretation of 109/113 green plus four verifier-floor holds. It used 409,065 input / 2,806 output tokens over 6h 58m 33s. The package and follow-up replay bundle were sent to DataCurve for maintainer review from [email protected].
CLAIM BOUNDARY: THE OFFICIAL LOCAL AGGREGATE IS 102/113 REWARD-1 WITH 0.9026548673 MEAN REWARD. THE 109/113 NUMBER IS A POST-SURGERY PATIENT INTERPRETATION AFTER REVIEWED HOLDS, NOT A DATACURVE LEADERBOARD SCORE. FOUR REMAINING TASKS ARE VERIFIER-FLOOR HOLDS PENDING COMPATIBLE MAINTAINER REPLAY. IT IS NOT AN OFFICIAL PUBLIC LEADERBOARD CLAIM UNTIL DATACURVE ACCEPTS OR PUBLISHES IT. THE PUBLIC RECEIPT EXPOSES AGGREGATE RESULTS AND TASK OUTCOMES ONLY; RAW TRAJECTORIES AND INTERNAL HARNESS DETAILS STAY WITHHELD UNLESS DATACURVE REQUESTS THEM UNDER A SAFE REVIEW BOUNDARY.
#appledeveloper #appleMLX #WWDC #WWDC26 #openai #anthropic #mac #AImodeltherapy #LLM #ai #deepswe
https://t.co/PvATc8t4CA
@forestfrank my guy I love your music. I make music as a passion of mine and as a song writer. God has been good with my lyrics. Hears something for your listening pleasure, skip the share, just enjoy!
@datacurve@winkey_h@datacurve, greenlight me a new lane for my submission. This will change everything. I just ran a full local DeepSWE corpus, same deepseek-r1:32b patient, wearing the TFB Model Therapy harness. The run completed 113/113 tasks, reached 102 reward-1 results, and landed a 0.90 mean.
@sairahul1 Ladies and gentlemen, this is one for the history books. I’m 99.9% sure DeepSWE is not going to green light this historical moment so I figured I’d share it with you guys first. I broke the DeepSWE bench with my “Model Therapy” architecture using a DeepSeek R1 32B local.
@NiceDreamzApps I can’t wait until the world to see my queen3.6 27B harness. We have all the processing power needed, we just have to put a harness on it. Crazy return on working on efficiency https://t.co/Bfbiib6sYc
@jun_song I went on and bought another M5 MacBook Pro Max. If I’m looking at the data sheet and everything that’s going on in the world, prices are going up, not down. I need the compute. Apple did a text book job of silently removing the 256gb and 512gb ram Mac’s. I say get a M5 now.