AI Engineer | Agentic Engineering · AI Agents
RAG · MCP · AI System Design · Healthcare AI - LLM Inference - AI RCM
360-days and counting GitHub streak
After evaluating GPT-5.6 in new ChatGPT desktop app, on our ChatGPT Business plan, the progress is notable.
Beyond strong reasoning, the plugin ecosystem enables high-quality automations and deep integrations with external applications. Features like Computer Use, combined with multi-agent orchestration, allow for sophisticated end-to-end workflows that significantly reduce manual work in professional environments.
A clear step forward in practical AI tooling.
@OpenAI @ChatGPTapp @thsottiaux@gdb
“Model orchestration is in many ways the natural outgrowth of agentic engineering”
Huge thanks to @AndrewYNg and @DeepLearningAI for the deep dive into Sakana AI’s Fugu and Fugu-Ultra. The article highlights how dynamic orchestration allows us to achieve near SOTA performance on benchmarks like GPQA-Diamond, LiveCodeBench Pro, and SWE-Bench Pro without being dependent on a single provider.
Full article: https://t.co/wk7Aukwzv1
2008
3rd rocket explodes.
Company on life support.
Everyone screamed “quit pack up.”
Yesterday, SpaceX IPO’d at $2.13 TRILLION
The largest in history.
"When something is important enough, you do it even if the odds are not in your favor." Elon Musk
Delusion is just vision the world hasn’t caught up to yet.
Stay obsessed. Stay delusional.
Grit always wins.
@karpathy dropping truth again Claude Fable 5 feels like the first model that actually crosses the "this could replace a solid engineer". The Jevons Paradox part is spot on. we're not going to code less, we're about to code way more weird, ambitious, personal software than ever
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time.
I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!
Claude 5 Fable is actually insane. Not just another benchmark flex.
it's built for the long, complex, agentic stuff where models usually fall apart.
Token efficiency + sustained focus over millions of tokens + real engineering wins (Stripe turning months into days)?
Cold take on what comes next:
- OpenAI will flourish
- Anthropic will continue to be profitable
- Google will not catch up to Anthropic or OpenAI
- no chinese company will catch up to Anthropic or OpenAI
- the highest tier of intelligence will become a luxury product that only companies and multi-millionaires/billionaires can afford
- most of the companies that invested massively in them will have massive returns
- SpaceX’s AI will be fine and on par with Google by end of year
- Nvidia will become the first 10T company
AI benchmark wars are starting to feel like Formula 1 for tech companies, everyone flexing numbers, everyone claiming dominance.
But beneath the marketing, the progress is actually insane. Models are becoming genuinely useful for coding, research, automation, and real-world work
Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors.
Available today at the same price.
Leveling up how we fine-tune AI models 🚀
Shared some practical insights on LLM fine-tuning, trade-offs, and what actually works in real-world setups. If you're building with AI.
👇 Full breakdown:
https://t.co/ePmZapGMJv
Curious to hear how others are approaching this 👀
Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
We made it 🏆
3rd prize Track Two: Qwen-Image-2512-Lora-Advertisement.
Massive shoutout to the Qwen team for shipping a base model this powerful, and @Ali_TongyiLab@Alibaba_Qwen for an unforgettable competition. One month of training, plenty of doubt, all worth it.
After one month of intense training and another month of rigorous judging, the Qwen-Image LoRA Training Competition has officially concluded! 🏆
We’ve been absolutely floored by the level of technical mastery displayed across every single entry.
From the delicate preservation of cultural heritage to the artistic magic of color grading and style transfer, you’ve pushed Qwen-Image to its absolute limits.
Massive congratulations to our winners. Check your DM and get in touch with us within one month.
First Prize (iPhone 17 Pro Max):
@ajie_run_CL@triaakkatuki
Second Prize (PS5 Pro):
@svntax@liang_zoey2364@ovi054@Playmaker_Tech
Third Prize ($800 Value Shopping Card):
@webdevdeep@ovi054@dx8152@galvinlim2001@Rayyan9477
A huge thank you to everyone who participated. Keep building, keep training!
My name on this list is something I'll remember forever 🏆
3rd prize, Track Two: AI For Production Qwen-Image-2512-Lora-Advertisement.
The talent across every team was insane. Thank you @ModelScope2022 for an incredible competition 🙏
Builders, let's keep going 🚀
After one month of intense training and another month of rigorous judging, the Qwen-Image LoRA Training Competition has officially concluded! 🏆
We’ve been absolutely floored by the level of technical mastery displayed across every single entry. From the delicate preservation of cultural heritage to the artistic magic of color grading and style transfer, you’ve pushed Qwen-Image to its absolute limits.
Massive congratulations to our winners. Check your DM and get in touch with us within one month.
First Prize (iPhone 17 Pro Max): @ajie_run_CL@triaakkatuki
Second Prize (PS5 Pro): @svntax@liang_zoey2364@ovi054@Playmaker_Tech
Third Prize ($800 Value Shopping Card): @webdevdeep@ovi054@dx8152@galvinlim2001@Rayyan9477
A huge thank you to everyone who participated. Keep building, keep training! We'll be sharing their wonderful work later this month — stay tuned! 🎨
Surreal moment 🤯
The goal was simple: democratize commercial-grade ad photography so anyone with a single GPU can create studio-quality assets.
Honored to see it featured by @Ali_TongyiLab as a 3rd prize winner. Try it, break it, build on it 🙏🚀
📸 Elevate your brand with Pro-Level Ad Photography.
Presenting our 3rd prize winner: Qwen-Image-2512-Lora-Advertisement by @Rayyan9477.
Trained on 12k+ curated images across 40 advertising categories—all on a single consumer GPU, this 1.7GB LoRA adapter transforms standard photos into commercial-grade advertisement assets with ease.
📍 Try it now:
https://t.co/3u4aQmGj3q