The more I learn about AI agents, the less I think the agent itself is the main story.
**Note #3** is about the system around it: workflow, tools, sandbox, permissions, ownership—and why reliable agents need more than a smart model. https://t.co/bwudaASt93
AI Agent Learning Notes #2 is out.
This one goes deeper into the part that makes agents both powerful and unpredictable: the loop.
An agent can decide → act → observe → adjust → repeat.
That’s what lets it solve multi-step problems.
But every extra loop also creates another chance to misunderstand context, make a bad assumption, or slowly drift from the original goal.
In Note #2, I’m sharing what I learned about:
• how the agent loop actually works
• why agents drift
• how to keep them grounded
• better ways to prompt an agent for longer tasks
Still learning this from a beginner’s perspective — and documenting what clicks along the way.
Note #2 ↓
The same loop that makes AI agents powerful can also make them drift.
**Note #2** is about what happens inside the agent loop — how an agent decides, acts, observes, and keeps going — and why repeated loops can slowly pull it away from the original goal. https://t.co/fIfi3W3J0R
The same loop that makes AI agents powerful can also make them drift.
**Note #2** is about what happens inside the agent loop — how an agent decides, acts, observes, and keeps going — and why repeated loops can slowly pull it away from the original goal. https://t.co/fIfi3W3J0R
That distinction really clicked for me while learning this. Tool use gets most of the attention, but a model calling a tool once doesn’t necessarily make it an agent.
The interesting part is: observe → evaluate → decide → act → repeat.
I’m curious how you think about the “evaluation” step in production agents — rules, another model, external feedback, or some combination? I felt like evaluation often leads to drifting. It should include necessary steps in the loop to prevent drifting!
I thought an AI agent was basically ChatGPT with tools. This week, I’m learning what actually makes something “agentic.” My beginner takeaway so far: **Goal + Choice + Tools + Loop.** Here are my learning notes ↓ https://t.co/1cM1ixr1Vz
Exactly — choice + loop might be the real “agent test.”
Goal + tools = “I can help you do something.”
Choice + loop = “I’ll figure out what to do next based on what just happened.”
That shift from responding → acting and adapting is what clicked for me.
Curious: where do you think agent development is heading next?
I’m starting a new series: AI Agent Learning Notes.
The goal is simple: learn AI agents from a beginner’s perspective, document what actually clicks, and share the notes with anyone trying to understand the space without getting lost in the hype or jargon.
Note #1: What is an AI agent, common misconceptions, and where agents are actually useful.
More notes coming as I learn, build, and test. ↓
I thought an AI agent was basically ChatGPT with tools. This week, I’m learning what actually makes something “agentic.” My beginner takeaway so far: **Goal + Choice + Tools + Loop.** Here are my learning notes ↓ https://t.co/1cM1ixr1Vz
@alvinfoo Changed the graph here to visualize the token difference per $50, which almost shows that you don’t quite get anything for $50 when using Claude. The comparison is insane!
@HealthRanger Changed the graph here to visualize the token difference per $50, which almost shows you don’t quite get anything for $50 when using Claude. The comparison is insane!
Three major Chinese AI releases over the weekend—three very different strategies:
⚡ @deepseek_ai V4 Flash: Optimize intelligence for speed, cost and local deployment.
🧠 @Alibaba_Qwen 3.8-Max: Push frontier reasoning, multimodal understanding and long-horizon agentic work.
🎬 @MiniMax_AI H3: Bring multimodal intelligence directly into professional video creation.
What a weekend for Chinese ai frontier models: Here’s the @MiniMax_AI new model launched roughly at the same time with @Alibaba_Qwen
MiniMax-H3: A multimodal model that understands text, images, video, and audio, generates 2K videos with native stereo sound, and is expected to have open weights is a big step toward making professional AI video creation accessible to everyone.
#AI #GenerativeAI #VideoAI #OpenSourceAI #MultimodalAI #MiniMax #ContentCreation #AICreators #MachineLearning #BuildInPublic
The release of @deepseek_ai V4Flash could be a major catalyst for the next wave of AI hardware.
A powerful model that can run locally with exceptional cost efficiency makes it much more practical to build AI-powered devices—from edge computers and smart home products to wearables with private, low-latency intelligence.
The AI race isn’t just about bigger models anymore. It’s about bringing capable models onto the device.
#AI #EdgeAI #OnDeviceAI #LocalAI #DeepSeek #GenAI #AIAgents #Wearables #AIHardware #OpenSourceAI #V4flash
DeepSeek V4 Flash 0731 is impressive.
It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient.
The model is available in Bionic, both for running locally and through the LM Studio Cloud!
https://t.co/6oxyTH0L2d
The AI race is no longer about benchmarks. It’s about how long a model can think, plan, and execute without human intervention.
Qwen3.8-Max looks like another step toward that future.
What caught my attention wasn’t the parameter count (2.4T). It was the shift in capability:
• 10+ days of autonomous coding
• Closed-loop planning and self-correction
• Native multimodal feedback during execution
• Production-ready work instead of isolated demos
This changes how we should evaluate AI.
Instead of asking:
“How good is this model at coding?”
We should ask:
“How much real work can it complete before needing a human?”
That’s the metric that will define the next generation of AI.
Open weights arriving next week could make this even more interesting.
We’re moving from AI as a tool → AI as a long-horizon collaborator.
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: https://t.co/iwjmQxLBof
✅ Qwen Studio: https://t.co/4V2pFvDovG
⚡ API: https://t.co/gAGqaLQGbN
CODEX SKILL THAT TURNS CUSTOMER FEEDBACK INTO A ROADMAP!
Most feedback analysis stops at positive or negative.
I made a Codex skill that turns support tickets, interviews, surveys, reviews, sales calls, and churn notes into evidence-backed product priorities and credible customer proof.
Give it your feedback files and Codex maps recurring pains, ranks what to build next, and links every recommendation back to exact evidence IDs.
It keeps four things separate:
-what customers said
-the underlying problem
-what to validate or build next
-what you can honestly claim
-> feedback clustering with source traceability
-> product opportunity scoring
-> Now / Next / Later roadmap
-> customer proof + case-study candidates
-> objections, churn risks, and unsupported claims
-> searchable Startup Customer Evidence Map
-> privacy-first quote and consent handling
-> one-command install
What would you analyze first: support tickets, interviews, or churn notes?
Install: npx --yes startup-feedback-engine@latest
100% open source.
Repo in Bio.