SpaceXAI engineer, Lauren Tan:
"GrokBot is the most powerful agentic tool we have ever shipped, but only 1% of users use it correctly|
at SpaceXAI, I'm running a team of 15+ GrokBot agents. I have a Chief of Staff agent, 3 managers and 11 workers - that's the new engineering setup in 2026"
In a 1-hour talk, a SpaceXAI engineer reveals how to get 100% of every agentic tool you are using
worth more than a $500 agentic course on the internet
skip Netflix and watch today, it will change the way you use GrokBot forever, then read the article below
I think Silicon Valley is making a fundamental mistake with personal agents… optimizing entirely for the outcome and forgetting that, for a lot of things, the process is a very core part of experience.
Travel is the obvious example. People will almost always never want “book me a trip.” They want to browse, compare, daydream, change their mind, send options to friends, and eventually book.
And the average person probably has far fewer recurring tasks worth delegating than the AI industry seems to think.
That’s why personal agents feel more like a feature that gets absorbed into existing products than a standalone category.
I just published a manifesto for all the developers out there who use an AI coding tool, but feel like they're not shipping much faster
I wrote 10 principles for changing how you build software with AI, based on what I've seen work for teams across Amazon
https://t.co/smq7Rt5QkZ
ripwire, from Red Hat Emerging Technologies, is a remarkably substantial new approach to giving coding agents repository context without embeddings, a vector database, an LLM indexer or a daemon. The zero-dependency C++23 binary parses 21 languages with Tree-sitter and builds a deterministic structural map of a codebase, ranking symbols for the task while attaching call relationships, complexity, git churn, change amplification and test coverage.
An agent can ask what matters for "incremental cache invalidation," for example, and receive the relevant symbols, their callers, likely blast radius and tests to run in a token-budgeted response instead of grepping and opening files repeatedly.
GitHub Repo: https://t.co/EHfXTtUssS
For the life of me, I am unable to do the multi-agent coding thing (when multiple agents are running on the same project) 🥶
I've tried it multiple times. But every time I get frustrated because:
1. I start to lose track of what each agent is doing.
2. Agents overwrite each others code and then drama ensues between them.
3. End up forgetting to check over aspects of the bug fix, feature or component because there have been multiple changes at once.
How do you guys do it? 🧐
I'm just not made for the madness that's called multi-tasking.
Will be sticking with what I know and work on a single thing at a time. My brain just copes better with that 😅
Respect to anyone who can to the 10-coding-agents-at-the-same-time thing 🫡
Most of your "Opus tasks" aren't actually Opus tasks. File reads, status checks, basic code changes don't need the $15/$75 model.
I added one line to Claude's persistent memory and now it coaches me on cost automatically:
"Call out when I'm using the wrong model tier. Lookups on Opus = waste. Architecture on Sonnet = underpowered. Quick nudge, not a lecture."
~95% of my work runs fine on Sonnet. Opus is for the actually hard stuff. 50x price spread between tiers.
One line in https://t.co/mp7EuYwUqv, Project instructions, or any system prompt. Set it once, every session benefits.
I've been using Opus 4.6 for a bit -- it is our best model yet. It is more agentic, more intelligent, runs for longer, and is more careful and exhaustive.
For Claude Code users, you can also now more precisely tune how much the model thinks. Run /model and arrow left/right to tune effort (less = faster, more = longer thinking & better results).
Happy coding!
In Claude Code v2.1.30, we introduced /debug, a built-in skill for Claude to read your session's debug logs and troubleshoot your session. Great for chatting through issues like "/debug why didn't my hook trigger?" or "/debug why did my tool call fail?"
How did /debug come about? Last week, I was observing a group of users onboarding onto Claude Code. I saw that when something unexpected occurred, it wasn't easy for us to figure out the cause from the limited TUI display. Meanwhile, it was a chore to locate the logs. What if we simply gave Claude access to it from within the TUI? That way, it can use its tools to scan the logs and/or pull in the claude-code-guide subagent. The user can also /rewind when they're done debugging.
This feature is only as useful as what's in the debug logs. We're open to suggestions about what else we should add to them! Then, /debug should get better over time.
We've added a new command to Claude Code called /insights
When you run it, Claude Code will read your message history from the past month. It'll summarize your projects, how you use Claude Code, and give suggestions on how to improve your workflow.
Introducing GLM-OCR: SOTA performance, optimized for complex document understanding.
With only 0.9B parameters, GLM-OCR delivers state-of-the-art results across major document understanding benchmarks, including formula recognition, table recognition, and information extraction.
Weights: https://t.co/vqIBgBCXYi
Try it: https://t.co/Ld7H8Pasls
API: https://t.co/xVLNG0XSfP
I'm Boris and I created Claude Code. I wanted to quickly share a few tips for using Claude Code, sourced directly from the Claude Code team. The way the team uses Claude is different than how I use it. Remember: there is no one right way to use Claude Code -- everyones' setup is different. You should experiment to see what works for you!
Introducing LM Studio 0.4.0 🎉✨
It's the next generation of LM Studio.
🪄 Deploy on servers, in CI, or anywhere
🚄 Parallel requests for high throughput use cases
🔨 New stateful REST API: use local MCPs
🎨 Complete UI revamp
See what's new in this release👇🧵
Just added Best Practices for Claude Code to our docs! 🥳
Always looking to add more from the community though, what setups/patterns have been working well for you?
https://t.co/nnmRu8iz16
Low-key websites I quietly rely on
1) https://t.co/FDnurfhwge
Gives you a brutally clear learning path for roles like frontend, backend, DevOps, etc
No fluff, just “learn this → then this → then this”.
2) https://t.co/1xhB1Us0oz
An online playground to quickly test HTML, CSS, JS without setting up anything locally
Perfect for quick experiments and debugging ideas
3) https://t.co/d80zVxq6TY
A collection of reusable React hooks with real use cases
Saves time and helps you avoid rewriting the same logic again and again
4) https://t.co/UgqeLiqese
Concise cheat sheets for languages, frameworks, and tools. Ideal when you forget syntax and don’t want to read a 20-minute blog
5) https://t.co/OehnjnfVix
Turns messy JSON into a clean visual tree
Makes understanding large APIs and configs way easier than staring at raw text
6) https://t.co/mcMPeqEFWJ
Lets you generate and preview color palettes instantly
Useful when you want decent UI colors without guessing or copying blindly
7) https://t.co/9yEsuJrWCB
Build, test, and debug regex step by step with explanations Honestly, the fastest way to stop hating regex
8) https://t.co/7eM95WZ8cJ
Shows how big an npm package really is before you install it
Helps you avoid bloating your app with “tiny” libraries
9) https://t.co/Za87baZsBk
Tells you which CSS/JS features actually work across browsers Essential before using shiny new features in production
10) https://t.co/YeYh94AX4R
Google’s own diagnostics tools for DNS, email, headers, and network issues
Surprisingly useful for debugging real-world problems
👉 Which one of these do you already use and which one did you not know existed?
Document Index for Vectorless, Reasoning-based RAG!
PageIndex is an open-source RAG framework that removes vector databases and chunking from document retrieval.
Most RAG systems rely on semantic similarity. They chunk documents arbitrarily, embed them into vectors, and retrieve based on what looks similar.
But similarity ≠ relevance.
Professional documents like financial reports, legal filings, and technical manuals require multi-step reasoning and domain expertise. Vector search falls short when every section contains similar terminology.
PageIndex takes a different approach.
It builds a hierarchical tree structure from documents, similar to a table of contents but optimized for LLMs. Then it uses reasoning-based tree search to navigate and retrieve information the way human experts would.
Two-step process:
Generate a tree structure index of the document → Perform reasoning-based retrieval through tree search.
The LLM can "think" about document structure. Instead of matching embeddings, it reasons: "Debt trends are usually in the financial summary or Appendix G, let's look there."
Key features:
• No vector database infrastructure or embedding pipelines
• No artificial chunking that breaks context across boundaries
• Traceable retrieval with exact page-level references
• Reasoning-based navigation that mirrors human document analysis
PageIndex powers Mafin 2.5, achieving 98.7% accuracy on FinanceBench for financial document analysis.
The best part?
It's 100% open source.
Link to the GitHub repo in the comments!
We just open sourced the code-simplifier agent we use on the Claude Code team.
Try it: claude plugin install code-simplifier
Or from within a session:
/plugin marketplace update claude-plugins-official
/plugin install code-simplifier
Ask Claude to use the code simplifier agent at the end of a long coding session, or to clean up complex PRs. Let us know what you think!
I'm Boris and I created Claude Code. Lots of people have asked how I use Claude Code, so I wanted to show off my setup a bit.
My setup might be surprisingly vanilla! Claude Code works great out of the box, so I personally don't customize it much. There is no one correct way to use Claude Code: we intentionally build it in a way that you can use it, customize it, and hack it however you like. Each person on the Claude Code team uses it very differently.
So, here goes.