SpaceXAI just released a free workshop on how to run a team of Grok Bots
This 1-hour session on running a team of agents:
05:09 - why spawning 100 agents fails if you can't trust one
10:43 - no verification skill and you become the bottleneck
24:00 - a judge agent of a different model scores every sub agent
35:14 - grok bot launches: agents with their own identity
56:26 - the same agents now run product and design, not just code
Nobody adds a decision layer to a team of agents
Which bot goes next, is the evidence good enough, is this safe to ship
A frontier model writes a paragraph for every one of those
Jev only does that single job. 20-200x faster, $0.042 per million input tokens, and it never writes a sentence
LLM makes the work → Jev decides what happens next → code enforces it
Most people scale agents by opening more tabs
Watch this and save it before XAI ships another agent update
AI memory is getting f...cking illegal
10 open-source GitHub projects that stop agents from starting from zero EVERY new session
01 Mem0
▸ https://t.co/ifCHrrB19c
→ 66K+ stars
02 Hindsight
▸ https://t.co/HmCGciXrGj
→ retain → recall → reflect
03 memU
▸ https://t.co/vfrs9hesyQ
TURN MEMORY INTO KNOWLEDGE
04 Cognee
▸ https://t.co/Zp6UkoMdeK
→ documents + code + conversations → knowledge graph
05 Graphiti
▸ https://t.co/mNcepop8y6
→ remembers how facts change over time
06 OpenViking
▸ https://t.co/4DUaXPYLVe
MAKE THE AGENT STATEFUL
07 Letta
▸ https://t.co/KK06ZBNsZH
→ memory + identity across sessions
08 Letta Code
▸ https://t.co/fDy0CZtJPG
REMEMBER ACROSS THE STACK
09 OpenMemory
▸ https://t.co/VqyG5fOxNE
10 Agent Memory Benchmark
▸ https://t.co/QePFzW2FnC
the loop:
experience → remember → connect → retrieve → act → update
3 builds I'd actually test:
coding:
Hindsight → Cognee → Letta Code
personal agent:
Mem0 → Graphiti → Letta
company brain:
Cognee → Graphiti → Hindsight
and this is where the sh...t gets interesting:
bigger context isn't the same as memory
a model can hold 1M tokens and still meet you for the first time every morning
these projects change that
save this before your agent forgets it ⭣
I want to share my opinion on @GuidedHacking
When GuidedHacking acted as a sponsor for vx-underground I received small amounts of criticism for it, primarily because GuidedHacking was accused of "stealing content". I can assert with 100% confidence this is not true. It is not "stolen content".
I recently paid for a 2 year subscription to GuidedHacking for $150. I did it because, after I memed PirateSoftware and his followers got insanely rustled at me, I got curious enough to want to briefly explore anti-cheat systems and how they work fundamentally.
I have virtually no background in video game cheating, omit maybe modifying files or doing ghetto stuff with CheatEngine. I also did not want to bounce around forums like UnknownCheats to flounder around looking for information. So, I decided to pay for a more structured format.
I did not pay for it because I want to make game cheats. I'm not really a gamer, I don't find the concept interesting. At the end of the day I wanted to compare anti-malware technologies to anti-cheat technologies. From what I've learned is (if you're curious) is that they're both fairly similar except their implementation and approach is far different.
Anyway, GuidedHacking acts as forum, which it kind of is not anymore, because comments on most posts are disabled. Regardless, it is still indeed "guided" in its approach to discussing cheat and anti-cheat.
I think the criticism it receives is because many posts contain a "read more" area, or a citation to another website. I don't have a problem with this and I don't consider this "stolen content". This section allows readers to explore more which may be outside of the scope of the initial "guide". I think many websites use content or information from other websites but DO NOT cite the resources it got it from which is probably a much larger problem that can discussed at a later time.
Other criticisms people may have is that the content is very relaxed. It is written as if it's a conversation, not a technical paper. I personally enjoy this but others may not. It boils down to an opinion.
The pricing for 2 years isn't bad. I don't believe I'll use it for 2 years in totality. I will probably only use the website for a few months at most. I've been slowly reading the material on the website, it's difficult to lock in and really focus when you've got a small human being screaming in your face.
Anyway, if you're a nerd who is curious on how anti-cheats work in comparison to anti-malware or EDR systems, I recommend the website. It's interesting, entertaining, and the "read more" section or citations allow you explore beyond the website. I want to note the site does not really discuss how anti-malware systems work, but if you already know they work then it's a breeze to read. If you don't know how anti-malware systems work, the material on the site will act as a great starter point.
No, GuidedHacking did not request I write a review. But, I like the owner (Rake). He is a funny guy and I really believe his site is slept on, so I really wanted to give him the props and recognition I think he deserves more of
That's my Ted Talk
Last year, Anthropic was optimizing inference across 3 different types of compute (Nvidia, AWS trainium, Google tpus). There were mistakes made during these optimizations that resulted in actual degradation. Initially Anthropic denied it, but eventually they found, fixed and explained what happened. There have been no notable instances of degradation since.
Sadly, this one instance has made half of tech twitter’s brains fall out.
Separately, there’s the statistics side. You’re more likely to see the stupid spikes over longer windows.
If a model has a 1/50 chance of doing weird shit, and you do 20 prompts a day, there’s a ~30% chance you’ll have encountered weird shit on day 1.
By day 5, it’s closer to a 90% chance 🙃
Tl;dr, people are stupid, don’t understand non-determinism, and it happened once so they feel righteous
Gonna be really fun when this ships and confirms that the models aren't actually nerfed.
I'm sure that'll totally stop all of the stupid discourse, since people who talk about model nerfs all the time are incredibly rational, intelligent individuals.
We're open sourcing our company brain, a multi-player harness that acts as a teammate in a slack,
with frontier memory! Link below, you can set it up and self-host in 5 minutes with one click.
Ilya Sutskever said: learn these 30 papers and you know 90% of what matters in AI.
this repo rebuilt all 30 in pure NumPy, No PyTorch, No TensorFlow, Just notebooks you can run.
- RNN
- LSTM
- Transformer
- ResNet
- VAE
- AIXI.
Ilya’s reading list, now executable.
Every paper in Jupyter, Synthetic data included,
Zero DL frameworks.
- https://t.co/GEAzryNQxm
i knew i wanted to leave meta as early as last year. but i had no idea what was next. i wasn't having fun at work any longer and every hour spent there felt like agony.
during my 1 month recharge, i spent the whole time working on side projects. it was incredibly liberating building whatever i wanted, without worrying about aligning stakeholders and psc.
it was immediately obvious to me that if i couldn't stand an hour at work but could work 16 hour days on my side project and feel happy, that something had to change.
at @cursor_ai i voluntarily work 12 hours+ 7 days a week. why? because it's fun!
find something that makes you smile and lose track of time and you will do the best work of your life
This is an extremely good watch.
The things that felt novel/interesting to me:
1. Lock down your agents
Humans tend to like 'sharp knife' abstractions - that are powerful, but you can cut yourself if your use them wrong. Lauren says agents perform much better in extremely locked-down environments. Abstractions are designed so they can't screw up, and lint rules enforce it.
They built a whole internal framework (Dune) to keep the agent on track.
That helps optimise agents that don't have a large context window to work productively in your codebase.
2. Create verification infrastructure
To trust the results of any agent, you either need to sit and watch it OR have it provide evidence of its improvement. This has always made sense to me, but Lauren really pushes it hard here:
- Invest in custom CLI's that let the agent drive the app and measure its performance
- Make the app "factory ready" from the get-go - i.e. deployable to an environment where the agent can mess about with it
3. Feature Maps
Lauren's software factory (what she calls an 'outer loop') often requires the agent to break down vague bug reports from users and to turn those into potential fixes.
To aid that, they built a 'feature map' of all the main features in their application, which describe exactly how the app is supposed to function.
This has become essential for helping the agent navigate the codebase, and figure out quickly how things are supposed to work. It's maintained along with the codebase, and kept in sync via automations.
This is the kind of documentation I usually warn against. It goes stale quickly and can confuse agents if it's not kept up to date.
But Lauren's team are using it as critical navigation infrastructure, and it makes it possible for agents to explore faster and better - even on a large codebase. So it sounds like navigation docs like this are worth it if they enable new behavior.
Banger talk - watch the whole thing on 2x.
opus 5.5 is f*cking cracked at motion design
this entire video is code, 0 after effects
im open sourcing the prompt template for these motion designs
steal it to recreate these ↓
<inputs>
Ask me for: 8 to 12 UI states I want the shape to become (e.g. button, loader, player, slider, toggle, tabs, chart, command palette, toast), pure black and white or one accent color, and a royalty-free song around 120 BPM (e.g. Mixkit, free for commercial use).
</inputs>
<direction>
Dribbble-level UI motion. One shape, never cut: every state is the same element morphing its size, radius and color while its content swaps with a short blur. A cursor drives every change with real clicks and drags. Light warm-gray canvas, black and white components, one clean UI font (Geist). Springs everywhere, a tiny overshoot at most. The camera zooms so each state fills the frame. The last frame is the first frame, so it loops.
Banned: bouncy easing, particle bursts, glows, gradients on UI chrome, mismatched icon strokes, dead time, anything that looks like a template.
</direction>
<structure>
120 BPM, 7 bars, something happens on every beat.
Button → loader → check → dynamic island → music player with a play/pause morph → scrub the progress bar → it becomes a volume slider that stretches when dragged past max → a toggle flips on the beat → the knob becomes a liquid tab indicator → the tabs open into a chart that draws itself, with a tooltip on hover → it collapses into ⌘K → type to filter → enter → toast → back to the button.
</structure>
<build>
1. One HTML file, square 1440x1440. Every style is computed from time inside seek(t): no CSS transitions, no timers, no state carried between frames.
2. Springs are closed-form step responses. A value that changes target many times is the sum of one spring per change, so it stays a pure function of time.
3. The tab indicator's two edges ride different springs, so the leading edge stretches ahead of the trailing one. Same trick for the toggle knob.
4. Drags are direct manipulation: while the cursor is held, the value is computed from its position. On release it springs back from wherever it was.
5. Analyze the song with numpy for the beat grid and start on a downbeat. Place every UI sound by its measured peak.
6. Render with Playwright: 4 subframes per frame, blended with ffmpeg tmix for motion blur at 60fps.
7. Render one frame per beat before the full render. Fix anything off the grid, cramped or hard to read.
</build>
<gotchas>
Never put will-change on anything the camera scales or the text renders blurry. Text that swaps inside a morphing container needs its own enter and exit timing or it overlaps. Make the last frame identical to the first, cursor position and speed included, or the loop stutters.
</gotchas>
<start>
Ask me for the inputs, then show me the state list on the beat grid before you write any code.
</start>
Just trained Tev1 0.8B, a tiny Jev-like classifier.
Here it is running completely locally on my mac with @ollama & classifying some tasks.
It's extremely fast: only ~50ms E2E latency. Video is not sped up!
Releasing weights & benchmarks very soon so you can try it yourself :)
IBM built a retriever that hallucinates 65x less than fine-tuned RAG systems.
no vector database. no embeddings. no re-ranker.
Right now, if you want an AI to read a massive document, you use Retrieval-Augmented Generation (RAG).
But standard RAG does something brutal. It takes a beautifully structured 500-page manual and throws it into a blender.
It chops the text into arbitrary, fixed-size chunks. It strips away the chapters, the sections, the hierarchy.
It throws away the map and asks the AI to find the treasure.
A new paper just introduced STAIR, a method that fixes this massive blind spot.
Instead of shredding documents into random chunks, STAIR uses the document's actual structure, its Table of Contents, as an addressing scheme.
The generative retriever pulls information against the real hierarchy of the text. It remembers where things actually live.
The benchmark results are staggering.
STAIR hit an 82.6% Recall@1, completely destroying traditional methods like BM25 and standard Dense Passage Retrieval (DPR).
But here is the most important metric for any business running AI in production:
Hallucinations plummeted to under 0.05%.
Almost zero.
By giving the AI back the structural context, the system stopped guessing and started retrieving with lethal precision.
Jev Founder, Diogo Almeida, just released a 12-page PDF on how to use Jev with LLMs
It is more useful than most paid AI courses:
this is a 10-step blueprint on how to build a faster, cheaper and more controllable AI system around Claude, Codex, Grok or any other LLM:
step 1 → split the responsibilities: the LLM generates, Jev makes bounded semantic decisions and deterministic code keeps authority
step 2 → build the state: give Jev the current request, relevant evidence, policy and proposed action instead of sending the entire conversation
step 3 → choose the right primitive: Choice selects a route, Score evaluates an ordered rubric and Noul returns the probability that a statement is true
step 4 → replace giant evaluation prompts with atomic questions: intent, urgency, evidence, risk and scope become separate typed decisions
step 5 → put Jev before the LLM: select the context, tools, provider and workflow before paying for an expensive generative call
step 6 → give the LLM a bounded job: once Jev selects the route, the model receives only the instructions, files and tools required for that branch
step 7 → put Jev after the LLM: check whether the result answers the request, uses sufficient evidence and stays inside the permitted scope
step 8 → route by confidence: high-confidence low-risk cases proceed automatically, uncertain cases request more context and consequential actions go to review
step 9 → batch independent decisions: ask multiple Choice, Score and Noul questions over one shared state instead of creating another LLM call for every judgment
step 10 → record the complete decision receipt: state version, question, probabilities, selected route, model, latency, outcome and human override
most AI courses teach you how to write a bigger prompt
this 12-page guide teaches you how to build the control system around every prompt
the result: smaller contexts, fewer unnecessary LLM calls, safer tool execution and decisions you can actually inspect, test and improve
Send this PDF and the original Jev article to Claude Code or Codex and start rebuilding one expensive LLM decision at a time ↓
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task
At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6 points against Opus 5 (60) and 4 points against Claude Fable 5.1 (62).
Anthropic has cut Opus pricing to $4/$20 per million input/output tokens, from $5/$25 for Opus 5, and cache reads to $0.20 from $0.50. Even with those reductions, Opus 5.5’s Cost per Task is $13.04, above Opus 5’s $10.79, because it uses substantially more tokens.
Key takeaways:
➤ Improves across all three Coding Agent Index evaluations: Terminal-Bench 4.0 rises to 63.1% from 54.5% for Opus 5, DeepSWE v1.1 to 68.4% from 62.5%, and SWE-Atlas-QnA to 66.4% from 62.1%. The largest gain is on Terminal-Bench, at +8.6 percentage points.
➤ The top score comes at the highest Cost per Task: Opus 5.5’s Cost per Task is $13.04, up 21% from Opus 5 at $10.79. It uses about 15.6 million tokens per task against 11.4 million for Opus 5, including about 2.4× as many output tokens.
➤ Extends the Coding Agent Index vs Cost per Task Pareto frontier: No lower-cost model in our comparison matches Opus 5.5's score. It moves the frontier upward at its high-cost end.
Other model details:
➤ Pricing: $4/$20 per million input/output tokens, down 20% from Opus 5. Cache reads cost $0.20 per million, down 60% from $0.50.
➤ Evaluation setup: Claude Code at max effort, measured on DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA. The Coding Agent Index gives each evaluation equal weight.
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: https://t.co/zwi9JOHKGx
💻 Code: https://t.co/rsHRYCGR8I
🗣️ Discord: https://t.co/Uqtdefvo3J
🤗 Data & Models: https://t.co/wdSWGGO3hu
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
@Av1dlive Ive buid a versión of this that You can plug into any harness, it gived auto mode With all the things You describe that You can plug. With also dangerous action protección. I'm gonna release it today hope You can check it.