Anthropic engineers just shared the agent setup they use at Anthropic to keep up with their work
It’s a reference architecture for any agent that runs on a cron with nobody watching.
Here’s how it works: on a schedule, an agent on Claude Managed Agents wakes up, reads every source you gave it, and posts a short brief with only what needs your attention today.
The interesting part is how they managed to avoid the usual ways these automations break.
Memory: every run starts in a fresh sandbox, so everything lives in two stores. Preferences are read-only and re-read every run, so your rule changes apply the next morning. State is the agent’s own: a bookmark with the newest item it saw per source, and a ledger with a stable ID for everything it already reported.
Reading: it starts from the bookmark, never from “the last 24 hours”. With a time window, a late run leaves gaps and an early one sends duplicates.
Credentials: the sandbox only holds placeholders. The real tokens get swapped in as the request leaves, and only for the hosts you allowed. Nothing inside the agent can leak a key it never had.
Broken sources: if a source is down or a token expired, the run still starts, just without that source. So its bookmark stays put and the output says what couldn’t be read, instead of pretending it was a quiet day.
Posting: before sending, it checks whether today’s output already exists. A send counts only when the destination confirms it, and only then do the bookmarks move. If the result is unclear, the run gets marked “maybe posted” and nothing else changes.
What I’d copy into any agent: bookmarks plus that “maybe posted” state.
Elon Musk gave Grok Bot engineers one task: make bots that work without you. they did.
In this live talk they show what that looks like in practice:
00:46 - Pen and Lauren join the Grok Bot team
07:05 - Grok Bots debate and delegate to each other
22:10 - a Grok Bot books an entire multi-city trip solo
28:31 - agent prs auto-merge before manual review
44:03 - "it's like hiring virtual employees"
The bots don't need more prompts
They need a decision layer to route work and stop asking you for every yes
If you still approve every step, you don't have a team of agents
You have a very expensive autocomplete
Save this before you spin up your next Grok Bot
Lauren Tan, the engineering lead behind Grok Bot, doesn't use AI like the rest of us
she runs an entire team of bots that manage, build, review, and even improve each other
this is f*cking insane...
i turned Lauren's exact playbook into one prompt:
a chief of staff routes tasks, a lead delegates, workers execute and verify, and a reviewer turns mistakes into reusable skills
you don't manage every bot. you manage one.
send this prompt to Grok Bot or Claude and thank me later
full setup, workflow, and examples below ↓
SpaceXAI engineer Lauren Ian:
“99% of people use GrokBot at just 1% of its real potential. They run a single agent without loops or graphs.
I run a team of 20+ fully autonomous GrokBot agents. I have a Chief of Staff agent, a PM agent and 20+ workers. That’s the new engineering stack.”
In a one-hour session, a SpaceXAI engineer explained how to build an effective team of AI agents from scratch.
This workshop is worth more than a $500 agentic engineering course.
Watch it today, then read the article below to learn how to build your own fleet of GrokBot agents.
These sites are essential for Data Analysts
1. Mockaroo (https://t.co/0ysck85810) → generates realistic test data in seconds. I used to spend hours creating fake datasets to practice with. This thing spits out thousands of rows based on whatever parameters you need.
2. SQL Fiddle (https://t.co/uwtkA31Op8) → test your queries before running them on actual data. Saved me from crashing our database more times than I care to admit.
3. Regex101 (https://t.co/EccBz5wsdk) → makes regular expressions actually make sense. That alone is worth it. I used to copy paste regex patterns and pray they worked.
4. Our World in Data (https://t.co/BsUNZSGUUY) → clean, reliable datasets on basically everything. When your boss asks for "industry benchmarks" at 4pm, this is where you go.
5. Datawrapper (https://t.co/6yBJfM4flJ) → creates charts that don't look like they're from 2003. Your stakeholders will think you hired a designer.
6. Mode Analytics (https://t.co/TV3B14G8go) → runs SQL, Python, and R in the same place. No more switching between five different tools to finish one analysis.
These tools don't make you a better analyst, they just stop you from wasting time on things that shouldn't take time in the first place.
CC: Goodness Nwadibie. I hope it helps beginners in DA.
If you learn AI Infrastructure Engineering, you will never be unemployed again.
Stage 1: Systems & GPU Foundations
- Learn: C++, Rust, CUDA memory hierarchy, PCIe vs NVLink, thread blocks and warp scheduling.
- Practice: Write a custom CUDA kernel for matrix multiplication from scratch and profile it against cuBLAS.
- Why: You cannot optimize what you do not understand at the silicon level.
Stage 2: Transformer Inference Physics
- Learn: Prefill vs Decode phases, arithmetic intensity, memory bandwidth bottlenecks, KV-cache math.
- Practice: Profile a HuggingFace model with Nsight Systems to identify the exact memory bottleneck during token generation.
- Why: LLM inference is almost always memory-bound, not compute-bound.
Stage 3: Modern Serving Engines & Batching
- Learn: vLLM, SGLang, PagedAttention, continuous batching, chunked prefill.
- Practice: Deploy a 70B model and tune chunk sizes to maximize throughput without starving decode requests during long-context prefills.
- Why: Naive batching leaves 60% of GPU VRAM wasted. PagedAttention fixes this.
Stage 4: KV-Cache & Memory Optimization
- Learn: Prefix caching, KV quantization, CPU offloading, multi-turn cache reuse.
- Practice: Build a routing proxy that directs requests with identical system prompts to the same replica to share KV blocks.
- Why: Reused cache is free speed. It can slash Time-To-First-Token (TTFT) by 80%.
Stage 5: Quantization & Compression
- Learn: FP8, INT4, AWQ, GPTQ, Sparsity, TensorRT-LLM calibration.
- Practice: Serve a model in FP8 vs FP16 and benchmark the exact perplexity drop against the latency and VRAM gains.
- Why: Quantization is the only way to fit frontier models on edge GPUs and protect your margins.
Stage 6: Kernel-Level Engineering
- Learn: Triton, FlashAttention, CUDA graphs, operator fusion.
- Practice: Write a fused Triton kernel for RMSNorm or Softmax and benchmark it against the native PyTorch implementation.
- Why: Python overhead kills inference. Fused kernels save milliseconds that compound at scale.
Stage 7: Distributed Inference & Parallelism
- Learn: Tensor parallelism, Pipeline parallelism, Expert parallelism (for MoE models).
- Practice: Shard a 405B parameter model across 8 nodes and measure the communication overhead between GPUs.
- Why: Single-GPU inference is dead for frontier models. You must master the shard.
Stage 8: Speculative Decoding
- Learn: Draft-target models, Medusa heads, acceptance rates, n-gram drafting.
- Practice: Build a pipeline where a local 3B model drafts tokens for a cloud 70B model to verify in parallel.
- Why: 2x decode speed at zero quality cost is the closest thing to a free lunch in inference.
Stage 9: Multi-Node & Hardware Interconnects
- Learn: NCCL, RDMA, InfiniBand, NVLink, Disaggregated Prefill/Decode architectures.
- Practice: Set up a multi-node cluster and profile the network latency of tensor parallelism across nodes vs within a single node.
- Why: Network latency is the new GPU bottleneck. Disaggregating prefill and decode is the 2026 meta.
Stage 10: Cluster Orchestration & GPU Scheduling
- Learn: Kubernetes GPU operators, Ray, Slurm, MIG (Multi-Instance GPU) partitioning, KEDA.
- Practice: Build a queue-based autoscaler that spins up spot GPUs based on pending inference requests and drains them when empty.
- Why: Idle H100s burn $3+/hr. FinOps and scheduling are now core infra responsibilities.
Stage 11: AI Gateways, Routing & Observability
- Learn: TTFT/ITL SLOs, semantic routing, DCGM metrics, OpenTelemetry for LLMs.
- Practice: Build a gateway that routes simple queries to a quantized local model and complex reasoning to a frontier API based on prompt complexity.
- Why: Routing protects your margins and DCGM metrics tell you when your GPUs are silently throttling.
Stage 12: Public Benchmarks & Teardowns
- Learn: Reproducible methodology, latency/throughput Pareto curves, cost-per-token analysis.
- Practice: Publish a teardown comparing vLLM vs SGLang vs TensorRT-LLM on your specific hardware with full configs.
- Why: Public proof of hardware mastery gets you hired instantly by top AI labs.
Wrappers are a commodity. AI Infrastructure is the physics of scale.
The modern AI Infra Engineer builds the silicon nervous system for global intelligence.
Bookmark and Repost!
Meet Agent Buddy! 🤖 A desk buddy for your AI coding agents.
It shows when GitHub Copilot, Claude, Codex, Grok and others are working, need you, or are done. Run it on an ESP32 device, or use the new desktop app. No device needed!
Website: https://t.co/OVO9S0mxdN
Repo: https://t.co/feaBdDmkl0
Special thanks to @darrenjrobinson and @bradygaster for their PRs!
This is the first time I've actually preferred with drive as Codex as main and Claude as secondary for an extended period of time. Amazing how good Sol 6.1 is!
This week's Azure update is up!
📽️ https://t.co/Z21GSAU2sc
📄 https://t.co/eiWH7s8Qfv
00:00 - Introduction
00:13 - New videos
00:49 - AKS bare metal on Ubuntu
01:43 - AKS Anyscale
02:23 - AKS NAT GW standardv2 default
03:00 - App Service Java 8,11 and 17 retirement
03:12 - ADE retirement
04:18 - Microsoft Dev Box retirement
04:52 - App GW WAF IPv6
05:20 - WAF exceptions
06:44 - API-M A365 integration
08:54 - PostgreSQL flexible in East US 3
09:06 - PostgreSQL Elastic major version upgrades
09:52 - Enabling the Bulk admin role for SQL Server on Linux
10:25 - PostgreSQL flexible & elastic Azure Backup support v2
11:39 - OneLake catalog table discovery
12:20 - SQL Server on Azure VM in Azure Bleu
12:41 - Azure HorizonDB additional regions
13:17 - Azure SQL DB always encrypted with Intel SGX enclaves retirement
13:41 - MAI Code 1.1 Flash local
15:56 - Close
#azure #cloud #microsoft #cloudcomputing #microsoftazure #azurecloud #azureadministrator #azurearchitect #microsoftcloud #azureupdate
Scale cloud native workloads in Azure Linux. In this episode of #AzureFriday, @shanselman and Poorvi Narang explore Azure Linux across WSL, Azure VMs, and AKS, plus its minimal footprint, built-in securities, and Azure Linux 4.0.
Watch now: https://t.co/aiVxK02UE2
PerformanceMonitor 3.10.0 is out now. Free.
The web dashboard gets FinOps, job history, mute rules, and query plans and deadlock graphs right from the grids. One time picker everywhere: type "last 3 days" and go.
https://t.co/DmoRBuyqsP
We're excited to have Aater Suleman, Co-founder of Vixul, Hannah Marlowe, Senior Director of AI Forward Deployed Engineering at Databricks, and Dan Maloney, CEO of LandingAI, joining the panel "Secure AI Needs Forward Deployed Engineers" at the Agentic AI Conference! 🎙️
Secure, compliant AI doesn't happen through principles sent to an organization. It happens through people embedded in it. Vendors ship guardrails, threat models, and compliance checklists. Enterprises run legacy systems, unique risk appetites, and their own approval processes.
Somebody has to turn one into the other, and increasingly that somebody is a forward deployed engineer. It's why the role keeps appearing at the biggest names in enterprise AI.
Together, our panelists will dig into how FDEs translate security and compliance requirements into real-world deployments, what the job demands day to day, and whether it can scale. Expect a candid debate on whether forward deployed engineering is here to stay, or a stopgap until the tooling catches up.
📅 November 9–10, 2026 | 9 AM – 2 PM PT | Virtual
🔗 Register now to save your spot: https://t.co/7BX8lCSH1r
#datasciencedojo #agenticaiconference #futureofdataandai
it's built on top of Claude Managed Agents, doing code generation in the cloud sandbo, using Sonnet 5.5 each run costs ~$0.75
you can try it yourself here: https://t.co/Q2zzOYeX9k
whoops I hit send and went to a meeting, didn't realize the repo wasnt public- it is now!
it's a chrome extension I use every day now https://t.co/OroAeTafj0