Stop sending your coding agent into a repo blind.
Agent Dispatcher is a repository retrieval and context pack for Claude Code and Codex users working inside existing codebases.
It helps you explore, debug, build, or review changes with more relevant context by combining code paths, symbols, source text, code relationships, and Git history into a bounded context packet.
Key features:
• Relevant-code retrieval – selects source excerpts using paths, symbols, source text, stack-trace frames, code relationships, and Git history
• Project knowledge reuse – uses source-linked project maps and private caches to reuse unchanged evidence
• Task-to-role matching – routes substantial work to specialist guidance while keeping small, obvious edits direct
• Visible verification – records passed checks, failures, zero-test runs, and work left unverified
• Flexible setup – installs as a Codex skill or a Claude Code plugin, with no additional API key needed for normal use
It’s open-source (MIT license).
Link in the reply 👇
A skill that works once isn’t necessarily helping your agent.
Caliper is an open-source evaluation tool for builders who want to test agent skills with Claude Code, Codex, Pi, or Hermes.
It helps you measure whether a skill earns its context by running it repeatedly in a real agent, then comparing the results with the same skill removed.
Key features:
• Repeated real-agent runs – run each task k times to track a skill’s success rate
• Separate activation scoring – see whether the agent chose the skill separately from whether the task succeeded
• Skill ablations – rerun the same tasks without a skill to measure what it adds over the bare agent
• Neighbour-skill tests – declare competing skills and assert which one should handle a prompt
• Task-level comparisons – diff full and ablated runs task by task, including token use and wall time
It’s open-source (MIT license).
Link in the reply 👇
demo walkthroughs are a great job to hand your Hermes Agent.
how to:
~ the clicks: browser automation and computer use drive the pages and apps your demo walks through, with any tool-capable model
~ the edit: the terminal tool runs ffmpeg, so it can cut, join and caption the clips from your instructions
~ the delivery: mp4 and mov files come back as attachments in Telegram, Discord and the other chat apps
~ the repeat: once a take works, it can save the steps as a skill and run them the same way next time
DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video
Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim
https://t.co/89t8PsgEId [𝚌𝚜.𝙲𝚅 𝚌𝚜.𝙲𝙻 𝚌𝚜.𝙻𝙶]
💬Code: https://t.co/hEElUjeHSa
Stop manually optimizing inference. Let an agent bash its head against the benchmark to rewrite forward passes. Faster tokens, less compute.
Grab the SIE repo and try it yourself: https://t.co/1EsK5MDKCB
#AI#MachineLearning#LLM
Your best regression tests are the ones you build from real failures.
Record real runs → review them → cluster failures by frequency × severity → write one cheap evaluator per failure class.
Then rerun those cases after every fix or change.
Your regression suite grows from how your agent breaks rather than how you imagine it.
Benchmarked speculative decoding on a live GLM-5.3-Flash deployment spanning 2× DGX Spark, without taking it down.
NVIDIA's specdec_bench loads the model itself, so it can't measure a server that's already running. So I forked it and added a client mode. 🧵
Ability of a Structural World Model to Detect Cryptic Pockets from Apo Structure
https://t.co/aZJCOKk1JV
Summary:
This study presents a method for identifying cryptic drug-binding pockets directly from a single apo protein structure, without requiring molecular dynamics simulations, conformational sampling, co-folding predictions, or external pocket-detection tools. The approach uses a proprietary structural world model that generates a per-residue latent representation from protein coordinates. This latent state is read out as a cryptic-lining score, suggesting the model implicitly captures conformational flexibility that other methods must explicitly simulate.
Predicted pockets are constructed as residue sets rather than fixed geometric spheres. High-scoring residues seed candidate pockets, which are expanded into distinct, non-overlapping predictions through a residue-growth strategy and geometric non-maximum suppression.
On CryptoBench (231 test proteins), the model achieves 84.8% top-1 and 95.2% top-5 localization accuracy. Similar performance is observed on a CryptoBank subset, with 84.6% top-1 and 99.0% top-5 accuracy. Residue-level classification is strong (AUC = 0.8465), although exact pocket-boundary recovery remains more challenging under stricter overlap criteria.
The method generalizes well to unseen proteins, recovering the known WRN helicase allosteric site at rank 1 across multiple apo structures after all WRN proteins were removed from training. A key advantage is its ability to place the true cryptic site at rank 1, addressing a common limitation of ensemble-based approaches.
Compared with single-structure baselines such as P2Rank, DeepPocket, PocketMiner, and fpocket, the model achieves substantially higher top-1 hit rates. It also complements the ensemble-based method OpenDDE, identifying many cryptic sites missed by OpenDDE while retaining all of OpenDDE’s top-5 hits. Overall, the work demonstrates that latent representations learned by a structural world model can effectively detect and rank cryptic pockets from apo structures alone, with performance improving as training data increases.
#DrugDiscovery #CrypticPockets #BioAI #AIforScience #ProteinStructure
Stripped binaries hide the context you need. This repo helps put it back.
Kong is an open-source agentic reverse-engineering tool for developers analyzing stripped binaries.
It helps you turn opaque decompiler output into a more readable program database by building Ghidra context for LLM-guided analysis and writing recovered names, types, and signatures back into Ghidra.
Key features:
• Full analysis pipeline – triages functions, analyzes them, cleans up results, synthesizes context, and exports output
• Call-graph ordering – analyzes leaf functions first so callers inherit already-resolved context
• Rich analysis context – gives the LLM decompilation alongside cross-references, strings, and caller/callee signatures
• Agentic deobfuscation – includes a pipeline for identifying and removing supported obfuscation techniques from decompiler output
• Evaluation harness – scores recovered symbols and types against ground-truth source code
It’s open-source (Apache License 2.0 license).
Link in the reply 👇
vLLM vs SGLang isn't a one-time pick. It depends on how much your prompts repeat.
At 80% prefix overlap (shared system prompt), SGLang cut time-to-first-token 37%: 310ms to 195ms.
With unique prompts, no shared text to cache:
- vLLM: 1,850 tokens/s
- SGLang: 1,920 tokens/s
That's a 5% gap, within noise. Pick SGLang for shared prompts. Either works for unique ones.
Your voice agent runs on one GPU, streams STT and TTS through the same process, and starts replying while you are still talking.
Fusion-runtime wraps Whisper, an LLM, and Kokoro TTS in one pipeline.
- 490 ms processing on an RTX 3090 with a 7B model
- 991 ms stopwatch from your last syllable
- Interrupts honoured mid-sentence
- One agent file, two commands to run
https://t.co/jB9QhhG6jz
Long-running agents need more than a chat history.
Active Graph is an event-sourced reactive graph runtime for builders developing durable, stateful agentic systems.
It helps you inspect, resume, and compare agent runs by recording every mutation in an append-only event log while reactive behaviors work against a shared graph.
Key features:
• Event-sourced runtime – every graph mutation becomes an event, creating an audit trail for a run
• Reactive behaviors – function, class, LLM-backed, and relation-attached behaviors can respond to graph events
• Fork-and-diff workflow – branch a run at any event and structurally compare the fork with its parent
• Policy controls – define which behavior capabilities need approval and which mutations the runtime refuses
• Fast local start – install with pip and run the bundled quickstart against recorded fixtures without an API key
It’s open-source (Apache License 2.0 license).
Link in the reply 👇
we launched the most comprehensive ai performance engineering repo in the world
follow and save to keep up with the series. links in thread 🧵
part 6: Transformer Inference Arithmetic
Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s.
Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation:
- prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase.
- batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths.
- memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency.
- KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate.
- precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform.
- tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency.
- kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor.
- check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target.
Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Jiapeng Li
https://t.co/Yivk0jqvuR [𝚌𝚜.𝙻𝙶 𝚌𝚜.𝙰𝙸 𝚌𝚜.𝚂𝙴]
The current balance of power in open models.
https://t.co/X4bdW8pklV analysis of open source AI:
- Who leads in open model landscape
- Major players and their positions
- 34 points on Hacker News
- Detailed breakdown available
https://t.co/JscoWFXPzu
I made an update to how I run my benchmarks with @VulcanBench, specifically to the task timeouts.
A little while ago I decided to increase the timeouts to 10 hours, and well, I regret doing that.
This has made it take a lot longer to get benchmarks out, and I've found in so many cases, if a model is going to overthink, in a detrimental way, it's going to do it, pretty much forever.
There are a lot of cases where I wish I cut off a model earlier vs. letting it just spin on a task for 10 hours.
Also, I think that we're now at the point, with the current frontier, where good strong models clearly can, and should, be able to complete all my tasks a lot faster than ten hours.
So I have reduced my timeouts to 3 hours. This will both help me get benchmarks out faster, and also give models the accuracy hit that I think they deserve, if they go into a doom loop and just keep thinking and burning tokens, I think we as engineers should be a lot less tolerant of this kind of behavior because it costs us both money and time.
I also think even three hours might be too generous, so will be doing a deeper dive into the traces for my Frontier v4 runs to see what average time per task is and likely even get more aggressive on timing for some set of tasks.
The reality is, when choosing a model, you don't want to just decide based on accuracy, you want a combination of accuracy and speed/token efficiency.
More to come, but I do feel very lucky that as an independent benchmarker, I can just make decisions like this. I have nobody I have to bounce it off of, I can just look at the data, and decide, with my only motivation being giving other engineers like me, real signal to make decisions around which models and effort levels to use.
Live long and benchmark 🖖
AI agents are officially getting hands.
Microsoft just open-sourced Fara1.5 — a family of computer-use agents that can actually SEE a screen and operate it with mouse + keyboard.
No accessibility tree.
No separate parser.
Just:
Screenshot → Think → Click / Type → Observe → Repeat.
And the crazy part?
Fara1.5 comes in 4B, 9B and 27B models.
The 27B model hits:
→ 72.3% on Online-Mind2Web
→ 89.3% on WebVoyager
It can search the web, fill forms, compare products, book things, find jobs and execute multi-step browser tasks.
But the really interesting part is FaraGen1.5.
Microsoft built a scalable pipeline to generate training data using:
Environments + Solvers + Verifiers
That produced roughly 2M training samples for teaching these agents how to actually use computers.
And the repo includes the full evaluation stack too.
This feels less like “another AI model”…
and more like a glimpse at what happens when AI stops just answering your computer requests and starts operating the computer itself.
REPOO👇