I just published An End-to-End NLP Question Answering task: From Data to Deployment. https://t.co/7Xoqo77ywR
This article contains a step-by-step approach to solving an NLP (question-answering) problem on Kaggle.
Agentic workloads generate long, expanding context across repeated model calls, tool execution, and parallel subagents. Much of this context remains stored in the KV cache while an agent waits on external tools or pending operations.
Efficient serving relies on preserving and reusing this PyTorch tensor state while maintaining sub-second latency targets across concurrent sessions.
Read our latest blog from @nvidia to learn how NVIDIA Dynamo integrates its router, inference engine, and KV-cache manager to deliver seamless end-to-end serving for agentic workloads: https://t.co/sf3urp1vzL
Today we release Open d1: two open-weight multimodal models in our d1 decision model family.
> d1-3B: text + vision
> d1-omni-600M: text + image or text + audio
> Real-time decision making anywhere, from data centers such as @nvidia DGX to RTX workstations to Jetson at the edge.
1/
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5.
Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta.
The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API.
Claude now works inside Google Docs, Sheets, and Slides, and those files also open inside Claude.
In Google Workspace, Claude sits in a sidebar next to your file, reads what you have open, and edits it in place. You can approve each edit before it lands.
Introducing EmbeddingGemma 2, a new open multimodal model that sets the standard for on-device efficiency.
- our first open, natively multimodal embedding model
- handles text, code, image, video, and audio tasks within a lightweight, modular 740M parameter form factor
- ideal for offline, privacy-first RAG when paired with Gemma 4
- outperforms some specialist models more than twice its size
Weights available now on Hugging Face.
If you live in the terminal, this one’s for you.
Watch how to start tasks by voice, manage agents across projects, and explore another direction in a separate worktree with the Codex CLI.
Announcing d1 with vision. 👁️👁️ Our first decision model now supports images, text or both as inputs. We tested d1 against GPT-6.1 Sol and Claude Opus 5.5 on six real applications, from filtering support tickets to inspecting circuit boards. d1 matches or beats GPT-6.1 Sol on four of them. It costs 19x to 200x less than both models and answers significantly faster on every task.
> probabilities for yes/no, choice, or score questions
> one forward pass, without generating tokens
> text decisions in 200 to 300 ms
> Liquid API: https://t.co/HxYWoaUnAU
🧵
We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view.
pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench.
https://t.co/tkBUEjWVio
Introducing Cohere Embed 5: our new state-of-the-art family of embeddings models.
Get frontier capabilities with Embed 5 Pro or low-latency performance with Embed 5 Fast.
Embed 5 Pro is our best retrieval model yet. It achieves the strongest average score of any model we measured on ViDoRe V3, beating out Voyage 4 Large, Gemini Embedding 2, and Jina Embeddings v5.
It also excels at image retrieval, parsed PDFs, and is trained on over 100 languages, meaning enterprises receive a wide range of capabilities within a single strong model.
Your LLM endpoint works. But how does it perform when traffic increases?
NVIDIA Dynamo AIPerf helps you measure TTFT, ITL, latency and throughput at scale, then test with realistic traffic patterns you can reliably repeat.
Read the blog: https://t.co/Z9TroqIcOK
I tried using chat gpt after a long time for learning some concepts and I feel it's explanation is much more easier to understand than claude
Claude is too much verbose and using complicated english whereas gpt sounds far more human and easier to comprehend
we launched the most comprehensive ai performance engineering repo in the world last week
now we'll be posting every single resource
this is Wafer's ai performance engineering series
save this as your starting point. links in thread 🧵
part 1: "All About Transformer Inference" from How To Scale Your Model.
- the authors cover the computations, memory traffic, and serving decisions behind transformer inference:
- arithmetic intensity of linear layers and attention across prefill and decode.
- kv cache sizing by layer count, kv heads, head dimension, sequence length, and precision.
- the compute/hbm bandwidth crossover and how batch size and quantization shift it.
- decode latency and throughput bounds from parameter bytes, kv bytes, and hardware bandwidth.
- weight reuse through batching and diminishing throughput gains as kv traffic grows.
- gqa, kv quantization, and PagedAttention, including the memory costs each addresses.
- model sharding, kv placement, and collective communication overhead.
- continuous batching, prefix caching, and disaggregated prefill/decode.
the worked problems connect model dimensions to deployment decisions like memory capacity, workload distribution, and expected performance under the stated assumptions.
We finished the Training Agents series. Six live sessions over six months, from evaluating agents to training them inside real environments. All of it is on the Hugging Face YouTube channel and all of the code is open.
Here's what we did and who made it happen:
1. Agentic Evaluations WorkshopWhere agent evals actually stand, and why benchmark scores don't match what people see in use. With Avijit Ghosh and Nathan Habib (Hugging Face), Arvind Narayanan (Princeton), Pierre Andrews (Meta), J.J. Allaire (UK AI Security Institute) and Mahesh Sathiamoorthy (Bespoke Labs).
2. RL for Agents Workshop Environments, rollouts, reward design and the inference bottlenecks that appear when you move from RL for LLMs to RL for agents. With Lewis Tunstall (Hugging Face), Will Brown (Prime Intellect), Ofir Press (Princeton) and Alex Zhang (MIT CSAIL).
3. Training Agents 1: SFT on agent traces Public coding-agent traces turned into prompt/completion data, a TRL + LoRA fine-tune on Hugging Face Jobs, metrics in Trackio, and an honest look at what the first eval numbers can and cannot tell you. Joined by Sergio Paniego and Quentin Gallouédec.
4. Training Agents 2: Distillation Off-policy, on-policy and self-distillation for moving capability from a teacher into a smaller coding agent.
5. Training Agents 3: Reinforcement learning GRPO after SFT: group sampling, verifiable reward functions, reading the reward/KL/length curves, and three experiments, one of them with a deliberately gameable reward so we could watch the hacking happen.
6. Training Agents 4: From reward functions to environments The reward stops being a function and becomes a place the agent acts in. We walked the reset()/step() contract from Gym to LLM agents, built an OpenEnv environment and pushed it to the Hub, plugged it into TRL's GRPOTrainer, then trained a real coding agent (OpenCode) through Harbor with AsyncGRPOTrainer on Hugging Face sandboxes.
The series has passed 300k views. Thank you to every speaker, to the TRL team, and to everyone who showed up live with questions.
Playlist: https://t.co/DYE1pvnywV