We just added Skills to the Perplexity Agent API.
Agents aren't defined by a single system prompt. They are assembled from many capabilities that developers extend and compose.
For example: pair our built-in office/pdf skill with your own inline design skill and your agent produces a finished, custom-formatted PDF research report in one go.
Try our built-in skills (extended Skills Directory 🔜) or bring your own today.
Docs: https://t.co/8Xi0e0m8xp
Research PDF cookbook: https://t.co/30yfhPTTKb
A huge amount of the Anti-AI code sentiment massively overestimates the quality of human code outside of a very small set of open source and high quality company codebases. Human Slop is everywhere and can trivially be improved on by any opus level model.
We open-sourced WANDR (Wide ANd Deep Research) benchmark, check it out!
"WANDR starts from de-identified patterns observed in production usage rather than synthetic prompts. This keeps the benchmark close to the work people actually delegate: open-ended discovery, repeated enrichment, cross-source checks, and structured comparisons at professional scale"
blog: https://t.co/1a03pn1s8E
pdf: https://t.co/oBkosFTe3I
Soft scores give partial credit. Hard scores require every component for a member to be complete and correct.
The benchmark tasks and evaluation harness are available at: https://t.co/o2HX0a4RTQ
We’re open sourcing WANDR.
WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer.
https://t.co/gp2BWjFK4d
today we're open-sourcing an eval/RL environment for measuring agentic search performance. importantly, these environments were synthesized from production traces, offering a real-world distribution, with weak human supervision.
internally, we've been using these RL environments to train capable models (more on that soon), and we're scaling this paradigm to other domains and tasks, drawing on use cases from our products to cover the entire knowledge work distribution.
more details here: https://t.co/YhCVg8cVhM
The GPT-5.6 model family is now available on Perplexity's Agent API.
Every gain at the frontier compounds through Perplexity's API stack. The GPT-5.6 family sits on the Pareto frontier of our agentic research evals: higher accuracy at lower cost.
As such, we've updated our Agent API Presets. xHigh, High, and Medium (our pre-configured presets for long-running research workloads) are now better, faster and cheaper on the GPT-5.6 model family.
Check it out today: https://t.co/khbWMfgO6i
@MParakhin Tangle is a great platform for autoresearch and, more broadly, for giving agents compute for data processing and experimentation. Thank you!
ICYMI our Tangle and Tangent demos at @ICMLconf this week, Linux Foundation just published the writeup. Link in🧵
Tangle: our open-source, platform-agnostic ML experimentation tool. Drag-and-drop pipeline builder, a caching layer for fast iteration, every run and log stored permanently. Any containerized CLI can be a pipeline component.
Tangent: an autonomous agent built on Tangle that runs a Karpathy-style autoresearch loop. Gated checkpoints at each stage. Persistent memory across runs.
We used Tangent to rebuild a reranking model. Iterating on its own hypotheses, it pushed recall at 90% precision from 67.3% to 75.6%, no human in the loop between runs.
This is the best tool to run ML experiments at scale (if I say so myself :-)). It has autoreseach (Tangent) built in! We are also thinking about adding popular recipes, like finetuning Qwen, ready out of the box. Would that be useful?
The DGX Spark is an incredible piece of hardware. Running it to almost full GPU and RAM utilization and it still doesn’t have any issues with the heat. Unified memory is the way to go for maximizing token value per watt.
Perplexity's Agent API presets have been refreshed: faster, smarter, cheaper while maintaining full computer and usage transparency.
Most agentic API endpoints charge a fixed per-request rate in the name of simplicity.
This masks what's actually happening under the hood. Providers can use as little computer as possible on your request and you pay the same (regardless of what actually ran).
On Agent API, it's simple and transparent:
> Pick preset mode: fast / low / medium / high depending on task complexity
> Pay only for the inference and tools you actually use (direct model provider pricing)
> No manual tuning + no model selection required
Try it for yourself: https://t.co/khbWMfhlVQ
What did we get done this week ?
1. Perplexity Computer
2. Samsung Galaxy S26 integration
3. Upgraded voice mode in Comet
4. State of the art search embeddings
this has been the most exciting project to work on. a lot of foundational infra we've built over time (search engine, browsing cloud, memory engine, and more) all came together in one platform.
this was also built by a small group of people and a lot of AI coding agents. testing, debugging, evals, development, all automated with AI. i've never seen such rapid iteration speed before, and i'm excited to see how fast we're going to make improvements over the next days and weeks. stay tuned!
Probably one of the most exciting moment, today we (@perplexity_ai ) have an office in Potsdamer Platz, Berlin 🇩🇪!
4 MTS onboarded today! A lot of efforts from legal/HR team and much appreciate to @dnlkwk , @denis_bykov and @denisyarats for making this happen!