About six months ago, Jean-Baptiste Tristan mentioned an idea that really caught my imagination: a policy language for AI agents based that can reason about history. About order, time, and rate.
Today, we've made that idea real, in Dogwood: https://t.co/ALV9xUgMw6
Nsight Python 1.0 is here 🎉
GPU performance analysis in a few lines of Python: sweep kernel configurations, collect Nsight Compute metrics, compare implementations, derive custom metrics, and generate DataFrames, CSVs, and publication-ready plots.
https://t.co/9V006bqXeV
Finding 1: The longer models reason, the more incoherent they become. This holds across every task and model we tested—whether we measure reasoning tokens, agent actions, or optimizer steps.
@clare_liguori Thanks for all the work you folks are doing. I woke up this morning and bumped into strands-vllm ...before that, strands-evals was released without any fanfare
Devin Review is currently free and works on public or private GitHub PRs. You can use it in three ways:
1. https://t.co/cXMHbmFi2K
2. Swap github for devinreview in the PR url
3. npx devin-review {pr-link} - run this command inside the repo of the PR you want reviewed
Check out the docs for more details: https://t.co/DrRD3EjF57
I have spent the past year assisting customers in building production-scale LLM applications. @sh_reya and @HamelHusain have done an excellent job of conceptually explaining the path to production, which involves navigating the "Three Gulfs."
Congrats to DeepSeek on producing an o1-level reasoning model! Their research paper demonstrates that they’ve independently found some of the core ideas that we did on our way to o1.
New post re: Devin (the AI SWE). We couldn't find many reviews of people using it for real tasks, so we went MKBHD mode and put Devin through its paces.
We documented our findings here. Would love to know if others have had a different experience.
https://t.co/DDqzoAXKkl