the biggest lesson that i took away from our jev webinar w/ @allietheicon is that jev actually enables REAL TIME inference
this is because jev is fast and cheap. you can:
- do content filtering on a stream in real time
- add guardrails that don't really impact ux
- monitor (and steer) agent trajectory w/o prohibitive cost
there's a lot more to say here but i find these examples in particular very compelling
curious what others are doing in real time w/ jev!
webinar here icymi: https://t.co/CRtxdvV4cJ
i normally would not post this late, but wanted to share what @huntlovell and I have been jamming on.
@typesafeai just released Jev, a new type of classification model. it's really popular right now because it's ridiculously fast and cheap (up to 200x / 400x reductions vs comparable LLMs on classification tasks).
learn all about Jev and how to build it into your harness!
We’re looking for someone to lead SmithDB at LangChain!
SmithDB is the database we’ve built from the ground up to power LangSmith. It’s optimized specifically for the access patterns that emerge when you’re storing and querying enormous volumes of agent traces.
It’s serving production traffic at significant scale with great performance, but it’s still early. The problems we are solving are genuinely fun and very hard: indexing, query execution, compaction, ingestion specifically for agent observability + making all of this fast and cost-efficient at massive scale.
We’re looking for an exceptional engineering leader to take ownership of this critical project. You’re likely a great fit if you’ve built database or distributed systems, led strong engineering teams through technically demanding projects, and have exceptional execution and project management skills.
Apply here or DM me directly: https://t.co/Pqy0AfYCFO
Learn more about SmithDB here: https://t.co/bjus8tSSfG
Life update: I’m doing @ycombinator S26!
Building stock market infra for agents.
4 years ago: I started building in public with the goal of shipping + improving by 1% daily.
2 years ago: I started @findatasets as a side project to fix stock market data.
5 months ago: I took a leap of faith, quit my job and went full-time on Financial Datasets.
Today: hedge funds and developers building agents use us as their market data layer.
Future is bright.
Thank you for being on this ride with me❤️
Special episode with @bentannyhill on the Max Agency podcast. A month ago, his team shipped LangSmith Engine, our agent that hunts through your agent's failures, prioritizes issues, and drafts the fix.
We dive into the architecture decisions -- from how we used sandboxes to how we created subagents. And we also discuss an interesting challenge -- how we were able to build evals for an agent that never stops running.
Check out the full conversation ⤵️
⏯️ YouTube: https://t.co/2yEjgUNUu4
🎧 Apple: https://t.co/V69kYf7wTm
🟢 Spotify: https://t.co/lQWG9UMNVl
my fave question, talked about this coding agent Eval+Improvement loop infra + UX in my AIE talk yesterday!
biased but LangSmith is the best spot to Eval + continuously improve your coding agents, and we want to make it better so would love any feedback :)
we eval all of our coding agents there --> supports Codex, Claude Code, OpenCode, Deep Agents, Pi, etc all into Tracing, sandbox infra for running evals, metrics + datasets for storing everything, and
imo the hardest parts of doing coding agent evals are:
1. having infra to easily store, update, run, and share coding agent evals
2. building a clear understanding of agent behavior & failure modes across all of the rollouts. you can use a coding agent + the langsmith-cli to look through all the data
or if teams want a managed, they can use LangSmith Engine to read every single trace from every eval, see what went wrong, prepare a report for you, and propose new changes directly in the code to fix issues
3. build new and better evals for your coding tasks. there's many ways to do this, but we find that looking through existing failure modes from eval runs and prod is a really good grounded way to measure where agents lack today and turn that data into
we recently launched a LangSmith x @harborframework integration to double down on making it super easy for LangSmith users to improve their coding agents over time with built-in infra for running large-scale containerized evals so you can fully reproduce your tasks and read all of the traces
https://t.co/LNloqsbt1k
i think most evals in the future will be shaped as environments and letting agents do real work in them, because agents are doing way more complex things and we need the eval shape to mirror how they work with us
if there's any experience you're looking for would love to chat
we have a ton of work underway on making each part of coding agent evaluation + improvement easier over time (as teams eval their coding agents for months and years)
as an aside --> evals are literally the training data for agents. the behaviors that we measure and reward in the evals directly get transferred as model/agent behavior as we hill-climb them. so i think there's nothing more important than making it easy to build good evals over time :)
Building voice agents can come with tradeoffs.
💬 Better convos w/ speech to speech models
Vs
🥪 More reliable harnesses w/ sandwich architecture
How to build a voice research agent with both:
✅ Gemini Live: low-latency, natural-sounding convos
✅ Deep Agents: long-running research tasks
✅ LangSmith: full tracing & observability
another banger from Sydney! i think this whole hierarchy of loops is still super early but some primitives we know work
ex: verification as a primitive is so ridiculously important for non-slop semi-long-horizon work, it’s worth spending days to weeks making sure the distribution of outcomes you want from your agent are verifiable in practice by your system
Detecting issues in production agent traces is hard. You have to do it cheaply (because of volume) but also accurately (or too much noise)
We post-trained our own model for this. SOTA accuracy, at ~10-100x cheaper rates than frontier models
Try it out: https://t.co/xSK0Cd8fxt