To truly understand AI, you need to understand evals. And they're not as scary as you think.
I had @Vtrivedy10 take me to school on Evals and it all clicked.
Here are notes from our call:
Easy Mode: WTF is an eval
- An eval is a test to see if what an ai agent did is right
- The sauce in “right” is encoding what you / your org thinks is right into software
- Example Viv took me through: GTM agent that logs the customer meeting, writes follow-up draft and stops before sending
- In this context an eval can be checking if the customer meeting was logged, if the follow-up draft was written, if the agent paused before sending
- Evals are actually not new, they’re just an evolved version of unit tests, something that’s always existed in software
- Recipe for setting up an eval is actually quite simple: agent + system + tasks + verifier
- In the GTM agent example, you have your agent, read/write access to Salesforce, a set of tasks (like update lead status, log call, write follow-up), and a verifier (tool whose job is to determine if task was done right or wrong)
- Verifiers can be a simple Python function or content/prose checks
- Rubrics are a great way for codifying what good looks like for non-verifiable tasks so verifiers can still work
- Once rubrics created if agent doesn’t perform, root cause could be: checklist incomplete, model not smart enough, bad/incomplete context or prompt not good enough
Hard Mode: WTF are Agent Environments
- Environment = place an agent does work and we measure how it performs
- Environments became necessary as we shifted from ChatGPT era of text in text out to agentic era of prompt in, real world work out
- Example: for Claude Code/Codex, when they do work on your computer, the environment is files, browser, email, apps, permissions
- When you build agent environments today = recreations simulations of existing tooling that look real to the agent but aren’t real (won’t overwrite prod)
- Basically you want to be able to test how performant your agents are without f’ing with live systems/data in your company
- Needs when building agents + evals: (1) manage tasks by team (GTM ≠ SWE), (2) team-specific environments, (3) run at scale without building all infra yourself
- Popular tooling to set up agent environments is an open source framework called Harbor, which provides all of the agent environment abstractions needed to run evals so you don’t manage everything, but can go low-level when needed
- Harbor primitives cover: environment building, verifier building, instruction building, sandbox infra
- Beginner path: grab an existing Harbor-format eval → ask Claude Code/Codex to explain → tweak for your agent
- LangSmith Engine: make evals at scale; UI for people who can judge good/bad outputs and want to give feedback without living in code
- Validate environments by running multiple real agents/models in them (e.g. frontier + GLM + small Qwen) and looking for weird patterns
- Environment engineering is iterative like agent engineering; humans build and review outputs
- To do evals at scale you need an eval suite. An eval suite is a collection of agent tasks that you’re running as a company. Each task is comprised of an input prompt, a verifier, and the environment.
- Anytime new model comes out, you’ll run your eval suite (or a selection of tasks from the suite), and weight the tradeoffs of cost vs. speed vs. intelligence
God Mode: WTF is a self-improving agent
- Continuous cycle of agent in the real world → production data → turn production into evals (and use evals to improve) → continuously improve the agent
- Hot take: if your team is really good at turning production data into evals, you can continuously improve agents
- Evals exist to find failures; then change the agent so that failure never happens again
- Two improvement buckets: harness engineering & fine tuning a small model
- Harness engineering: (cheap place to start): tweak prompts, tool definitions/descriptions, model choice, model combos (e.g. Sol+Opus, GLM+Opus)
- Fine-tune (on your narrow production task): open models + fine-tune infra exist; can match Opus on *that* task at ~10x cheaper; GTM agent doesn’t need frontier math
- God mode = collect real production data → turn into evals/environments → choose how to improve (harness first; later own fine-tuned intelligence for specific tasks)
- Bottleneck often human updating verifiers/rubrics
- Two questions for autonomy: how to spin faster; can it run without a human in the loop?
- Goal: minimize human contact points on the loop
- Agent Traces are the first practical step for orgs using agents
- Trace = log of every action: tool calls
- Humans can’t reason well about what an agent *will* do; looking at live behavior (aka traces) makes right/wrong obvious
- See a bad outcome → read traces (receipts) → find the wrong step → change agent so it doesn’t repeat
- Human touch points shrink with every smarter model release
- Today’s human role #1 (up front): explain to the mining/improvement agent what good vs bad looks like (prompt of behaviors to look for); mostly upfront with some edits over time
- Then an agent can find failed tool calls, instruction-following failures, etc., and propose improvements (new prompt, new/different tool, etc.)
- With a good eval suite or good prod bad-behavior measurement, an agent running overnight can be very good at proposing fixes
- Human role #2: when turning production into evals/environments, review what the *company* cares about
- Highest leverage organizational act: humans writing down what good looks like
I tried @jack's Buzz.
It's like Slack + OpenClaw + Herdr + but with some really unique features that people are sleeping on.
The video below shows how it works, and some of my thoughts on the process and platform, e.g.:
- Create and interact with agents on top of any harness (claude code, codex, pi, etc.)
- Choose which models agents use, including local ones
- Agents can delegate work and work in parallel in git worktrees
- Agents are first-class citizens and work like humans (creating channels, delegating, access to chat history)
- You can share AI compute within a community
- It's completely open-source and decentralized
Things I like:
- Delegating work in chat feels natural: tag an agent, it replies in a thread with status updates as it e.g. compiles, commits, and deploys.
- Shared compute: relay owners can share local compute with members, so a community could pool funds for one beefy machine running a local model and everyone uses it.
- It's built on Nostr, an open protocol already tied into Bitcoin Lightning so I can imagine communities tipping each other or paying for compute/agent tasks with instant zero-fee micropayments in the future.
- It ties together things like OpenClaw, an agent manager, and Slack-style chat into one tool.
Things I didn't like:
- You can't see what the agent is doing in a terminal. The activity view exists, but if you're used to watching a session run, this UI feels a bit abstracted. A terminal view would be great.
- It feels slower than running a session in Claude Code, though no evidence to back that up. For that reason I found myself doing one-off tasks in the terminal instead.
Verdict:
- I really like it so far and can genuinely imagine working with a team this way.
- It doesn't feel ready for big, complex tasks yet. For shallower tasks, it's perfect.
- The shared compute + Nostr/Lightning angle is what really separates it from every other agent manager for me, and I think that future is coming.
Introducing OpenSEO, the open source alternative to Semrush and Ahrefs.
I started building it out of spite because the existing tools were too expensive, bloated or scammy.
Since then, we've gotten 4k stars on Github and are launching on Product Hunt today.
Hope you support!
Just launched a new website where I’m sharing all the AI agent-specific components I build as ready-to-use code snippets. 12 are already available.
👉 https://t.co/ySHwCL1yBw
Still in beta, and a lot is WIP, but I’ll be adding a lot more over the next few weeks to build out a solid free library.
Want to contribute? Feel free to DM me 🫡
New tools I keep coming back to. Save this 👇
- https://t.co/wQrrKyScQy - AI prompts for animated websites
- https://t.co/FgTPEqdeQ4 - free animated component library
- https://t.co/DBIDZdVb86 - pure CSS text animations
- https://t.co/Hk6ovLEr1l - spring physics UI motion
- https://t.co/wv2C0KXUCr - build terminal UIs with TypeScript
- https://t.co/nHTq0N41dq - name any UI element instantly
- https://t.co/NRWoHNzeI0 - logo, color & gradient tools
- https://t.co/CLb1khhxQZ - 3D product visualizer in browser
- https://t.co/JtOAiFEBJy - Claude skills for web design
What am I missing?
Introducing ARDY — real-time, controllable 3D human motion generation
Streaming text prompts + flexible kinematic constraints — root waypoints, full-body keyframes, joint targets — all at interactive speed
To be presented at SIGGRAPH 2026, Paper & demo
https://t.co/ujS94CQAto
1/5
🚀 Introducing Lingbot-World-2, our newest real-time world model.
Richer interactivity, sharper visuals, and advanced camera controls that make exploring what you create feel truly alive.
Available today via API, exclusively on Reactor.
Try it now: https://t.co/rRGqMrhvea
I updated EasyPeasyEase so it detects the BPM in uploaded tracks and automatically sets the duration of each segment to land on the beats.
Its a free browser tool for adding speed curves/easing to videos. That's all it does.
Link below.
New interactive world model called AlayaWorld.
Generates playable video worlds with real-time camera control and editing with prompts.
720p/24 FPS, 60s+ rollouts, and long-term memory that helps places stay consistent.
https://t.co/e3W9ZB60oQ
Introducing the Cinemagraph LoRA for LTX-2.3.
Turn a still image into a continuous looping video where only one element moves, everything else stays locked.
Most video models fight to keep backgrounds still. This one was trained to do it by default.
day 20: shipped agentic onboarding
after a few tries with @stedelmanto, this one finally clicked: clear, clean, and a little magical
find me a better SaaS onboarding flow and I’ll buy you coffee
@linear doesn’t count. @emilkowalski doesn’t count. we're designing a different game, but huge respect to their flow
📢WorldMesh is accepted to #ECCV2026, and we're releasing the code today! 🎉
Led by @mschneider456: navigable, multi-room 3D scenes from a text prompt, with a mesh scaffold conditioning image diffusion for global consistency + photorealistic detail
👇
https://t.co/8fXCl2flIu
Klein 9B Sun Direction LoRA
Nice LoRA that lets you move the sun to a specific spot in the sky just by pointing at a reference ball.
https://t.co/rHJJ0o0Zz7
📢 GenRecon Code Release 📢
Few images in → complete, high-fidelity 3D scene out!
GenRecon builds a generative prior on full scenes, resulting in unprecedented 3D reconstruction quality.
🔗 https://t.co/1jvzgC7aBO
🌐 https://t.co/sbz1Nb0ptN
📄 https://t.co/VqL3flGBGj