we rolled out jev for @HeyMiloAI's candidate search over the weekend:
- 5x faster than gpt-5.6-sol
- 22x faster than claude-fable-5.1
would be interesting to see a couple of things next :
a) do recruiters find the results more accurate
b) how will jev scale for customers with 5M+ candidates
Astra cleared almost every skill on Runebench. the RS bot market in 2008 with this would have been insane.
imagine waking up and your PK bot has looted you 100M from wildy overnight 💀
you can sell that irl for well over the token costs
the true value of an RL environment is its ability to turn expert judgment from messy, unstructured, undocumented exceptions into situations agents can experience, learn from and be measured against.
essentially translating tribal knowledge into something structured and repeatable.
a strong benchmark deliberately introduces those conditions. the env recreates the work, the tasks surface the exceptions, and the verifier turns expert judgment into measurable outcomes, resulting in a score telling us whether the agent can handle the decisions that make the work difficult, rather than whether it can reproduce the normal path.
Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees.
In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn.
Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.
We tested 20 AI coding agents on 50 ATS integration tasks against live APIs.
Fable 5.1 scored 69.3 out of 100.
None cleared 70.
Today we're releasing Integration Bench: https://t.co/SqDnlfwCHY
The struggle to integrate 100s of SOR platforms over the past couple of years with a 4 person team was real.
Today, we’ve got a powerful harness that enables a suite of agents to implement these integrations effectively.
For our customers: integration costs/timelines ⬇️
We tested 20 AI coding agents on 50 ATS integration tasks against live APIs.
Fable 5.1 scored 69.3 out of 100.
None cleared 70.
Today we're releasing Integration Bench: https://t.co/SqDnlfwCHY
i’ve always drawn the line when it comes to codex/claude permissions by scoping to subfolders and access to executables on my mac
with Astra having more capabilities on complex multitasking I’m tempted but wondering what is being sent to servers and what the retention on it is
"Can AI use the tool?" is very different from "Can AI reliably do the work?"
That gap gets much bigger once a task spans multiple systems, requires decisions, and has to be completed end-to-end.
We built a benchmark to measure exactly that.
Launching tomorrow. Feedback + collaborators welcome!
100+ platforms integrated at HeyMilo taught @pasannnnnnnn and I one thing: an integration can pass every call and still write the wrong record.
Integration Bench grades the record, not the calls.
Out tomorrow.
first time setting up a ceiling projector, and I’m quite amazed
was wondering why I feel so much more relaxed so I asked Claude:
- lying flat = sleep posture, not couch posture… completely different nervous system response
- soft reflected light, not a glowing panel aimed at your face
- no bezel, no furniture - it reads as a skylight, not a device
- dim + textured = window vibes
i just found out my bestfriend broke up with her sf boyfriend over the weekend
i asked what happened
she goes "i found out he had an ai agent managing the relationship"
turns out the thoughtful texts, the check-ins, the good morning messages, the conflict resolution, dinner date plans...all automated
my friend was emotionally attached to a cron job