we rolled out jev for @HeyMiloAI's candidate search over the weekend:
- 5x faster than gpt-5.6-sol
- 22x faster than claude-fable-5.1
would be interesting to see a couple of things next :
a) do recruiters find the results more accurate
b) how will jev scale for customers with 5M+ candidates
I use IG Reels to this day to learn about new concepts/ideas specifically about AI (unsurprisingly). Realized how powerful it was in helping me learn when I was working on the team a few years ago.
That said, there's times in the day when I go in and my brain is focused on learning (i.e during the day) and then there's times when I'm looking to relax (i.e at night).
Unfortunately the algo doesn't pick that up still. If it did, I would likely spend significantly more time on the platform at night instead of closing the app after a day of working/hearing/using AI :)
We process a new candidate nearly every minute of the day at this point through @HeyMiloAI
One of the most interesting things to see is the types of roles candidates are being screened for.
Might be interesting to publish how these roles have changed over the past few years
This is ridiculous scheduling.
I get that today's matches have basically all gone long, but you can't have the biggest match of the tournament so far beginning at 11 PM ET at the earliest.
The US Open has to do better.
Shelton and Alcaraz will basically have no east coast audience.
@kevinwng Congrats! Curious how youโre measuring the gains in design aesthetics and the more subjective parts of knowledge work. What do the tasks and judging criteria look like?
@ArtificialAnlys just added GDP.pdf to their Intelligence Index.
Which means their definition of intelligence now includes something deceptively simple: can models understand the documents you deal with on a normal Tuesday at work?
Leases, invoices, dosage tables, and financial reports.
Astra, the best model, still solves just under 1 in 3.
Master the boring, master the frontier.
Scaling @HeyMiloAI into some of the largest staffing agencies and enterprises in the world relies heavily on the depth and robustness of our integrations with their applicant tracking systems (ATS).
Despite how good coding agents have gotten, weโve consistently found that getting these integrations production-ready still requires a human engineer in the loop.
That gap motivated Integration Bench.
We turned 3+ years of our engineering teamโs real integration work into tasks designed to test whether coding agents can handle the work required to ship a successful production integration.
Would love feedback from folks working on coding agents, evals, and post-training on the benchmark.
Amazing work @0xpasan, @yapramie, @andrunder
We tested 20 AI coding agents on 50 ATS integration tasks against live APIs.
Fable 5.1 scored 69.3 out of 100.
None cleared 70.
Today we're releasing Integration Bench: https://t.co/SqDnlfwCHY
@Swarooprm7 Not sure majority of current benchmarks best represent actually important economic work. Need benchmarks that better represent economically important work and tasks people do in the workforce and I feel thatโll be better measure of intelligence
"The fastest way to trigger neuroplasticity isn't playing brain games or reading research papers.
It's forcing your brain to fail repeatedly at something you have zero natural talent for. high frustration is the literal chemical signal that forces synaptic rewiring."