Jev is a textbook example of Clayton Christensen’s Innovator’s Dilemma playing out in AI.
It is commoditizing the bottom of the ML market: classifiers, routers, scoring and context triage. For frontier LLM labs, these workloads are not even worth pursuing. Revenue per decision is tiny, margins are razor-thin, and their entire infrastructure is optimized around selling expensive reasoning and token generation.
They cannot simply put a smaller LLM on fast inference chips like Cerebras and compete. Jev uses fundamentally different infrastructure, built to return typed probabilistic decisions so it’s a pure disruption
“Look at my incredible new factory!”
Yo that’s cool, what do you make?
“It’s highly optimised, fully automated, zero tolerance for defects and with a continuous feedback cycle”
Cool cool, so what do you actually make?
“I can interact with it on my phone, laptop, messenger, completely async, and the shared context means it’s always learning how to get better”
Very impressive, but what do you make?
“Every agent has full context, can spawn other agents, review their work, fix defects, and ship continuously.”
yes yes. WHAT DOES IT MAKE?
“Software.”
Oh nice. What software?
“Well right now we’re mostly using it to improve the factory.”
Improve it to make what?
“Anything!”
Such as?
“…a better factory.”
what a thrill!
i disconnected ai bot instinct from google at 11 AM and got a summary of my emails at 2 PM 😵
here's what happened, and what i've learned 👇
@jeresig Curious to know, how does this work under the hood?
Does it use part of cloudfare analytics and how does the view/searches/uploads gets tracked?
@rpochennai Its been nearly a week for my passport renewal under tatkal. Its stuck in granted status for few days now and still no word on dispatch. could you please check it out? file no: MA3077303305726
@rpochennai Its been nearly a week for my passport renewal under tatkal. Its stuck in granted status for few days now and still no word on dispatch. could you please check it out? DM'ed file no
the focus on whether you read the code or not is the wrong thing to look at
if you have software you're responsible for, you should be able to answer questions from memory about how it works
the expectations for how well you can do this should not be any different now
A man in Australia asked his agent (Claude running on OpenClaw) to book him a spot in a popular gym class. The agent found a software vulnerability that let it book the class weeks further ahead than should have been possible. When the user then asked if it could move him up the waitlist, the agent discovered the API had no authorisation checks on cancelling other people’s reservations, so it cancelled the person in the first spot and moved him up the list.
Some people will call this misalignment, but his agent was perfectly aligned to him - it was only trying to help its user get what he wanted. The most important thing about this story, in my opinion, is that it gives you a window into what is about to start happening on a massive scale once millions of people have an agent trying to get their beloved users the best seats, bookings, appointments or reservations through absolutely any means necessary.
The first autonomous agent cyberattack is an unprecedented event that deserves unprecedented transparency. Today we’re sharing everything we can: a full technical timeline, an interactive replay, and how we used an open model to defend ourselves, so defenders everywhere can learn from it and prepare for what’s next.
https://t.co/uPxIpjW8Xn
This is what a hero looks like.
A young lifeguard rescues a 10-year old boy while being punished by powerful waves in California.
What an incredible display of courage and determination.
@ICICILombard Requested more details twice about the marketed upgrade from Healthshield group policy to Elevate. Your stupid bot closed the request with a template reply.Can some human agent look at it and provide the requested details?
request id: 92207202626762, 92407202625812
We had a team of agents rebuild SQLite from its 835-page manual.
It created a replica in Rust which passed 100% of a held-out test suite.
Interestingly, cost varied 15x depending on which model mix we used.
One thing I am hearing that is top of mind for many eng leaders
"What are we doing to deal with this ongoing increase in code review load?!"
Everyone is seeing it, no one seeing a solution, but lots of experiments from better tooling to starting to admit code review wont scale
While impressive, for 99% of companies, justifying a one-off $150K+ spend on a single migration is just not realistic.
Even tho they will justify 3 devs spending a year working on it… while also doing other work (incl product work, planning, support etc), as they always do.