@chris_girard13 Live your life. She clearly didn’t seem to hold your opinion that she was out of your league. Now, you’re both in a league of your own. Congrats!
I hope the timing of this tweet makes it clear. I really don’t want to be the face. I truly, to my bones, don’t. Unfortunately that’s just how things are lining up. I will give 3 total interviews ever. Period. I owe nothing to anybody. Be good. Do good. That’s all.
Just for the record, when my model is announced in the next few weeks, please feel free to test and benchmark it, just don’t crash my servers. Real AGI doesn’t care. Those terms are fishy.
ARC-AGI-2 self tested after my recent vision and spatial reasoning additions came back at 40% of tasks got exact matches (correct) and it refused to make guesses at the rest. 23ms per task. Mac Mini M4. I’m not done teaching world knowledge, will submit for a prize soon after
@dotta Do you all use something like what pear(sp?) was (micro influencer marketing) or? The same stuff regurgitated by so many accounts it feels very much like that. Curious for my upcoming AGI announcement, clearly it works to get the word out.
This is not useful. It has limitations that all ML does and is not intelligence. It’s a neural rules engine which, IMO is a step backward. My guess is financials necessitated “we have to show something so do one of those fake documentary style videos and say something smart”
this is the easiest way to understand Jev:
LLMs generate answers.
Jev makes decisions.
that sounds like a small difference, but it actually changes the entire use case.
say you give a normal LLM this:
“here’s a user, their account history, payment behavior, support chats, device data, etc.
tell me if this looks risky.”
the LLM might reason through it and return:
“yes, this looks high risk.”
maybe in JSON if you ask nicely.
with Jev, you define the possible decisions upfront:
risk:
* low
* medium
* high
manual review:
* yes
* no
and Jev returns something closer to:
risk = high (96%)
manual review = yes (91%)
that’s basically the product.
it’s not trying to be another ChatGPT.
it’s more like an AI-native if statement.
instead of:
if transaction > $10,000:
review()
you can start thinking more like:
if “does this behavior look suspicious?” > 95%:
review()
and that opens up a pretty interesting category of software.
a few assumptions I had at first that turned out to be wrong:
1. “so it’s just a classifier?”
kind of, but that undersells it.
the input can be messy real-world context, and you can ask multiple typed questions about that state at once.
fraud?
churn?
escalate?
eligible?
priority?
all from the same input.
2. “so it replaces GPT / Claude?”
not really.
I actually think the interesting architecture is:
Jev decides WHAT needs to happen
Claude / GPT reason or generate WHEN deeper intelligence is needed
normal code executes the deterministic stuff.
Jev becomes the routing layer.
3. “it can’t hallucinate?”
this one needs nuance.
if your allowed answers are:
LOW
MEDIUM
HIGH
Jev won’t suddenly invent:
“EXTREMELY HIGH 🚨”
the output structure is constrained.
but it can still be wrong.
HIGH at 92% can still be the wrong decision.
so “no hallucinations” doesn’t mean “always correct.”
4. “why not just force an LLM to return JSON?”
you can.
we already do this everywhere.
but you still deal with generation latency, schema validation, retries, weird outputs, confidence estimation and a lot of glue code.
Jev is designed around the decision itself rather than text generation.
5. “why should I care?”
because most software is ultimately a giant tree of:
if this → do that
if this → route here
if this → escalate
if this → reject
if this → ask a human
Jev is basically asking:
what if those if statements could understand messy human context?
that’s a much more interesting framing than “another AI model.”
I can see this being very useful for:
fraud / risk
support routing
moderation
PR / QA automation
lead scoring
compliance
workflow orchestration
agent routing
especially as the cheap + fast decision layer sitting in front of larger reasoning models.
early tech, obviously.
but the category itself makes a lot of sense.
Kinda like how threads was? Pumping numbers up numbers reasonably so by one tap to the download. Retention was the issue then, and what matters still. With how much they’ve spent on training I would assume it’s a decent app. Lots of new ways to get data too
Did I hear this right…. Portland Oregon Police chief just talked about “protecting … lesser people in our community” in his interview. @PortlandPolice any idea the particular group your boss was referring to?
If you aren’t constantly doing this, you’re not the engineer, you’re the tool and AI is the engineer. If you do, you’re the engineer and AI is the tool.
Has anyone ever actually challenged a decision made by Claude Code or Codex while coding?
Like, the AI suggests a way to do something and you just go:
“No, I think there’s a better way.”
Curious how often people actually push back on their AI coding tools.
@BernieSanders I am an expert. This is not a valid concern. AI fakes intelligence and intent, at least LLMs with tools bolted on. I’m releasing an AI soon that is actually AGI. I already have it built and learning and coding etc. even then, I am able to limit what it can do cleanly
@rileybrown I think this type of thing will happen across saas, on top of the 100s of new tools companies are building themselves with AI. That’s the whole concept behind https://t.co/GopZj9x5NH - I’m even building a slack like chat to build from: https://t.co/btLfx3kPup
@ben_sage Find a design system you like, or two, Fred that into whatever ai agent you use and just be sure to be specific when promting to then create a new design system “pay attention to borders. Don’t use pills. …” that alone gets you a long ways. Did that with https://t.co/U2NCz3erYa
My plan for Quickish: AWS & GitHub for AI Driven Dev Projects. The modern version of an intranet for orgs and geocities for the rest of us with all the services to power at scale built in.
@mil000 Built a different version of Claude / ChatGPT / Grok . Shows you how it’s not founders or ideas that matter when raising, it has been and always will be who you know.
@ClaudeDevs Or just tell your ai agent to use https://t.co/btsEv1BGAN dev tunnel, that’s all you gotta say and you’ve got a live reload web accessible private and shareable tunnel for free using any agent