Jev has been exploding in popularity recently.
If you already have access to the Jev API but aren’t sure how to start experimenting with it, just copy this checklist:
1. agent-desktop
Desktop automation. Read the system's accessibility tree, judge which button, menu, or input field to click next. https://t.co/ZtSEjUSPBF
2. typesafe-mario
Have Jev play Super Mario. No screenshots—just read the structured state in the emulator's RAM, then decide to run, jump, or dodge. https://t.co/GHttIjWQ3p
3. jev-drone
Use Jev to control a drone. The underlying flight control still handles stability and safety; Jev just does higher-level judgments like climbing, braking, and navigating obstacles. https://t.co/z0lYh9ykJq
4. OneVOneJev
1v1 FPS in the browser. Every decision tick, judge movement, view angle, aiming, firing, and jumping. https://t.co/aJiU0aaNcI
5. jev-trader
High-frequency market making on Monad testnet. Jev judges the next buy or sell based on spreads and trade direction, with model latency around 81ms. https://t.co/DaDRIrkJpO
6. Prism
Doesn't directly have Jev place orders. It judges states like toxic flow, market pressure, mean reversion, etc., then hands off to the original strategy. https://t.co/aim9lGRAP8
7. neo4jev
Stuff Jev into a knowledge graph. At each node, judge the most worthwhile edge to take next, then follow it all the way. https://t.co/9h0KXKdWj9
8. jev-curate
Use Jev to screen training data. For JSONL / Parquet, first judge quality, relevance, and risk, then decide which ones go into the next training round. https://t.co/yYV6aEdUtG
9. Canny
Prevents Coding Agents from stubbornly claiming they're done. Look at tool outputs, code diffs, and test results, then judge if the completion claim is reliable. https://t.co/H4jFT8hV0E
10. killmyidea
Input a startup idea, and Jev scores it from multiple angles, finally giving you KILL, FIX, or SHIP. https://t.co/XWo7JPOb6y
Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓
Holy sh*t, I wish I’d found this earlier.
someone put 300+ AI agent projects into one free GitHub repo
because half the time you think you need to build something new…
someone already built it.
agent memory? there.
browser control? there.
coding agents, sandboxes, evals, voice, multi-agent teams, local models?
all there.
so I went through the list and started connecting the pieces:
> RUN - Ollama runs the model locally
> BUILD - LangChain connects everything
> CODE - Aider + Continue work inside your codebase
> ORCHESTRATE - CrewAI + AutoGen manage the agent team
> ACT - E2B + Composio give agents tools and somewhere safe to run
> REMEMBER - Mem0 keeps memory between sessions
> WATCH - AgentOps shows what your agents actually did
> SHIP - Vercel AI SDK helps turn it into a real product
but here’s what clicked for me:
the list isn’t really a list. it’s a build-your-own-agent menu.
model -> framework -> tools -> sandbox -> memory -> evals -> product
missing something?
don’t spend the weekend rebuilding it.
find the category -> see what already exists -> grab 3 options -> check which one’s still alive -> start there.
and if you want to build on top of the list itself, they even give you the ecosystem data in JSON + YAML.
> no signup.
> no paywall.
> MIT licensed.
> 300+ projects in one place.
basically, before I build anything agent-related now, I’d check this repo first.
I pulled out the projects I’d actually start with + the best stuff from each category below
this Jev stack is f*cking crazy
5 open-source repos that let your agent make decisions for cents instead of burning frontier tokens.
each one handles a different part of the job:
ROUTE
> jev-for-all - Jev picks which skill to load and which tools your coding agent gets. 64 requests, 0 wrong picks, ~$0.0001 per decision
https://t.co/XreSBsVgTy
GUARD
> agent-jev - a 0.6B model that reads diffs, logs and traces and answers typed questions in ~60 ms. zero generated tokens. ships with a Claude Code hook that gates every tool call
https://t.co/wdcviYegzg
RUN LOCAL
> jev-rs - a Jev-compatible engine in Rust. plug in any local model, state gets prefilled once, every question is answered from probabilities. 75 ms per question on a Mac
https://t.co/Sjhaai0xcE
TRAIN
> typed-decisions - build your own Jev-style model. DeBERTa version hits 85.5% accuracy at 42 ms, and training cost $0.26
https://t.co/0KyH0OmIBE
SHIP
> jev-usecases - 36 production harnesses: support triage, fraud, invoices, guardrails, SOC agents. every decision lands in auto, confirm, human or block by confidence
https://t.co/C2yl2gDEeC
the cool part isn't any single repo.
it's that your agent can now:
pick the right skill -> gate its own tool calls -> run offline -> learn your decisions -> ship with confidence bands
your LLM doesn't need to think about every yes or no anymore.
the decision layer is already on GitHub. you just plug it in.
bookmark this before you build your next agent.
Send this Graph Engineering prompt to your coding agent before you build another multi-step workflow.
It stops your agent from grading its own homework, and catches the bugs walking back in twice.
Steal it before your bill finds out how many times it paid for the same mistake.
i genuinely don't understand why everyone isn't doing this yet
Grok Bot watches your X timeline, figures out why other posts blow up and writes my next one before i even know what to post.
i used to burn ~3 hours a day on this:
scroll -> save posts -> check views -> dump ideas into notes -> stare at a blank draft -> repeat
so i built a tiny X newsroom instead:
-> Scout tracks ~18 accounts, skips replies, drops anything under 5K views, stars 10K+
-> then scans ~20 fresh AI posts from the last 48h to catch what the whole feed suddenly talks about
-> it breaks down what worked: hook, keywords, structure, topic, tool, CTA
-> Draft turns those patterns into a new post + a short video idea
-> Board tracks everything: found, drafted, posted, still to do
-> i fix whatever sounds dumb and hit Post myself. nothing auto-posts
it doesn't invent trends. it reads what people already proved with their attention.
one Claude bug-hunting post from this workflow did ~2.1M views.
the math:
~3 hours/day back
~$1K/month not spent on assistants
1–2 posts/day without living in the feed
a little AI employee with one job:
find what works -> figure out why -> hand me something worth posting
it's public and free
Peter Steinberger, the creator of OpenClaw, open-sourced his entire agent setup
The core idea is one folder that every agent on his machines reads from.
A single AGENTS.MD holds his rules, and a script symlinks it into ~/.claude/CLAUDE.md and ~/.codex/AGENTS.md
This way Claude Code and Codex always follow the same instructions.
Every other repo gets one line at the top:
READ ~/Projects/agent-scripts/AGENTS.MD BEFORE ANYTHING.
Change a rule once and every project picks it up.
Skills work the same way. 69 of them live in one place, each with a short description the agent reads to decide what to load, and scripts/sync-skills links them into both agents.
Fork it, replace his rules with yours, and delete the skills you do not need. His AGENTS.MD is full of his own hosts and accounts.
6.6k stars, MIT - https://t.co/3XQj9GxI6b
WHOEVER DISCLOSED THIS IS ABSOLUTELY FEARLESS
A SINGLE TRADER IS RUNNING A FULL DECK OF 400 GROK AGENTS
I assumed this was just standard AI hype until I looked under the hood
First agent monitors order books and volume
Next one screens for high-conviction entries
Another parses breaking news and market mood
Another tracks smart money wallet flows and chain metrics
Another calculates position sizes and tail risk
A primary LEAD GROK orchestrates the whole flow, only pinging the operator when manual execution is required
Zero staff needed. Zero time spent staring at candles
Just an army of 400 Grok instances constantly passing context to each other non-stop
The wild part?
The creator claims the infrastructure runs almost entirely on autopilot now
Detailed instructions are provided in the article
This is absolutely insane.
A Stanford AI research group joints JEV with Claude Code to sort 100+ billion data points every 15 minutes.
the trick is stupidly simple:
JEV runs a cheap first pass on everything. Claude only gets the hard cases.
so instead of: everything -> Claude -> $$$
it’s: everything -> JEV filters -> hard stuff -> Claude thinks
the boring stuff never touches the expensive model.
Claude gets a tiny pile that actually deserves deeper analysis.
faster. cheaper. way less compute burned.
basically, JEV sorts the mail so the genius only opens what matters.
LLMs think, JEV decides
Read the full Jev setup guide in the article below:
This is absolutely f*cking insane.
A Stanford AI research group joints JEV with Claude Code to sort 100+ billion data points every 15 minutes.
the trick is stupidly simple:
JEV runs a cheap first pass on everything. Claude only gets the hard cases.
so instead of:
everything -> Claude -> $$$
it’s:
everything -> JEV filters -> hard stuff -> Claude thinks
the boring stuff never touches the expensive model.
Claude gets a tiny pile that actually deserves deeper analysis.
faster. cheaper. way less compute burned.
basically, JEV sorts the mail so the genius only opens what matters.
LLMs think. JEV decides. code does.
SpaceX AI hired this engineer at $725K-$1.1M a year to build teams of GrokBot agents that run a whole company
This 1-hour workshop shows exactly how to build one from scratch
step 1 → split the work by expertise - CloseBot, ProdBot, StalkBot, ProtoBot, one job each
step 2 → give every bot its own computer - it signs up, clicks and tests like a human, no API needed
step 3 → CloseBot preps every call and rereads every transcript - 99% of the customer journey runs without you
step 4 → ProtoBot turns X, support and email feedback into PRs - feedback to deploy in hours, not weeks
step 5 → one bot optimizes all the others - every week the whole team wakes up smarter
SpaceX AI calls this "automating a staff function"
Most people spend months hiring this team
You don't have to
Bookmark and watch it
Then read how to wire a team of agents into one graph below ↓
this is f*cking gold
How to build your first AI agent (Full guide)
this is basically everything you need to know about building with Jev
in the right hands, this changes everything:
1. We will keep accelerating. Our AI efforts are only 3 years old, vs 6 and 10 years old for Anthropic and OpenAI. If our second derivative remains strong, SpaceX will reach pole position in about 6 months.
2. Once you far exceed the caliber of intelligence needed for a class of tasks, additional intelligence is pointless. You don’t need (and it would be cruel to put) Newton-level intelligence in your toaster!
3. Hardware is hard. Bringing massive compute online rapidly is incredibly difficult. SpaceX has demonstrated exceptional ability in this regard and will only get better.
SpaceXAI engineer just released a free 1-hour course on mastering GrokBot
From one prompt to a fully autonomous team of agents:
0% → 00:00 - Introduction to GrokBot agents
25% → 12:31 - Build your first GrokBot team
50% → 29:18 - The 4 layers of the GrokBot stack
75% → 44:41 - Create and use GrokBot templates
100% → 52:08 - Run GrokBot in phone mode
GrokBot → Roles → Agent Stack → Templates → Phone Mode → Autonomous Teams
This 1-hour watch can replace 15 paid agent engineering courses
Skip Netflix tonight
Watch this course and ship a fully autonomous team of GrokBot workers by Monday
NVIDIA JUST OPEN-SOURCED A FREE AI ROUTER THAT TURNS EVERY MACHINE ON YOUR HOME NETWORK INTO ONE LOCAL INFERENCE CLUSTER.
ZERO CLOUD. ZERO API COSTS. ALL YOUR IDLE COMPUTE WORKS TOGETHER.
SpaceX built rockets this way, now the same playbook is building drone boats
Saronic’s head of manufacturing came from @SpaceX, and he had the same rule: Build the system that builds the system, then hunt the gap.
“When you have this linear flow, you can see there's actually no boat in that workstation.. basically, my kids could walk through this production line & identify where something isn't flowing appropriately & start asking questions on how to get it back to rate.”
Rockets first, autonomous boats next.
Writer: Val
Built a Self Executing Agent Loop that can run 300 agents through 4,000 steps without needing a human to restart the process.
- 300 agents.
- 4,000 steps per run.
- 5 live data feeds.
- 3 verification passes.
The important part isn't the agent count.
It's that the system can execute, verify its own output, update context, and keep running without waiting for the next prompt.
GPT-6 Astra + @threejs is insane. We are way beyond static models. We are engineering living worlds in code now. There are no "3D model files" in this demo. Both trains are generated at runtime from Typescript/Three.js code using dimensions, profiles, and geometry functions. Wheel motion and explode/reassemble animations are also entirely code-driven. Runs super smooth inside the browser.
To bring large AI data centers online requires building massive power plants, transformers, liquid cooling loops & chillers, as well as incredibly complex networking (especially for training clusters)