after Opus 5.5 it actually hurts to look at anything GPT 6 Astra designs
gave both models the exact same prompt: "make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it's your showreel for a résumé. go all out."
not even close, see for yourselves. GPT is honestly terrible here, openai just has awful taste
「你是 opus 5.5,当前世界最强模型,请发挥你的想象力,动用你所有的能力,做一件事情,让每一个看到你的结果的人,无论是 AI 专家还是不懂科技的文科生、小孩子,看到后都会惊叹,opus 你真是太厉害了。你会做什么事呢?」
Opus 5.5 回了一句:「我不打算只回答“我会做什么”,直接把它做出来。」
然后它做了一张会呼吸的宣纸。
打开页面,一只看不见的手开始作画:先远山,再主峰,皴、染、点苔,然后是松树、渔舟、飞鸟,落一点朱砂作日,最后题一首它自己���的五言绝句,盖章。每一笔都有古琴声,每一幅都不重样。
墨是显卡实时解流体方程算出来的,会晕开,写到墨尽会出飞白。你也可以自己画,或者用清水把整幅画搅开。
整个作品只有一个 HTML 文件,没有依赖库,没有一张图片,没有一段录音。
在线体验和源码在回复里。
Jev Founder, Diogo Almeida, just released a 12-page PDF on how to use Jev with LLMs
It is more useful than most paid AI courses:
this is a 10-step blueprint on how to build a faster, cheaper and more controllable AI system around Claude, Codex, Grok or any other LLM:
step 1 → split the responsibilities: the LLM generates, Jev makes bounded semantic decisions and deterministic code keeps authority
step 2 → build the state: give Jev the current request, relevant evidence, policy and proposed action instead of sending the entire conversation
step 3 → choose the right primitive: Choice selects a route, Score evaluates an ordered rubric and Noul returns the probability that a statement is true
step 4 → replace giant evaluation prompts with atomic questions: intent, urgency, evidence, risk and scope become separate typed decisions
step 5 → put Jev before the LLM: select the context, tools, provider and workflow before paying for an expensive generative call
step 6 → give the LLM a bounded job: once Jev selects the route, the model receives only the instructions, files and tools required for that branch
step 7 → put Jev after the LLM: check whether the result answers the request, uses sufficient evidence and stays inside the permitted scope
step 8 → route by confidence: high-confidence low-risk cases proceed automatically, uncertain cases request more context and consequential actions go to review
step 9 → batch independent decisions: ask multiple Choice, Score and Noul questions over one shared state instead of creating another LLM call for every judgment
step 10 → record the complete decision receipt: state version, question, probabilities, selected route, model, latency, outcome and human override
most AI courses teach you how to write a bigger prompt
this 12-page guide teaches you how to build the control system around every prompt
the result: smaller contexts, fewer unnecessary LLM calls, safer tool execution and decisions you can actually inspect, test and improve
Send this PDF and the original Jev article to Claude Code or Codex and start rebuilding one expensive LLM decision at a time ↓
Follow the link below to claim the credit or run /claim-credit in the CLI. You’ll need GitHub connected to start a session. Claim by Oct 7. Terms apply.
https://t.co/TzoO6fzJHB
At the top of my AGENTS.md:
- NEVER write unit tests after you write code.
- Highly prefer E2E tests as the sole testing mechanism. Use them to verify complex features work. At the end of E2E tests, produce a verifiable and repeatable artifact.
- If you must test a system in isolation, FIRST write all the ways it could fail, THEN write the code.
Jev is "Internet" moment for AI: up to 193x faster and 444x cheaper in tests with Claude Fable 5.1 and GPT-6 Astra
Whar is Jev, how to use and unlock its real 100х advantage in my 10-page research:
step 1 → meet Jev: LLM writes, agents act, Jev chooses the next move - split intelligence from execution
step 2 → turn every agent fork into three primitives: Choice selects one route, Score measures a defined scale, Noul returns the probability of yes
step 3 → build before getting access: use TypeSafe’s official adapter with OpenAI, Anthropic or xAI, then swap in Jev without rebuilding the graph
step 4 → setup first Jev: one state, three parallel decisions, risk-based thresholds and a real queue your agents can execute
step 5 → batch decisions instead of serializing them: 13 questions in one call ran 10x faster and 12.2x cheaper than 13 sequential calls
step 6 → place Jev at every bounded fork: choose the agent, model, tool, browser action or human escalation, then read fresh state
step 7 → benchmark the entire loop: Browser Use hit Google Flights in 7.1s, Every ran 777 checks in under 0.7s, Mobile Jev completed 9 actions in 21s
step 8 → rank wide, read narrow: Jev cut wrong Hermes skill loads from 16.8% to 7.3% and pushed legal Top-10 retrieval from 38% to 62%
step 9 → steal a system, not a prompt: Chief of Staff, model router, inbox firewall, research feed, browser controller and safety gate all use State → Questions → Action → Verify
step 10 → keep Jev out of math, writing and irreversible execution: code computes, LLMs create, Jev decides, fresh state proves the result
the result: one slow, expensive agent becomes an always-on decision machine that routes, scores and escalates in milliseconds
Copy the complete 10-page Jev blueprint - then read full 10-step roadmap below ↓