Looped Transformers seem to be the new model architecture the frontier models are using
so go watch hours of videos about it.
I am sure that will help you really understand how it works...
karpathy is goated for suggesting to ask models to write in the ASD-STE 100 style.
it has improved the response of each model and helps so much to understand depth of something in most efficient manner.
6 real-world use cases of Jev:
Many AI systems use generative models for decisions with a small, predefined answer space.
That is exactly where Jev fits. Give it the relevant state, define the questions, and receive typed decisions that your application can use directly.
Here are six practical ways to apply that pattern.
→ Input guardrails
Evaluate whether an incoming prompt is safe before passing it to an LLM. Jev can provide a focused semantic check while deterministic rules continue handling exact restrictions.
→ Model routing
Estimate how difficult a prompt is, then route it to the cheapest model capable of handling it. Straightforward requests can use a fast model while harder ones reach a frontier model.
→ Reranking
Judge which retrieved documents actually answer a query. The resulting relevance values can reorder candidates before the best context is sent downstream.
→ Tool-call gating
Decide whether a proposed action should run, require clarification, or be rejected. This semantic gate belongs after hard permission and policy checks, not in place of them.
→ Confidence gates
Evaluate whether a drafted response can be sent automatically, needs review, or should go directly to a person. The thresholds should be calibrated using representative examples before production use.
→ Output evaluations
Score a generated response against a defined rubric. Several focused criteria can share the same evidence, making it possible to evaluate grounding, relevance, honesty, or helpfulness without generating a written critique for every check.
The common thread is bounded judgment.
When your application needs a category, score, or probability from a known answer space, another round of open-ended text generation may be unnecessary.
Jev does not replace every LLM call. It gives us a focused primitive for the places where software needs a decision rather than another paragraph.
I explored the sixth use case in detail by building a Jev judge for agent evaluation and connecting it to an experiment workflow with Opik.
The full article is quoted below.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
These are the best visual AI resources for learning Transformers, LLMs, embeddings, diffusion, inference and model internals ↓
1/ Transformer Explainer - Watch GPT process text through embeddings, attention, MLPs and next-token prediction.
https://t.co/vkyxOqDVll
2/ Brendan Bycroft’s LLM Visualization - Explore an LLM from architecture down to tensors and operations.
https://t.co/ZFGH0pRtP3
3/ 3Blue1Brown - Visual intuition for linear algebra, neural networks, backprop, attention and Transformers.
https://t.co/plaENBG6aQ
4/ The Illustrated Transformer - One of the clearest visual explanations of embeddings, Q/K/V and attention.
https://t.co/tYrJCm17L6
5/ TensorFlow Playground - Watch neural networks learn as you change layers, activations, features and learning rate.
https://t.co/vva9dm1Gnv
6/ Google PAIR AI Explorables - Interactive explainers on LLMs, generalization, interpretability and model behavior.
https://t.co/vva9dm1Gnv
7/ Distill - Exceptional visual essays on t-SNE, feature visualization, GNNs and interpretability.
https://t.co/FuqpUdDwZD
8/ Abhik Sarkar’s Transformer Visualizations - RoPE, KV cache, FlashAttention, MQA, GQA and more.
https://t.co/vva9dm1Gnv
9/ Visual Guide to Attention Variants - MHA, MQA, GQA and MLA visually compared.
https://t.co/nbpKMa0cpT
10/ Visual Guide to Mixture of Experts - Routing, experts, sparse activation and load balancing.
https://t.co/3gu7xq5e3i
11/ Modular LLM Inference Handbook - Prefill, decode, KV cache, batching, quantization and speculative decoding.
https://t.co/itMIfZnPP5
12/ Apple Embedding Atlas - Explore clusters, neighborhoods and outliers in large embedding spaces.
https://t.co/hxKdXSp5I2
13/ Diffusion Explainer - Follow Stable Diffusion step by step.
https://t.co/YVHEG1zslC
14/ Neuronpedia - Explore features, activations, SAE latents and attribution graphs inside real models.
https://t.co/F8P8eNDH3W
15/ Seeing Theory - Probability, Bayes, distributions, regression and inference made interactive.
https://t.co/KeEhGNSMIQ
16/ CNN Explainer - Visualize convolutions, feature maps, activations and pooling.
https://t.co/sCwv59d2S6
17/ GAN Lab - Train a GAN in your browser and watch its generated distribution evolve.
https://t.co/G94RDWMNak
Save this. There’s a serious AI curriculum hiding inside these links.
These are the best visual AI resources for learning Transformers, LLMs, embeddings, diffusion, inference and model internals ↓
1/ Transformer Explainer - Watch GPT process text through embeddings, attention, MLPs and next-token prediction.
https://t.co/vkyxOqDVll
2/ Brendan Bycroft’s LLM Visualization - Explore an LLM from architecture down to tensors and operations.
https://t.co/ZFGH0pRtP3
3/ 3Blue1Brown - Visual intuition for linear algebra, neural networks, backprop, attention and Transformers.
https://t.co/plaENBG6aQ
4/ The Illustrated Transformer - One of the clearest visual explanations of embeddings, Q/K/V and attention.
https://t.co/tYrJCm17L6
5/ TensorFlow Playground - Watch neural networks learn as you change layers, activations, features and learning rate.
https://t.co/vva9dm1Gnv
6/ Google PAIR AI Explorables - Interactive explainers on LLMs, generalization, interpretability and model behavior.
https://t.co/vva9dm1Gnv
7/ Distill - Exceptional visual essays on t-SNE, feature visualization, GNNs and interpretability.
https://t.co/FuqpUdDwZD
8/ Abhik Sarkar’s Transformer Visualizations - RoPE, KV cache, FlashAttention, MQA, GQA and more.
https://t.co/vva9dm1Gnv
9/ Visual Guide to Attention Variants - MHA, MQA, GQA and MLA visually compared.
https://t.co/nbpKMa0cpT
10/ Visual Guide to Mixture of Experts - Routing, experts, sparse activation and load balancing.
https://t.co/3gu7xq5e3i
11/ Modular LLM Inference Handbook - Prefill, decode, KV cache, batching, quantization and speculative decoding.
https://t.co/itMIfZnPP5
12/ Apple Embedding Atlas - Explore clusters, neighborhoods and outliers in large embedding spaces.
https://t.co/hxKdXSp5I2
13/ Diffusion Explainer - Follow Stable Diffusion step by step.
https://t.co/YVHEG1zslC
14/ Neuronpedia - Explore features, activations, SAE latents and attribution graphs inside real models.
https://t.co/F8P8eNDH3W
15/ Seeing Theory - Probability, Bayes, distributions, regression and inference made interactive.
https://t.co/KeEhGNSMIQ
16/ CNN Explainer - Visualize convolutions, feature maps, activations and pooling.
https://t.co/sCwv59d2S6
17/ GAN Lab - Train a GAN in your browser and watch its generated distribution evolve.
https://t.co/G94RDWMNak
Save this. There’s a serious AI curriculum hiding inside these links.
This is absolutely insane.
A Stanford AI research group joints JEV with Claude Code to sort 100+ billion data points every 15 minutes.
the trick is stupidly simple:
JEV runs a cheap first pass on everything. Claude only gets the hard cases.
so instead of: everything -> Claude -> $$$
it’s: everything -> JEV filters -> hard stuff -> Claude thinks
the boring stuff never touches the expensive model.
Claude gets a tiny pile that actually deserves deeper analysis.
faster. cheaper. way less compute burned.
basically, JEV sorts the mail so the genius only opens what matters.
LLMs think, JEV decides
Read the full Jev setup guide in the article below: