the interesting part of Jev is the framing: RLHF rewards confidence, so models collapse onto one mode for assistance whereas automation needs honest uncertainty.
if that's the core argument, would be interesting to see calibration metrics, not just latency
1/ Jev, a decision model by @typesafeai, sparked a burst of projects and discussion. We tested it using Ori Eval against popular LLMs on OpenRouter at judging.
Jev was >5x faster than the next fastest model, and even its slowest requests beat every other model's median.
@siddsax@OpenAI@AnthropicAI maybe pre-training scaling hitting a hall, but i’d rather wait for test-time scaling on this. the preparedness evals here are also lower bound for potential capabilities. plus an interesting direction of improved eq
@prdeepakbabu@abacaj even if we add a quality metric for answering well, the model might still exploit proxy signals unless we have a joint reward that truly captures both instruction compliance and semantic fidelity
Multi‑Head Latent Attention(MLA) is such a clever strategy in DeepSeek‑V3. By compressing the KV cache, it slashes memory usage dramatically. With 64 attention heads(each of 128 dims), a compression dimension of 2048, and a positional dimension of 2048, MLA cuts memory by ~75%!
Have been using this LLM Consortium for sometime, nice to see @llama_index's implementation with orchestrating asynchronous agents
https://t.co/CBbs7xdeqZ
LLM1 writes code → LLM2 critiques → Feed critique back for self-improving iterations. Each iteration improves itself based on previous feedback. Anyone built a browser extension for this loop?
OpenAI Strawberry (o1) is out! We are finally seeing the paradigm of inference-time scaling popularized and deployed in production. As Sutton said in the Bitter Lesson, there're only 2 techniques that scale indefinitely with compute: learning & search. It's time to shift focus to the latter.
1. You don't need a huge model to perform reasoning. Lots of parameters are dedicated to memorizing facts, in order to perform well in benchmarks like trivia QA. It is possible to factor out reasoning from knowledge, i.e. a small "reasoning core" that knows how to call tools like browser and code verifier. Pre-training compute may be decreased.
2. A huge amount of compute is shifted to serving inference instead of pre/post-training. LLMs are text-based simulators. By rolling out many possible strategies and scenarios in the simulator, the model will eventually converge to good solutions. The process is a well-studied problem like AlphaGo's monte carlo tree search (MCTS).
3. OpenAI must have figured out the inference scaling law a long time ago, which academia is just recently discovering. Two papers came out on Arxiv a week apart last month:
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. Brown et al. finds that DeepSeek-Coder increases from 15.9% with one sample to 56% with 250 samples on SWE-Bench, beating Sonnet-3.5.
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. Snell et al. finds that PaLM 2-S beats a 14x larger model on MATH with test-time search.
4. Productionizing o1 is much harder than nailing the academic benchmarks. For reasoning problems in the wild, how to decide when to stop searching? What's the reward function? Success criterion? When to call tools like code interpreter in the loop? How to factor in the compute cost of those CPU processes? Their research post didn't share much.
5. Strawberry easily becomes a data flywheel. If the answer is correct, the entire search trace becomes a mini dataset of training examples, which contain both positive and negative rewards.
This in turn improves the reasoning core for future versions of GPT, similar to how AlphaGo’s value network — used to evaluate quality of each board position — improves as MCTS generates more and more refined training data.
Tried comparing 4o vs o1 on reasoning tasks involving compound rules: this UKCAT-style problem to differentiate two sets. Both models are failing, o1's explicit reasoning doesn’t seem to provide any advantage here.
Haven't come across a single useful scenario where o1-models perform better. They take much longer with repetitive output that doesn’t feel like actual 'thought.' In contrast, implicit Chain-of-Thought reasoning with 4-o seems more precise and definitely faster.
Haven't come across a single useful scenario where o1-models perform better. They take much longer with repetitive output that doesn’t feel like actual 'thought.' In contrast, implicit Chain-of-Thought reasoning with 4-o seems more precise and definitely faster.
LLM1 writes code → LLM2 critiques → Feed critique back for self-improving iterations. Each iteration improves itself based on previous feedback. Anyone built a browser extension for this loop?