@trq212 Have you had the experience where the model finishes its goal perfectly fine, all tests pass, but then you notice it had introduced either regressions or tried to be smart and build something never intended? I’d argue a plan at least still enforces a build phase!
Reading Foucault’s Discipline & Punish this week and wow did he get ahead of our times. A 1975 piece laying out the foundation of a teacher-student model akin to RL post training today
@TracerootAI@xinwei_97 1. In workspace settings -> model provider -> add OpenRouter as adapter
2. Paste your key, label "stealth/ox-alpha"
3. Try the model as LLM-as-judge
You can also completly self-host @TracerootAI on your own stack!
New stealth frontier model just dropped and nobody knows its true identity yet. 1M context window, Zero data retention, Multi-modal, Free to use, and essentially unlimited usage.
Ox Alpha is now natively supported in @TracerootAI 's BYOK through OpenRouter. Try LLM-as-judge now!
🥷 New stealth model: Ox Alpha
Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use.
- 1M token context window
- Text, image, and video input
Try it now and share feedback to improve the model! https://t.co/tU5lmZrO5z
Having confirmed the pipeline works, I then asked:
what does the geometry of "death" look like? a rather conceptual term to humans.
Full write-up: https://t.co/epjMBzV6Iu
I reproduced Geiger et al on Llama-3.1-8B:
Temperature @ L19: R²=0.993, Spearman=+0.995
Age @ L19: R²=0.978, Spearman=+0.989
The manifolds are beautiful geometric illustration.s Here's my side-by-side 🧵
Neural networks might speak English, but they think in shapes.
Understanding their rich *neural geometry* is key to understanding how they work – and to debugging and controlling them with precision.
Starting today, we’re releasing a series of posts on this research agenda. 🧵
One new detail I found in reproducing the paper:
"Geography" peaks at Llama's final layer — not L19 like the other concepts
A small finding, but it suggests spatial coordinates are resolved later in the forward pass than scalar quantities like temperature or age
It’s understandable Fable 5 is refusing lots of my inputs but why is Codex and 5.5 refusing such harmless prompt, meant to test my endpoint? @OpenAIDevs
@claudeai@AnthropicAI In my experience in the TUI, I can visibly notice Opus4.8 refusing or deferring certain tasks as opposed to 4.7 gladly taking on. I wonder if it’s bc of new alignment guardrails in pretraining🤔
@bcherny@trq212