I've found the following skills helps: https://t.co/OfOt93P5IX
But yes these recent Claude models are all horrible at being a language model.
Funny language models are no longer good at generating language. Hahaha
I wonder if this has to do with some weird attention sinks caused by the post-training.
Diagnosing what exactly causes this issue with recent models (and a possible fix) is a good research problem for someone to work on...
🤖🧠NEW PAPER🧠🤖
(The result of an 8-year project!)
LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it?
Our finding: LLM representations have implicit symbolic structure!
Link in thread ⬇️
1/n
Congratulations Levent, Tristan, and OpenAI! What a miraculous time to be alive!
I view this as the first successful achievement of Recursive Self Improvement or RSI. Models get increasingly better at Math and get increasingly used by mathematicians to solve all kinds of difficult problems in their domain, thereby contributing very rare and very hard tokens that are then used in the model training to further improve the model's math capabilities: generalizing and drawing connections across sessions and everything else the models learn from; which is then served back, and so on and so forth.
This happened first for math but will eventually happen for every domain.
Positional embeddings discard syntax, keeping only order & distance. We show that syntax can be reintroduced by adding two lightweight tag embeddings, looked up from small tables, into the model's input embedding, resulting in improved fundamental language modeling. Delighted that our paper on Syntax-informed Positional Embeddings (SiPE) was accepted to EMNLP 2026 main conference! 🥂 w/ Hyungji Kim & @msurd
#EMNLP2026 #InductiveBias #PreTraining
Positional embeddings encode order and distance, but are agnostic to syntax.
"Please move this large file to another folder."
To act, move must bind file (object) and folder (destination). Absolute PE: no link. Relative PE/RoPE: "six tokens apart." A dependency parse tree: one arc. Introducing Syntax-informed Positional Embeddings (SiPE) (1/n)🧵
Please welcome to the world a beautiful new geometric object, to do with a problem i’ve always loved. claude really contains multitudes:D Does S^6 admit a complex structure?
Yup
Introducing GEN-1.5, a one-shot learner.
It can learn new tasks in a few seconds. Show it what to do, and it generalizes.
This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
📣Call for contributions + co-authorship!
RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D.
Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress!
Contributors receive:
→ Co-authorship on the RSI Bench research paper
→ $2,000 per accepted task for the initial 50 tasks
→ Modal compute credits to build + iterate
→ Access to the RSI Bench research community
Register below for full requirements.
Pretty interesting rethinking paper on "discouraging" self-evolving loops.
So the background is that most current self-evolving loops run their search directly on the test set. It kind of makes sense as harness search needs accurate, verifiable feedback to make grounded edits.
However, that quietly turns self-evolving loops into a form of test-time scaling, which is exactly what this paper argues.
Specifically, the paper points out that a loop that repeatedly evaluates and revises candidates against task feedback, then reports on those same tasks, is logically a test-time search procedure. So its gains should be measured against test-time scaling under matched feedback and inference budgets — otherwise you can't tell whether it discovered a better harness or just spent more compute.
With that in mind, the paper runs four methods under the same budget:
1. parallel sampling — fixed harness, k independent trajectories per task, with a self-judge or unit tests picking the final answer
2. sequential refinement — fixed harness, k retries in a row. Each round summarizes the previous attempt into context and tries again (essentially prompt refinement)
3. harness evolution — the standard self-evolving loop. One shared harness, revised each round from feedback pooled across all tasks
4. harness scaling — the per-instance counterpart. Each task evolves its own harness
The results are very interesting.
Harness evolution doesn't beat plain parallel sampling, and without verifiable feedback it can even fall below single-attempt direct sampling with the initial harness.
Its gains also show up at pass@5 but barely at pass@1, which implies that the improvement comes from taking multiple attempts, not from the harness getting better.
And on a disjoint search/eval split the evolved harness transfers almost nothing to held-out tasks, which means the edits memorize task-specific fixes rather than distill reusable strategies.
There is one caveat, which is that the "unified budget" only counts inference on the tasks, not the compute spent generating harnesses. But this flaw kind of favors harness evolution, and it already loses.
So in my opinion this really shows that existing self-evolving loops might just be a different way of applying test-time scaling, rather than some new intelligence discovery.
And we should focus more on making generalizable self-evolving loops work!
RL against rubrics quietly reward-hacks: the training judge keeps scoring higher while true quality falls.
One-line fix: randomly drop part of the rubric each step, so the policy never optimize the same set of rubrics twice.
📄 https://t.co/BSrzyJOciF
"How to Steal Reasoning Without Reasoning Traces"
Hiding chain-of-thought may not stop reasoning distillation.
This paper introduce trace Inversion which reconstructs synthetic reasoning traces from only a black-box model’s answers and optional summaries, then uses them to fine-tune a student model.
These reconstructed traces transfer reasoning better than answers, summaries, or even the weaker surrogate model’s own traces.
Demonstrating that reasoning capability still leaks even if you hide them.
https://t.co/mEcaGqN22B
P.S. all these experiments were run on a tiny academic compute budget — if you’d like to make a generous contribution of GPUs & see it scaled to LLM pre-training… 👀 @nvidiaai@PrimeIntellect
Positional embeddings encode order and distance, but are agnostic to syntax.
"Please move this large file to another folder."
To act, move must bind file (object) and folder (destination). Absolute PE: no link. Relative PE/RoPE: "six tokens apart." A dependency parse tree: one arc. Introducing Syntax-informed Positional Embeddings (SiPE) (1/n)🧵
TL;DR Our results indicate that the age of (text-only) pre-training may not be over. Raw text holds more signal than tokens alone — syntax trees are just one structure we can model explicitly.
Work done w/ Hyungji Kim & @msurd.
End���.
Blogpost: https://t.co/SzukI3g9Ef
Paper: https://t.co/mMbGmPPYbq
Syntax in decoders is best injected at layer 1
Which layers should receive the prior? We sweep Transformer-XL, injecting into layers ℓ…16 for every ℓ.
Layer 1 wins. Skipping just the first layer drops SyntaxGym from 80.6 to 73.5; later entry points only get weaker. In a decoder, syntax belongs in the earliest layers. (11/n)
Dependency arcs also carry relation labels (nsubj, obj, …; 40 total). Does this richer signal beat using two coarse directional tags?
It does not. Added, concatenated, or swapped in for the terminal tag, every deprel variant scores below simply adding terminal + non-terminal tags (10/n)
How should the tag embeddings be integrated into the model? We experiment with four strategies on a simple model with absolute PE: RoBERTa. Avg. GLUE score; baseline 72.30):
addition of embeddings: 72.84
addition on skip connection: 72.40
learned α-interpolation: 72.05
concat + projection: 70.91
Plain addition, with no extra parameters, wins. (9/n)
SiPE improves fundamental language modeling. Pre-training with SiPE simply produces a better language model: consistent gains on the GLUE benchmark across four architectures (RoBERTa (absolute PE, encoder), DeBERTa (relative PE, encoder), ModernBERT (RoPE, encoder) and Transformer-XL (relative PE, autoregressive decoder)), up to +8.2% for the decoder, where all eight tasks GLUE improve. (8/n)