Stop Testing. Start Measuring: The Physics of LLM Reasoning
Current AI evaluation is stuck in a behavioralist loop—relying on curated benchmarks that are often leaked, biased, or shallow. We are proposing a shift toward the intrinsic physics of LLMs.
HypothesisForge: open-source, autonomous ML-research agents that hypothesize, run a real GPT-2 experiment, critique the stats, and write the report. Apache 2.0. Follow us for more.
https://t.co/svPYm2q3WB
One point I made that didn’t come across:
- Scaling the current thing will keep leading to improvements. In particular, it won’t stall.
- But something important will continue to be missing.
“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling.
Is the belief that if you just 100x the scale, everything would be transformed?
I don't think that's true. It's back to the age of research again, just with big computers.”
@ilyasut
We are happy to announce PLDR-LLM support for Huggingface Transformers library! We also release three high performing fine-tuned small size PLDR-LLMs for sentiment analysis, token classification and question answering.