@Cozy2054934 Facts g, the raw rank is only half the story. I’d also run the same tasks through two different agent scaffolds and see which model’s score collapses first and that gap usually reveals more about real reliability than the leaderboard does.
everyone checks the same AI leaderboards.
Chatbot Arena.
SWE-bench.
MMLU.
but there are some absolutely insane AI leaderboards buried across the internet that tell you WAY more about what models can actually do.
here are 6 almost nobody checks:
1. Remote Labor Index (RLI)
this might be one of the most important AI benchmarks on earth.
instead of asking an LLM trivia questions, RLI gives AI agents real remote-work projects originally sourced from professional freelancers.
240 projects.
actual economically valuable work.
the AI’s output is then compared against professional human work.
so instead of asking:
“is this model intelligent?”
you can ask:
“could this thing actually replace part of someone’s job?”
the current leaderboard is brutal too. even the leading agents are nowhere near solving everything. (Scale Labs)
2. HiL-Bench
this tests something almost nobody measures:
does the AI know when it DOESN’T know enough?
the benchmark intentionally removes critical information from tasks.
the agent has to recognize:
“I cannot safely infer this.”
then use an ask_human() tool to get clarification.
300 tasks.
1,131 hidden blockers.
models that look extremely capable when given perfect instructions can fall apart when they have to decide whether they should ask you a question.
this is basically a leaderboard for AI judgment. (Scale Labs)
3. MCP-Atlas
want to know which models are actually good at using tools?
check this.
MCP-Atlas throws agents into 1,000 tasks across 36 MCP servers and 220 tools.
they have to find the correct tool inside noisy menus, format parameters properly, chain multiple calls together, recover from errors, and still produce the right answer.
this is much closer to how future AI employees will actually operate than answering multiple-choice questions. (Scale Labs)
4. Aider Polyglot Leaderboard
most coding benchmarks ask:
“can the model generate code?”
Aider asks something more useful:
can it correctly modify an existing codebase without a human fixing its edits?
its Polyglot benchmark uses 225 difficult programming exercises across:
C++
Go
Java
JavaScript
Python
Rust
it even tracks whether the model correctly follows the required edit format.
if you use AI coding agents, this leaderboard is worth bookmarking. (Aider)
5. LiveBench
one giant problem with AI benchmarks:
models can eventually train on the benchmark itself.
LiveBench was built around fighting that.
its latest release evaluates models across 23 objective tasks and 7 categories, including reasoning, coding, agentic coding, math, data analysis, language, and instruction following.
the tasks are periodically refreshed, and it even shows cost per successful task.
that last metric is underrated.
the “best” model isn’t always the best model if another gets nearly the same result for a fraction of the price. (LiveBench)
6. HELM Capabilities
Stanford has its own leaderboard called HELM.
what makes it interesting isn’t just the scores.
HELM provides prompt-level transparency, so you can actually inspect what models were tested on instead of blindly trusting one mysterious number.
its goal is to measure multiple core model capabilities using reproducible evaluations. (Stanford CRFM)
the deeper you get into AI, the less useful the question:
“what’s the smartest model?”
becomes.
the better questions are:
which model can use tools?
which one knows when to ask for help?
which one can edit real code?
which one can perform economically valuable work?
which one gives me the most intelligence per dollar?
there isn’t one AI leaderboard.
there are hundreds of tiny battlefields testing completely different forms of intelligence.
and most people are only watching one of them.
most people are trying to fix AI hallucinations at the wrong layer.
they keep tweaking prompts and adding retrieval.
meanwhile, researchers are literally changing how models choose their next token, verify their own answers, and measure whether they might be making shit up.
these are 7 hallucination reduction techniques almost nobody talks about:
1. DoLa: Decoding by Contrasting Layers
different layers inside a transformer represent information differently.
earlier layers contain more basic representations.
later layers tend to contain more developed semantic and factual information.
DoLa compares the token probabilities produced by earlier and later transformer layers while the model is generating.
then it favors tokens where the later layers show stronger confidence.
no web search.
no external database.
no extra fine tuning.
you literally change how the model decodes its own internal knowledge.
researchers reported substantial TruthfulQA improvements on LLaMA family models.
2. Self-RAG
normal RAG works something like this:
the user asks a question.
the system retrieves documents.
those documents get added to the context.
the model answers.
Self-RAG adds another layer.
the model can decide whether retrieval is even necessary, whether the information it found is relevant, whether its answer is supported, and whether the response is actually useful.
retrieval becomes dynamic instead of happening blindly every time.
3. CRAG: Corrective RAG
RAG has a huge weakness.
bad retrieval can actually make hallucinations worse.
if your retriever pulls garbage, the model can confidently reason from garbage.
CRAG adds a retrieval evaluator.
retrieved information can be judged as useful, questionable, or bad.
if the retrieval quality is weak, the system can trigger another search.
it can also remove irrelevant sections and keep only the useful information.
basically:
RAG with quality control.
4. Context-Aware Decoding
this one happens during generation.
the model generates probabilities using the supplied context.
then the system compares those probabilities against what the model would have predicted without that context.
if your documents support one answer but the model memory strongly prefers another, the decoding process can push the model harder toward the supplied evidence.
this is useful when a model memory conflicts with the information you gave it.
5. Chain-of-Verification
instead of trusting the first answer, the model creates a draft.
then it generates questions that could verify the claims inside that draft.
then it answers those verification questions separately.
finally, it uses those answers to rewrite the original response.
the important part is doing the verification separately.
if the model constantly sees its first answer while checking itself, it can anchor onto the original mistake.
CoVe gives it another chance to catch the error.
6. SelfCheckGPT
this one does not even need external sources.
ask the model the same factual question multiple times with some randomness.
then compare the answers.
if the model actually knows something, its answers tend to stay semantically consistent.
if it is making shit up, the details often start drifting.
one answer says 1987.
another says 1991.
another names a completely different person.
the disagreement itself becomes a hallucination signal.
basically:
make the model testify against itself.
7. Semantic Entropy
this takes that idea even deeper.
normal uncertainty measures can get confused because two answers can use completely different words while meaning the exact same thing.
Semantic Entropy groups outputs by meaning.
the system samples multiple answers and asks how uncertain the model is about the actual semantic answer.
10 differently worded answers that all mean the same thing:
low semantic uncertainty.
10 answers making fundamentally different claims:
high semantic uncertainty.
once uncertainty gets too high, the system can retrieve evidence, run another verification step, escalate to another model, or simply refuse to guess.
that is where hallucination reduction gets way more interesting.
you can attack the problem through retrieval, decoding, verification, training, uncertainty estimation, and post generation checking.
the goal is not just making the model smarter.
it is building a system that can recognize when its own answer might be bullshit before you ever see it.
everyone knows AI hallucinates.
most hallucination fixes stop at retrieval.
the interesting stuff starts after that.
researchers have built some genuinely weird ways of stopping LLMs from confidently making shit up.
here are 7 hallucination-reduction techniques almost nobody talks about:
1. DoLa — Decoding by Contrasting Layers
this one is insane.
different layers inside a transformer represent information differently.
earlier layers contain more basic representations.
later layers tend to contain more developed semantic and factual information.
DoLa compares the token probabilities produced by earlier vs later transformer layers while the model is generating.
then it favors tokens where the later layers show stronger confidence.
no web search.
no external database.
no extra fine-tuning.
you literally change how the model decodes its own internal knowledge.
researchers reported substantial TruthfulQA improvements on LLaMA-family models.
2. Self-RAG
normal RAG works something like:
question → retrieve documents → put them in context → answer.
Self-RAG asks something more interesting:
does the model even need retrieval right now?
the model learns special reflection tokens that help it decide when to retrieve information, whether that information is relevant, whether its answer is supported, and whether the response is actually useful.
retrieval becomes dynamic instead of blindly happening every time.
3. CRAG — Corrective RAG
RAG has a problem:
what happens when your retriever pulls garbage?
now you’ve just given the model more garbage to confidently reason from.
CRAG adds a retrieval evaluator.
retrieved information can be judged as useful, questionable, or bad.
if retrieval quality sucks, the system can trigger additional retrieval like web search.
it can also strip out irrelevant sections and keep only the useful information.
basically:
RAG with quality control.
4. Context-Aware Decoding
this one happens during generation.
run the model with the provided context.
then compare that against what the model would’ve predicted without the context.
the decoder amplifies the difference.
so if your documents say X but the model’s pretrained memory strongly wants to say Y, CAD pushes the model harder toward the supplied evidence.
this is especially useful when the model’s memory conflicts with the information you’ve given it.
5. Chain-of-Verification
instead of:
question → answer
you do:
question
↓
draft answer
↓
generate verification questions
↓
answer those questions independently
↓
rewrite the original answer
the important part is independently.
if the model sees its original answer while checking itself, it can anchor onto the original mistake.
CoVe gives it another shot at finding the truth without constantly staring at its first answer.
6. SelfCheckGPT
this one doesn’t even need external sources.
ask the model the same factual question multiple times with some randomness.
then compare the answers.
if the model actually knows something, its answers tend to stay semantically consistent.
if it’s making shit up?
the details often start drifting.
one answer says 1987.
another says 1991.
another names a completely different person.
the disagreement itself becomes a hallucination signal.
basically:
make the model testify against itself.
7. Semantic Entropy
this takes that idea deeper.
normal uncertainty measures can get confused because:
“Paris is the capital of France”
and
“France’s capital is Paris”
use different tokens while saying the exact same thing.
Semantic Entropy groups outputs by meaning.
the system samples multiple answers and asks:
how uncertain is the model about the actual semantic answer?
10 differently worded answers that all mean the same thing:
low semantic uncertainty.
10 answers making fundamentally different claims:
high semantic uncertainty.
everyone knows ChatGPT, Claude, Gemini, Grok, and DeepSeek.
but there’s an entire layer of LLMs underneath the mainstream that almost nobody talks about.
here are 4 you should know:
1. Jamba
Jamba is built by AI21 Labs, and its architecture is what makes it interesting.
most LLMs you use are built heavily around Transformers.
Jamba mixes Transformer layers + Mamba layers + Mixture-of-Experts (MoE).
the idea is simple:
keep the intelligence and flexibility of Transformers while using Mamba to process long sequences more efficiently.
AI21 has continued developing the architecture into newer Jamba generations aimed heavily at efficient, private, enterprise AI.
basically: Jamba is an example of what LLM architecture could look like when companies stop assuming every model needs to be a traditional Transformer.
2. DBRX
DBRX was created by Databricks.
this thing is massive.
132B total parameters.
but because it uses Mixture-of-Experts, only around 36B parameters are active for a given input.
instead of activating the entire network every time you ask something, DBRX routes the request through a subset of specialized “experts.”
it was also pretrained on 12 trillion tokens of text and code.
DBRX is especially interesting if you care about coding, RAG, enterprise AI, or how MoE systems actually work at scale.
3. OLMo
OLMo might be the most interesting model here if you actually want to understand how LLMs are built.
it comes from Ai2.
instead of just releasing model weights and calling something “open,” Ai2 has released models alongside training data, code, recipes, evaluations, checkpoints, and research.
OLMo has now grown into an entire family of models.
you can basically look under the hood.
for researchers, builders, and anyone trying to understand what happens BEFORE you get the final chatbot, OLMo is a gold mine.
4. Yi
Yi is the LLM family from https://t.co/hzv5TMsHB5.
it has been built for things like:
reasoning
coding
text generation
chat
mathematical problems
general language understanding
and despite being far less recognizable to the average person than ChatGPT or Claude, Yi became another serious entrant in the open-model ecosystem.
the bigger lesson:
“AI” is much bigger than the 5 chatbots everyone talks about.
there are hundreds of models experimenting with different architectures, training methods, data mixtures, parameter counts, context systems, and ways of routing computation.
Jamba.
DBRX.
OLMo.
Yi.
go far enough down the LLM rabbit hole and ChatGPT starts looking like the surface.
most people using Claude Code are wasting tokens without realizing it.
type:
/context
it shows you exactly what is eating your context window.
then fix the actual problem:
huge conversation → /clear between tasks
long debugging session → /compact
bloated CLAUDE.md → delete unnecessary instructions
massive terminal output → use quiet flags
Claude searching for files → @mention the exact file
tools you don’t need → disable them
the sneaky one is terminal output.
if Claude runs a command that dumps 5,000 lines, that output gets added to the conversation and can keep getting carried forward.
clean context = fewer tokens + better Claude.
the new chatgpt update is way bigger than people realize.
you can now give ai a job instead of giving it prompts.
they’re called dots.
a dot is basically an ai worker you give a goal to, and it can keep working even after you close chatgpt.
you give it a job, connect the apps it needs, and tell it what you want done.
then it can:
use its own computer and browser
connect to 4,000+ apps
work on many tasks at once
remember how you like things done
learn from your feedback
watch for new info
react when something changes
ask you when it needs a decision
the easiest way to understand the difference is with an example:
instead of opening chatgpt every day and saying:
“check customer feedback”
then:
“find the biggest problems”
then:
“figure out how to fix them”
you can give a dot the goal once.
it can keep watching the feedback, find common problems, work on fixes, test them, and bring the finished work back to you.
you can also decide what it can do on its own and what needs your approval first.
that’s the big shift:
chatgpt used to be something you gave prompts to.
now you can give it a job.
80% of leveling up in life is literally just remembering what you told yourself you were going to do.
There are so many distractions now that people forget their own realizations.
That’s why u need to journal. Slap a fuckin sticky note to your forehead if you need to.
Because if u realize something important, then forget it a week later, you’re right back where you started.
I remember being one of the only kids in school to use Claude almost 4 years ago.
Felt like I had an infinity stone.
Now everybody uses it.
At the time it had such an upper hand in writing against ChatGPT so it was inevitable to get this big.
Good job @claudeai