@gabepereyra@harvey Ahh, so it’s less trying to train firm data for memorisation and more how the firm operates within the harness, that could be interesting. Will be interested to see the first evals etc on this as it is an ambitious project.
Interesting, I’m wondering where you guys see the benefit of say a firm specific model vs something like your Tenet model. From my understanding training knowledge so to speak into a model doesn’t beat regular retrieval type systems. So what is the use case compared to say a firm agnostic post trained legal model?
We built a 9B model that matches frontier systems at knowing when not to answer.
Introducing Cairn Scout: the first model in Clava's Cairn lineup and the grounded-answer engine inside our product, Helm.
Most RAG systems share the same quiet failure. When the retrieved documents don't contain the answer, the model answers anyway. When two sources disagree, it silently picks one and states it as fact. In legal, finance, insurance and healthcare work that isn't a minor flaw. It's the entire risk.
Scout is built to do the opposite. It answers strictly from the sources in front of it, says plainly when the answer isn't there, flags genuine conflicts instead of resolving them for you, holds up against prompt injection, and returns a structured, machine-readable result every time.
So we tested it properly.
The test: 452 held-out grounded-answering tasks across 63 industries, scored by blind LLM judges reading each answer against the source. Format-agnostic: models are judged only on whether they reached the right conclusion.
The field: DeepSeek V4, Gemini 3.1 Flash Lite, Gemini 2.5 Flash, GLM 5.2, Kimi 2.6, and gpt-oss-120B.
The results:
→ 90.5% judged correctness. Top cluster, within 1.4 points of the best model tested, ahead of systems more than 13x its size.
→ 100% output-contract compliance. Scout was the only model that returned every required field of the structured output a real application needs, on every single item. Every other model: 0%.
→ 97% on source conflicts, where most large models scored between 17% and 78%. Where they silently pick a side, Scout flags the disagreement. That one behaviour is why this model exists.
→ All of it self-hosted on a single 24GB GPU, at around a tenth of the price of the large models it was measured against.
It's not perfect: Scout is currently too cautious on the simplest extractions, and that's the first fix in v2. But the thesis holds. For grounded, high-stakes work, a small model trained to be honest about its limits beats a large model that will confidently tell you anything.
Clava builds Cairn models to run Helm. Scout is the first. Full benchmark methodology and eval data publishing soon.
If you're building RAG somewhere a wrong answer actually costs something, talk to me.
Our model thesis
The industry default is to rent the biggest model available and prompt it into behaving. For sensitive work, we think that's backwards.
A frontier model is expensive to run, hard to inspect, and hard to constrain. It can do ten thousand things, which means it can fail in ten thousand ways. In regulated work you don't need ten thousand things. You need one thing done the same way every time, with evidence.
A small model trained hard on one task has a smaller job, a clearer input, a clearer output, and fewer places to fail. It runs on hardware you control, so your data never leaves your environment. It costs a fraction per call. And when it fails, it fails in ways you've measured and can route around.
That's Cairn: our family of task-specialised models. Each one handles a single part of the workflow. Ranking evidence. Drafting the grounded answer. Checking claims against sources. Redacting sensitive data before it goes anywhere.
None of them tries to be everything. That's the point.
The interesting question isn't "how smart is the model." It's "what happens when it's wrong." Small, scoped, and measured beats big, general, and confident. Receipts coming soon.
This is Helm.
Most firms don't need another chatbot. They need the repetitive, document-heavy admin to move through the business faster, with the right checks in place.
One Helm workflow can take a messy intake email, identify the client, pull the right records, spot what's missing, draft the next step, route it to the right person, and log every bit of it.
Emails become case notes. Files get checked before anyone wastes time reviewing them. Replies get drafted from approved material. Summaries land ready for the meeting.
The important word is controlled:
• Models only touch approved data sources
• Sensitive actions wait for human approval
• Every prompt, source, output, reviewer, and override is recorded
Because in legal, financial, insurance, and healthcare work, "the AI said so" is not an acceptable answer. You need to know what the system used, what it ignored, what changed, and who signed it off.
Helm deploys in our secure cloud or entirely on your infrastructure, up to fully air-gapped. Your workflows, your data, your logs.
This is the difference between AI as a toy and AI as part of an operational system.
Introducing Clava.
We're an independent AI lab based in Inverness. The first in the Scottish Highlands.
We build private AI systems for organisations that handle sensitive information: legal, financial, insurance, healthcare, and other work where a wrong answer actually costs something.
Two things make us different.
We train our own models. Cairn is our family of small, task-specialised models. The thesis: a compact model trained hard on one job beats a giant generalist on that job, at a fraction of the cost, on infrastructure you control.
We build for control, not demos. Helm runs AI workflows with approved sources, human review gates, and a full audit record. In regulated work, "the AI said so" is not an answer.
First model drops soon, with a public benchmark and the full eval data to check our numbers.
Own the stack. Control the boundary. Keep the evidence.