Hourly billing is 72% of legal pricing today.
GCs expect that share to fall to 44% within two to three years.
That drop is already in their planning.
Firms still haven't put it in their proposals.
Deloitte Legal survey, side by side:
- 79% of GCs raised AI budgets this year
- 61% already in deployment
- 66% expect to insource more work
- 4% say outside counsel showed them any AI benefit
Four lines. One panel conversation.
A model can top a legal reasoning leaderboard and still choke on the file in your intake folder.
Typo in the defendant's name. Injury date buried on page four. Nobody cleans that up before it hits a paralegal, let alone a machine.
Benchmarks score clean questions. Your docket runs on messy ones.
the play:
1. Ignore the single leaderboard screenshot in the deck
2. Pull 10 real files from last month's docket
3. Write 5 prompts your team already types
4. Partner-grade every output (invented cite = automatic fail)
5. Buy the tool that survives your files
90 minutes.
Cheaper than one bad pilot.
41.1% of the time, models fill in missing details instead of asking.
Small typos drop performance 20–33%.
Reasoning tweaks cut abstention ~24%.
Real case Qs hallucinated 58–88% in one 2024 study.
Leaderboard day is the easy day.
Intake day is the job.
A model can top a legal leaderboard and still choke on the PDF in your intake folder.
Clean exams measure legal reasoning.
Your docket runs on messy client files.
Full breakdown in the article.
Demos test charm.
LAB and LAB-AA test multi-step legal work.
May 2026 gave the market a shared scoreboard: Harvey’s LAB, plus independent LAB-AA runs (~120 private tasks, 24 practice areas).
Stop buying on vibes.
Start buying on task fit and supervision load.
Legal AI scoreboard, use it in 60 seconds:
1. Ask which underlying model they use
2. Check that model on LAB-AA
3. Ignore overall rank if your practice area is weak
4. Run a bakeoff on one real low-stakes matter
5. Test for invented facts and cites first
Scoreboard = shortlist filter.
Your files = the real exam.
Take this into the Monday partner meeting:
The best models on the new public legal-agent scoreboard still full-pass under 15% of the time.
Buy tools.
Require a human check on every client-facing draft.
Best full-pass rate on this public multi-step legal test: about 14.2%.
That’s Claude Fable 5 on LAB-AA.
Next names sit nearer 7.5%.
Full-pass means every requirement met.
Not “looks fine in a demo.”
Treat that as the real weather report on legal AI right now.
Expert AI intake in 60 seconds:
1. Will you use AI to form opinions or select materials? Which tool?
2. Closed enterprise system, or consumer product?
3. Are prompts, uploads, and material outputs logged and preserved?
4. Would you be comfortable producing every prompt if ordered?
5. Does the engagement letter and ESI protocol name AI by name?
6. Who writes and edits the prompts: expert, assistant, or counsel?
7. How will AI use be disclosed in the report?
Run this before the first session, not after the motion to compel.
They had a Rule 29 deal that covered expert “notes, drafts, or communications.”
They tried to call the AI prompts “notes.”
The court said agreements that block discovery have to be quite clear. Silence on AI is not a shield.
If you want prompts protected, write the words “AI prompts” into the paper before anyone opens the tool.
When the other side’s expert used AI, ask for:
1. Every prompt, query, and system instruction
2. Every document uploaded as input
3. Tool name, version, and input retention terms
4. Whether counsel wrote or edited any prompts
5. Logs of what the model returned vs. what made the report
#5 is the cross. If the model flagged bad facts and the report dropped them, the prompt trail is the map.
A testifying expert used AI to cull documents for their report.
A federal magistrate ordered the prompts produced as methodology under Rule 26.
Order is stayed on objection. The logic is already in every firm litigation update.
Treat every expert prompt like a document that can land in a deposition.
Full breakdown in the article below.