Tokenomics in action: In KYB applications, the most expensive LLM (Opus) can be the cheapest model overall. In our latest update on KYBench, we test the current generation of frontier models on a task every financial institution cares about: Onboarding businesses that have a clean footprint. No criminal record, no ongoing investigations, etc.
Compared to our initial evaluation, even more models clear the test with flying colors: We find that 19 of 23 models detect 100% of the problematic businesses. And even relatively cheap and cost-efficient models such as Claude Sonnet 5 and GPT-5.6 Luna do really well.
The remaining difference in performance comes from how much unnecessary analyst work the models generate: false positives an analyst has to clear, and findings the model misses and leaves for a human to dig up. The total economic cost takes into account not just the token cost but also the human cost of that work. Since human labor is expensive, so are both kinds of mistake. And that’s how the model with the priciest tokens among the front-runners (Opus) ends up being cheapest overall: In addition to flagging the bad guys, it characterizes them most completely, so it leaves analysts the least to chase down.
Read the full benchmark on the Taktile Labs website: https://t.co/mMEZxDdX0u
Prompt injection is a real AI risk in financial services. But our new @Taktile Labs benchmark shows the risk is lower than you’d think:
We just published PIBench, the first benchmark of prompt-injection resistance specifically targeting agentic underwriting. 16 frontier models tested across three providers and five attack vectors, with and without our untrusted-content tagging layer.
TL;DR: Several frontier models can now resist every attack we designed; the risk is real but it is a solvable problem. Our top three findings:
(1) Provider predicts vulnerability better than model tier. OpenAI's GPT family resisted virtually every prompt injection in PIBench (99.6% average across models), while Gemini models averaged 85.3%. Even OpenAI's smallest model, GPT-5 Mini, defended every attack.
(2) Tagging works, and it's essentially free. Adding explicit boundary tags around untrusted content lifts average defense success from 93.2% to 97.7%, and does that without producing any false positives.
(3) The residual risk is concentrated in specific models and attack vectors. After tagging, nearly every attack vector is solved, across almost all models. The remaining vulnerabilities lie in sophisticated prompt injections embedded within uploaded documents.
Huge thanks to Koen Roelofs and Jakob Schmitt for collabing on this!
Read the full benchmark: https://t.co/m5JywutLRq
We raised a $110M Series C to power AI transformation in financial services.
Led by Goldman Sachs, with participation from @balderton, @IndexVentures , @dig_ventures, @tigerglobal, @VisionariesVC, and @ycombinator.
Great news for banks and insurers.
Tough news for fraudsters and money launderers.
Tough news for credit processes that make businesses wait weeks for an answer.
Tough news for onboarding workflows that reject good customers.
Tough news for claims processes where a single review takes days.
We’re serious about solving our customers’ problems, so we hired the toughest people on the planet to get the word out:
New Yorkers.
If they didn’t convince you yet, here’s what we want everyone in our industry to know:
AI can now reliably automate the decisions that define performance in banking and insurance. This was not true just 6 months ago - and it creates huge opportunity for efficiency gains and better customer experiences.
But model capability is not the full solution.
The challenge is making every AI-driven decision reliable, controlled, and ready for the most regulated industry.
That’s what we’re building at @taktile_org.
From New York to the world.
To our customers, partners, investors, and team: thank you for getting us here. Very excited to collaborate with the @GoldmanSachs Asset Management team on the next chapter.
Learn more: https://t.co/T6BmNFc7gm
Latest benchmark is out. Doing adverse media searches manually will look very strange from the perspective of someone looking in a Q4/26. Compliance and risk are changing for the better.
Taktile Labs and @p0 just published a new benchmark that puts agents to the test on a core KYB task: adverse media search. The results might surprise you →
We tested an adverse media agent built on 7 frontier AI models across 47 real businesses.
The takeaway: on this task, AI agents can now outperform human analysts - but collaboration is still essential.
→ Evidence quality is higher when using AI agents.
On a 20-point rubric spanning evidence/source quality, entity identification, and risk assessment, humans averaged 13.50/20 vs. 14.60/20 for AI agents.
→ A hybrid human x AI is the ideal setup.
An “agent screens first, human reviews uncertain cases” model reduces analyst workload by 93% while keeping compliance in check.
→ Lower-cost models still hold up.
Less expensive configurations can match human performance, and the cost of high accuracy will continue to fall as new models are released.
→ Hallucinations are less of a risk than many believe.
We saw a 0% hallucination rate in this dataset (no fabricated sources). False positives primarily came from the agent flagging too eagerly.
Huge shoutout to David Ahn, @maximilianeber, @paraga and Sahith Jagarlamudi driving the research working together to give banks the evidence they need to deploy AI with confidence.
Read the full study: https://t.co/0O7Vbo3LTF
Pretty cool results from our first publication benchmarking agent performance for banking and insurance: Machines can now analyze financials.
https://t.co/1iYeMzOuPs
“nothing says Wall Street like lobsters because somebody is always getting cooked.” Thanks @LizClaman at @FoxBusiness for breaking the story of why @taktile_org put a giant lobster on Wall Street.
wasn’t to scare the bulls!
Taktile Labs is here to keep the bulls running with reliable AI.
see our latest research via link in comments
Something special happened in Q4/25:
The latest generation of LLMs — Opus 4.5, GPT-5.2, Gemini 3 Pro — is finally strong enough for most tasks in banking and insurance. We are launching Taktile Labs to bring the evidence.
https://t.co/fo7XUI0KBW
Time to reveal who let the 🦞 out ;) Today, @Taktile launches Taktile Labs. We dropped the lobster on Wall Street to ask the question: are banks ready for autonomous agents?
With our applied AI research institute, we aim to bridge the gap between what frontier models can now do - and what regulated institutions need in order to trust it.
Our first benchmark shows the latest models can beat human accuracy on very complex banking tasks: 96%+ vs. 89% in financial spreading.
The models are ready. Now the industry needs evidence, benchmarks, and practical frameworks to ensure they work reliably at scale.
That is what Taktile Labs is built for.
AI is coming to financial services - let's make sure we can trust it.
Excited to drive this with a stacked internal team and many incredible individuals on our Research Council and Advisory Board.
Thanks to Bradesco’s Fagner Abreu, Parallel’s @paraga , Founder, Investor, and Morgan Stanley Lead Director Tom Glocer, Harvard Business School Professors Robin Greenwood and Karim Lakhani, Harvey’s Ben Liebald, Camunda’s Daniel Meyer, Cursor’s Jonas Nelle, ROC Partners’ Tina Reich, Equifax’s Harald Schneider, Suno’s @MikeyShulman, Intuit’s Henry Venturelli, Allianz Partners’ Pieter Viljoen, Flexcar’s Michael Zambrano, and Varo Bank’s Jill Zucker Sheckman.
Learn more at: https://t.co/wV9UVEPIES
(nothing AI generated about it btw, we worked with NYC artist @AndrewLoganAMW to build the lobster from scratch)
Time to reveal who let the 🦞 out ;) Today, @Taktile launches Taktile Labs. We dropped the lobster on Wall Street to ask the question: are banks ready for autonomous agents?
With our applied AI research institute, we aim to bridge the gap between what frontier models can now do - and what regulated institutions need in order to trust it.
Our first benchmark shows the latest models can beat human accuracy on very complex banking tasks: 96%+ vs. 89% in financial spreading.
The models are ready. Now the industry needs evidence, benchmarks, and practical frameworks to ensure they work reliably at scale.
That is what Taktile Labs is built for.
AI is coming to financial services - let's make sure we can trust it.
Excited to drive this with a stacked internal team and many incredible individuals on our Research Council and Advisory Board.
Thanks to Bradesco’s Fagner Abreu, Parallel’s @paraga , Founder, Investor, and Morgan Stanley Lead Director Tom Glocer, Harvard Business School Professors Robin Greenwood and Karim Lakhani, Harvey’s Ben Liebald, Camunda’s Daniel Meyer, Cursor’s Jonas Nelle, ROC Partners’ Tina Reich, Equifax’s Harald Schneider, Suno’s @MikeyShulman, Intuit’s Henry Venturelli, Allianz Partners’ Pieter Viljoen, Flexcar’s Michael Zambrano, and Varo Bank’s Jill Zucker Sheckman.
Learn more at: https://t.co/wV9UVEPIES
(nothing AI generated about it btw, we worked with NYC artist @AndrewLoganAMW to build the lobster from scratch)
We’re thrilled to announce that Jason Mikula (@mikulaja) has joined Taktile as our new Head of Industry Strategy for Banking and Fintech 🎉
Please join us in welcoming him to Taktile!
Learn more about Jason's new role here👇
https://t.co/PzogoUcuNe
We’re thrilled to be featured in @Siftedeu ultimate list of the 100 most promising B2B SaaS startups valued under $1bn in 2024! 🚀
Check it out here: https://t.co/nMIs7NgVWi