Most companies are experimenting with AI. Few can show measurable ROI.
One of the main issues is the lack of operating evidence: a clear view of how operations actually run and where AI can create value.
๐๐๐ฆ๐ฉ๐ขโ๐ฌ ๐๐ ๐๐ซ๐๐ง๐ฌ๐๐จ๐ซ๐ฆ๐๐ญ๐ข๐จ๐ง-๐๐ฌ-๐-๐๐๐ซ๐ฏ๐ข๐๐ ๐๐ฅ๐จ๐ฌ๐๐ฌ ๐ญ๐ก๐๐ญ ๐ ๐๐ฉ.
Within days, @lampi_ai analyzes authorized context across relevant company data to map how the business operates.
It builds digital twins and turns signals into actionable intelligence:
โ An operating digital twin showing how operations run
โ A customer twin identifying revenue opportunities
โ ROI prioritization showing the highest-impact AI opportunities
Each opportunity is assessed by return, feasibility, data readiness, and time to value.
The strongest opportunities are converted into governed AI agents and deployed inside existing workflows.
Every agent operates within defined permissions, approval paths, and ownership rules, with human oversight, traceability, and value measurement built in.
Find the value. Deploy the agents.
For more: https://t.co/UVlbMm3QLP
@gordon_cassie The real issue is the gross margin of any AI application layer that relies on LLM inference. The valuation dynamic is interesting, but it comes down to how much of that revenue the application layer actually gets to keep
๐๐๐๐ ๐.๐ ๐๐๐๐๐๐๐๐๐ ๐๐๐๐ ๐๐๐๐๐๐๐.
The coding numbers are strong, but the knowledge work benchmarks are what matter most to me.
On GDPval-AA v2.1, Opus 5.5 scores 1846, ahead of Fable 5.1, Opus 5, and Astra.
At lower effort, Opus 5.5 seems to produce:
- fewer false positives in review tasks
- better structured financial answers
- more industry-standard slides
- more relevant citations
- better handling of hidden assumptions
- less tendency to stop at the first plausible answer
That looks promising for serious knowledge work in enterprise AI.
Testing it.
Looks like today may be a record day for token volume % of open models on Vercel AI Gateway:
๐ฆ Open 78.4% ๐จ Closed 21.6%
While spend ๐ฒ usually tells a different story, #3 and #4 today are Moonshot AI & DeepSeek. Adding Zโ .ai, their combined spend surpasses OpenAI (#2).
(Do note that's the spend for inference of the model across providers (mostly in the US), not revenue going directly to the open weight labs.)
JUST IN: Anthropic is considering releasing a new AI model to counter OpenAIโs GPT-6 Astra momentum, just days after CEO Dario Amodei called for the industry to โslow the paceโ of AI development. โ Reuters
Iโm convinced that Jev, or the next "System One model", will play a big role in AI orchestration.
When building an agent harness, you quickly realize that a lot of the work around the LLM is not open-ended reasoning, but rather bounded decision-making: supervising workers, detecting issues, routing tasks, evaluating progress, selecting tool calls, filtering context, triggering retries, escalating, etc.
Jev will not replace frontier LLMs. It will sit around them.
- #Jev: fast, deterministic decisions.
- Frontier #LLMs: expensive, intelligence work.
Thatโs the interesting split
With all the hype around #Jev, we wanted to test it.
We tested it againt gpt-5.4-mini, gpt-5.6-sol, and gpt-6-astra across few use cases: detecting real deal signals, routing user requests, selecting the right tools, deciding what to remember (memory), and workflow discovery.
We know that LLMs are an expensive way to make deterministic choices. The question is wheather Jev might change this economic.
First, the cost gap was brutal: ๐๐๐ฏ ๐ข๐ฌ ๐๐๐ฑ ๐๐ก๐๐๐ฉ๐๐ซ ๐ญ๐ก๐๐ง ๐ฆ๐ข๐ง๐ข, ๐๐๐จ๐ฎ๐ญ ๐๐๐๐ฑ ๐๐ก๐๐๐ฉ๐๐ซ ๐ญ๐ก๐๐ง ๐ฌ๐จ๐ฅ, ๐๐ง๐ ๐๐๐จ๐ฎ๐ญ ๐๐๐๐ฑ ๐๐ก๐๐๐ฉ๐๐ซ ๐ญ๐ก๐๐ง ๐๐ฌ๐ญ๐ซ๐.
But the result was more interesting than โJev is cheaper.โ
Across our tests, Jev performed particularly well on narrow, high-volume gates: deal-signal detection and workflow discovery, while GPT-5.4-mini was stronger at tool routing, and the larger models performed better on complex memory judgments.
The key takeaway: ๐๐๐ฏ ๐ฅ๐จ๐จ๐ค๐ฌ ๐ฅ๐ข๐ค๐ ๐ ๐ ๐๐ญ๐๐ค๐๐๐ฉ๐๐ซ, ๐ง๐จ๐ญ ๐ ๐ ๐๐ง๐๐ซ๐๐ฅ๐ข๐ฌ๐ญ.
Instead of asking one expensive LLM to make every decision, you can use a cheap, calibrated model to filter the massive volume of simple decisions, and reserve the expensive reasoning models for the cases that actually require them.
Cheap gate โ rich reasoning โ action.
We are definitely going to keep testing this approach.
Many private equity firms want to adopt #AI. Far fewer have a clear strategy for turning it into measurable ROI.
That is why our approach at @lampi_ai has always been use-case first.
We identify specific tasks, understand the existing deal-team workflow, and focus implementation where AI agent can make a concrete difference.
In this article, we share what that looks like in practice through use cases we have implemented with PE firms โ from back-office tasks to deal screening, due diligence, and portfolio value creation.
This is exactly how our automation already runs in production:
Typed workflow specs.
Deterministic validators.
Confidence thresholds.
Review gates.
Code owns the loop.
The llm judges the evidence.