Recently left my role as a L/S Tech Analyst to go build @InfoFloHQ. I have a vision of the future where most smart investors will be using voice AI agents to go do diligence on their behalf by tapping into the “tacit” knowledge in the heads of industry professionals. I think this tacit knowledge is only going to become more valuable as AI commoditizes all publicly available data. The pitch with InfoFlo is we use voice AI agents to help you due diligence so you can build more conviction, move faster, and spend less.
Reading the Infinity Machine you see how much Demis dedicates his life to getting to AGI, must be incredibly frustrating to be leapfrogged by Dario, Sam, and now Mark.
$GOOGL'S GEMINI 3.5 PRO IS REPORTEDLY MONTHS LATE
Bloomberg says Google delayed the model while trying to improve its coding capabilities, as OpenAI and Anthropic continue to advance.
The smart CEOs know just the threat of insourcing is enough to get a discount at renewal, even if they never intend to actually do so. Especially for public software cos where they know any mention of uptick in churn is going to permanently handicap their multiple, better off just passing along the discount while still retaining the customer.
Flexport is already pulling roughly 20% out of its SaaS vendors at renewal by threatening to rebuild the tool in-house.
Ryan Petersen (@typesfast), Flexport's founder and CEO, says the next negotiation with Salesforce is going to look nothing like the last one. He has a procurement team building PowerPoint case studies of the SaaS tools Flexport already ripped out and how fast it did it.
The pitch to every remaining vendor is simple. Cut your contract rate or we vibe-replace you. He thinks 20% comes out of almost everybody.
He's honest that some of it is a bluff. He doesn't want to burn engineering talent rebuilding software instead of building product, so he can't do it for all of them. But the leverage is real and it's in procurement conversations today.
Selling SaaS to companies that can build it themselves is about to be a tough business.
Key takeaways already in your email via @PodWireHQ
Source: 20VC with @HarryStebbings
Sharing some of our monthly learnings below. Our early hedge funds clients find these insightful given we are deploying frontier AI in production faster than most.
Building AI Expert Network: January 2026 Learnings
“There are decades where nothing happens, and there are weeks where decades happen” – Vladimir Lenin
Pushing Voice Agent to its Limits
- We tested the Interviewer Agent for hours to figure out where follow-up question ability starts to degrade. We discovered there is an optimal balance between total context window size and recency-weighted attention, and the resulting ability of the Interview Agent to perform with relevant follow-ups
- After hundreds of tests we landed on a unique orchestration, allowing for compressed context and a reset attention budget which is producing more consistent follow-up quality throughout long interviews
Claude Cowork/Opus 4.5 for Admin Work
- Given our simple 2025 P&L, I had Claude take the first pass at preparing our Form 1120/Form 5472 for US tax returns. It made 4 small mistakes (PDF readability issues, not reasoning errors). After some back-and-forth and telling it to triple-check its work and think harder, it got to the correct version – saved us a few thousand dollars
- We’re using Claude to generate most of our marketing/sales enablement material. I provide a rough outline, it generates the material in our brand colors, and I iterate with edits. This is faster than coordinating with a freelancer and the output quality is higher (I continue to think ADBE and FVRR are shorts)
Future of Software Engineering
- My technical co-founder has shifted to mostly reviewing code since Opus 4.5 was released. Agents do the writing, he does the judgment calls and review
- We tested the limits of how far we could push this approach by having agents pick up our Linear tickets and make changes directly to the codebase. We reversed course because the agents went off the rails too often. We’ll continue to push the frontier until this is possible
Rewriting System Prompts
- Given the underlying model capabilities improved a lot over the last 8 weeks and inspired by Dario’s recent essay “The Adolescence of Technology”, we decided to rewrite some of our core system prompts to use high-level principles with rich reasoning instead of with a rulebook with checklists – like how Anthropic wrote Claude’s constitution
- The result was higher scores across our evals while reducing system prompt size by >50%. Earlier models needed well-defined guardrails, these newer models reason from principle much better
From “Vibe Evals” to Systematic Evals
- Moved from evaluating prompt changes by gut feel to building a full evaluation system on Langfuse (amazing platform)
- We now use synthetic datasets with different client personas, LLM-as-a-Judge scoring, and prompt version controls so we can now effectively test how minor changes in our system prompt impact output quality
Misc.
- Both a16z's "State of Markets" and Avenir's "Future of SaaS" decks released recently – worth flipping through. The a16z deck has good data on private AI infra/app companies
Moving from vibe-based prompt evals to a real production system with datasets and systematic scoring is like putting on glasses for the first time. Happy to retire "this feels better" as a metric.
Older cohort data suffers from the same
problem that eventually that you switch to the newer model. But if you stay within the same family of models (Anthropic) then it doesn’t matter as much. Agreed that the constant need of spend to developed the newest model and how you allocate for that in COGS vs R&D is a question mark
Thoughts on recent SaaS narratives from a former hedge fund software analyst now building an AI software company
- From an analyst/investor seat, it’s very easy to say “you can vibe-code anything now.” That is a very surface level take and there are a lot of nuances underneath. Once you’re actually building software for enterprise customers, you quickly realize how much of the work sits outside writing code; security reviews, compliance, data integrity, edge cases, debugging production issues, and long-tail requests. Coding was always the easy part but planning and building the system to produce a good product is where the bottleneck is.
- Building to feature parity is very different from building a company. You’re still selling to humans, and history is full of software companies that won not because they had the best product, but because they had better distribution. Enterprise buying still means navigating budgets, security and compliance reviews, procurement, and multiple stakeholders with different incentives.
- When a core input to building something gets significantly cheaper, you tend to do a lot more of it. The closest historical parallel to what’s happening with coding today is spreadsheets and financial analysis. Excel didn’t eliminate analysts, it expanded the amount of analysis done. The role evolved, expectations changed, and more people could contribute. The definition of a software engineer will change, but many more people will be shipping code.
There are pockets of software I wouldn’t touch, specifically where the underlying workflow is being fundamentally reshaped by AI and can be done better, faster, and cheaper by agents ( $ADBE). When the core job shifts from “help me do X” to “do X for me”, the newer age AI companies are in a better position to win.
I'm very bullish on vertical software. These companies win on domain expertise and that knowledge compounds. They're selling outcomes, and AI expands the set of outcomes they can deliver (look at $IOT and $APPF AI products).
Spent a lot of the weekend thinking about the future ramifications of AI Inbox/Overviews in Gmail. Gmail has >70% market share in the U.S. for personal email and >40% for business (skewed SMB), if this becomes the preferred way of people interacting with their inbox that likely means cold outbound conversion structurally declines. Big ramifications for a lot of B2B, software, and email marketing driven industries that are going to have to figure out how to reach customers. This is on top of AI personalized-outbound already starting to clog all outbound channels.
Claude’s browser extension is definitely a step-up vs OpenAI’s Operator for computer vision and being able to navigate around web pages. We’re testing and deploying agents internally which use computer vision to automate some parts of our workflow. It’s still finnicky, but fascinating what it can already do. If the Q1’26 LLMs trained on Blackwell chips are indeed a step-function improvement in computer vision, there are a lot of workflows internally we can agentify. TAM for probabilistic workflows is infinitely larger than deterministic workflows.
@JaredSleeper % split of successfully deployed agents built internally vs purchased from 3P vendors
% of deployed agents with a human approval gate vs. fully autonomous
8/ Yes, I realize I left out a bunch of material with the old vs new CFO, Non-Atlas accounting vs Atlas, GTM changes which started paying off, etc., but wanted to just get the main points across!
1/ Thought I would share a fun case study on one of my biggest PnL wins at my prior L/S hedge fund seat, how it led to the idea to go build @InfoFloHQ, and the market structure creating these outcomes.
MDB is a NoSQL database provider that I had covered for years. The stock dropped more than 25% on Q4’25 earnings and bottomed in Apr 2025 after a ~45% drawdown. The crux of the issue was they materially guided below street for FY’26, and investors believed the bear case of losing share to PostgreSQL (an open-source database provider) was playing out. In tech, once you’re deemed a legacy vendor on the losing side of a platform shift, the multiple investors assign de-rates materially.
7/ Conclusion/InfoFlo: While I was doing the >40 expert calls, the one thing that kept bugging me was all the friction in the process with the scheduling, prepping, and conducting each call over the span of weeks. I was paying very close attention to what was happening in the AI application layer and had the thought of building a platform where investors scope out a diligence project, the platform sources the most relevant industry experts, and then voice AI agents do the expert calls and come back with insights. This idea led me down a rabbit hole and I eventually left my seat to go and build @InfoFloHQ with the goal of helping investors reduce friction to insight, build more conviction, and move faster. I will continue sharing my thoughts/learnings building the business.