@StevenBartlett@saylor Must listen for founders and investors. We weigh this same conviction when backing LATAM founders. Technology, capital and career mobility are one bet. What separated his early infrastructure bets from the rest?
Generative AI compresses due diligence and structuring cycles in our own deals. Saylor's $15B is proof at scale. Are you using these tools with your lawyers and bankers yet?
@saylor used ChatGPT to make $15 Billion last year? 🤯
He says he used ChatGPT alongside lawyers, bankers and investors to help solve complex problems and design new financial products, resulting in $15 billion of credit being sold.
#bitcoin
Another legal scholar makes clear why EU Commission should propose a ban on trade with illegal Israeli settlements as a matter of living up to EU's foundational principles. By Nicolas Angelet, professor of international law https://t.co/rz5fTiUqTq
MOST RAG EVALUATION TOOLS TELL YOU WHAT BROKE. THIS PAPER TELLS YOU WHY AND HOW TO FIX IT
RAGXplain (arXiv:2505.13538) introduces a framework that does what every eval tool promised but never shipped: it translates quantitative metrics into actionable guidance.
### The Metric Diamond
Evaluation structured around 6 diagnostic dimensions connecting user input, retrieved context, generated answer, and ground truth. not just a score. a map of where the pipeline failed.
### 3-Stage Pipeline
1. **metric calculation** - LLM-as-judge approach, customizable metric suite
2. **insight generation** - synthesizes patterns across metrics into human-readable narratives
3. **action item recommendation** - prioritized, practical recommendations for system enhancement
### What Makes It Different
- **explainable metrics** - transforms abstract scores into detailed explanations of specific weaknesses
- **prioritized fixes** - identifies whether the failure is in retrieval or generation, and suggests what to change
- **multi-level** - works at instance level AND dataset level, unlike most frameworks that only do one-shot per example
- **LLM-agnostic** - compatible with various LLMs as judges
### Results
- strong alignment with human judgments across public QA datasets
- applying RAGXplain's recommendations measurably improved RAG performance
- bridges the "trust gap" - users understand not just what the system output but why it failed
this matters because most teams use aggregate Apache scores without knowing WHERE the bottleneck is. is it retrieval precision? generation hallucination? context contamination? a single number does not answer that. RAGXplain does.
the article covers the production engineering behind this: why golden sets under-represent real users, why BLEU and ROUGE are dead for LLM evaluation, why hybrid judge (rule-based + human + cross-family LLM) beats single-judge, and why evals-first teams move 3-5x faster on prompt changes
@milesdeutscher Models are improving faster than the infrastructure. I would watch permitting velocity and unit economics before calling this the era of robotics. Where do you see the first commercially viable margins?
Sam Altman on why an entire company's worth of work from the early YC days now takes seven minutes:
To understand how radical this shift is, Altman first takes you back to what startups actually looked like twenty years ago.
"If it were possible to get as far from this moment as I can imagine, it was: startups are not cool at all. We were like hiding out in this little building in Cambridge. PG was making us dinner."
Just a small group of founders grinding in obscurity, too broke and too heads-down to do much else.
Here's the number that reframes everything: what took three months to build back then, the kind of product any single company built over the course of the entire YC program, can now be done in about seven minutes by a coding agent.
Not seven months. Seven minutes. That gap spans just twenty years.
@sama says there are two ways founders can respond to that:
"You could either be sad and be like, 'Oh man, a Codex prompt is a whole startup,' or you could be like, 'I can go start the world's most ambitious crazy company. I can have experts in every field working together. I can do these very hard technological things that were just impossible.'"
His verdict is unequivocal: "It's going to be amazing."
But he refuses to romanticize the old days just because the tools were primitive. Back then, he says, "it was very difficult to get anything to work." There was no sense that any of it was building toward something historic.
And he gives credit where it's due for the ecosystem that made this possible in the first place:
"I think Paul Graham is probably the most important force in startups of the last few decades and no question."
Shiller is right. In our deals, the term sheet aligns or exploits. Founders focus on valuation. Structure is where risk lives. What terms shift incentives most?
A Yale economist opened his financial markets course by saying the quiet part Wall Street avoids.
Finance is a technology, for good or evil.
Robert Shiller is standing alone on a huge stage, the words glowing behind him like a warning label.
Not “how to get rich.”
Not “how to beat the market.”
A technology.
That is the frame most people never get. Finance is not just money moving around. It is a machine for moving risk, promises, time, fear, trust, and power between people who usually do not understand the same side of the trade.
Used well, it funds homes, businesses, insurance, retirement, medicine, cities.
Used badly, it packages confusion and sells it to the person least able to survive the downside.
That is why the same instrument can build a hospital or blow up a pension fund.
The object is not good or evil by itself.
The design, incentives, and user decide what it becomes.
Shiller put that warning at the start of the course.
Most people skip straight to the part where money is made.
That is usually where the damage starts.
@caspr_exe This is immediate for us in UAE and LATAM. We are building AI infrastructure, but frontier capital and model weights remain concentrated elsewhere.
BREAKING: In a rare prosecution of settler violence, Yinon Levi will face manslaughter charges for the killing of activist Awdah Hathaleen. Maya Rosen reports for @JewishCurrents
https://t.co/5YZv6eWQ0m
🤝 UAE builds first Humanitarian Index — TRENDS Barometer & ICRC sign MoU in Abu Dhabi Aug 5. Measures public awareness of international humanitarian law, access to health/education/food/housing in conflict zones & displaced-people data. Uses geospatial analysis & AI. Signed by Bernasconi & Ali Abdullah Al Ali.
Read: https://t.co/2hyGuzXo15
Financial independence is often about flexibility rather than escape. Options can reduce pressure and improve decisions.
#FinancialPlanning#PersonalFinance
@shawnchauhan1 We see the same trap in our portfolio. Feeding closed models proprietary data slowly transfers competitive edge. Abu Dhabi built sovereign compute to prevent exactly that.
@Veltrxai This shows up in our pitch meetings every week. Deals that split our partnership deserve real attention. When everyone agrees, the window has usually closed.
@pdhsu My bet is experimental speed stays the hard floor. Models can't replace the bench. Someone still has to culture the cells. Where do you see the first chance to compress that cycle?