KCEX запускає «Схожі K-Line»!
BTC/ETH на 1г, 4г, 8г, 1д графіках. Історія допоможе знайти тренди.
Реєструйтесь з 41U2V1, низькі комісії та бонуси!
https://t.co/hSDAs3bfOU
@0xCortexx "Same pattern with every AI benchmark claim honestly. '95% accuracy' means nothing until you know what the other 5% looks like — is it edge cases or is it half the real-world distribution?"
"AI will contribute $15.7 trillion to the global economy by 2030."
Every major projection like this models the output layer. The model, the application, the productivity gain.
None of them model the input layer.
No annotation labor costs. No data acquisition budgets. No synthetic data pipeline infrastructure. No quality verification workforce. No storage and versioning overhead at training scale.
The economic analysis is measuring what comes out of the machine while treating everything that goes in as free.
That is not an oversight. It is a structural blind spot in how economists are trained to look at software businesses. When the marginal cost of copying software is near zero, you model the output. But AI is not that kind of software. It is a data-intensive manufacturing process with real upstream costs that compound at scale.
The DATA supply chain has its own economics. And right now it is almost entirely invisible in the projections that get quoted in earnings calls and policy papers.
"AI will automate 300 million jobs." Goldman Sachs, 2023.
That number gets cited in every boardroom deck. What it skips: who is building the training pipelines for the specialized domains where those jobs actually live.
Medical coding. Legal document review. Financial compliance. The annotation infrastructure for those verticals barely exists at production scale.
You cannot automate a job your model has never seen labeled correctly. The $7 trillion GDP upside assumes the DATA problem is already solved.
It isn't. Most enterprise stacks are years away from a labeled dataset that's actually production-grade in their core domain. The economic model is running ahead of the pipeline reality by a wide margin.
"The AI cost curve is bending. Inference is getting cheaper every month."
This gets repeated constantly. It is true. It is also the wrong thing to track.
Inference costs are an engineering problem. You optimize them once, and they stay optimized. The cost that does NOT bend on the same curve: acquiring, cleaning, and maintaining the training data that makes a model worth running in the first place.
A 50% drop in token pricing tells you nothing about why one model outperforms another on a specific domain task. That answer is upstream. Always.
The labs spending on annotation pipelines, data contracts, and quality review right now are buying capability that will not show up in a public benchmark for another 12 months. That is the actual capex story in AI.
Inference pricing is the income statement. Data infrastructure is the balance sheet.
Everyone is analyzing the wrong line item.
"AI will add $15.7 trillion to global GDP by 2030."
That figure is from a PwC model published in 2017. Before GPT-3 existed. Before anyone ran inference at enterprise scale and saw the actual cost structure.
The projection borrows adoption curves from previous tech waves. It does not account for:
1. Annotation labor capacity as a hard ceiling on fine-tuning throughput
2. Data quality degradation when pipelines scale past a few hundred annotators
3. The gap between "pilot" and production deployment that most enterprises are still stuck in
The economic upside from AI is real. The timeline assumes problems that have not been solved yet.
What's actually constraining the curve isn't model capability. It's the rate at which verified human feedback can be collected, cleaned, and made usable at scale.
That part never shows up in the GDP models.
"The cost of intelligence is falling to zero."
That's a compute claim. Inference per token is down roughly 99% over two years. True.
But the cost of DATA selection isn't following that curve.
Knowing which 40 billion tokens belong in a training corpus, which instruction-tuning examples are actually good, which human raters have consistent quality across a 6-month labeling contract, that scales with expertise, not with Moore's law.
The cheap part gets cheaper. The bottleneck moves upstream.
Every serious lab is spending MORE on data curation pipelines year over year, not less. The teams doing quality verification at scale are not getting smaller. The tooling for provenance tracking, deduplication, and domain filtering is not getting simpler.
AI economics are not uniform across the stack. They compress at the inference layer and harden at the data layer. If your model for this industry only looks at GPU bills, you are reading the wrong line item.
"The cost of intelligence is approaching zero. Every business will benefit."
Running inference: yes, cheaper every quarter.
Acquiring clean, domain-specific labeled data at scale: not cheaper. Not even close.
The businesses celebrating falling API costs are optimizing the wrong variable. OpenAI dropping prices does not change what it costs to build a proprietary CORPUS. Those are two separate markets, and only one of them is commoditizing.
The moat was never the model. It was the data that made the model worth running.
Everyone tracks inference cost per token. OpenAI cuts prices, Anthropic follows, the graph goes down and to the right.
That's the visible layer of AI economics. Here's what the pricing pages don't show:
1. Expert human raters. A credentialed domain specialist doing RLHF at quality costs $50-$150/hour. You need thousands of hours per capability domain, per training run.
2. Rejection sampling overhead. For every example that clears quality filters, somewhere between 4 and 10 get discarded. You pay for the rejects.
3. Structured red-teaming before release. Not a one-time audit. Iterative, adversarial, human-intensive. It doesn't compress.
Inference cost falls because it's mostly hardware and software optimization. Data labor cost is sticky because it's human expertise at a specific level of domain depth that can't be arbitraged away.
The labs understand this arithmetic. It's the actual reason synthetic data investment is accelerating: not primarily to improve model quality, but to reduce exposure to the one input that doesn't obey a learning curve.
The public AI economics story is about chips and tokens. The private one is about finding enough people who can tell a good reasoning trace from a bad one, at scale, reliably. That problem has no clean solution yet.
@TheFutureIRL "24/7 operation" is doing a lot of work in that math. Batteries need charging, joints need maintenance, and most of these systems are still supervised, not autonomous. Real capacity right now is a fraction of theoretical capacity, that gap is the entire story of this industry.
@0xCortexx Genuine question: if this really is robotics' "GPT-1 phase," what's the equivalent of the internet-scale text data that made LLMs suddenly work? Robots don't have a trillion-token dataset sitting around waiting to be scraped.
@0xFinch1 My number: 3-5 years for narrow tasks like this (pouring, rebar, repetitive site work), 10+ for anything requiring judgment calls on an unpredictable site. The gap between "can do a task" and "can run a site" is where all the actual time goes.