@SemiAnalysis_@dylan522p am I reading this right? if the pro 500 plan for OAI is a 150% percent jump in pricing but the true additional token limit increase is only around
20% for GPT 6 Astra, that should be positive sign for their subscription margins right but not the most efficient subscription for the user lol
Also Claude subscription only being 10% of the revenue and utilizing 42% of the compute sounds scary, at this rate I would believe they will have to do a massive overhaul of the subscription model
@Futurenvesting
Seeing so much hate on $SOFI on X recently, I kind of want to enter into a position again. Guys chill out , I know we have companies selling chips doing 100%+ growth but sofi is doing pretty damn good given the industry they are in
For my simple brain @RealMattMoney
am I thinking this right?
$NEXT bull case: long-term LNG contracts, first production targeted for H1 2027 and substantial future cash flow.
Bear case: costly debt, partial project ownership and dilution risk.
Dropping tomorrow at 8:30a EST!
The Most Asymmetric Bet in the Stock Market Right Now | $NEXT NextDecade
Let me know what you think.
https://t.co/q58jVeNyi8
@pequityresearch we will soon see that 2028 is also sold out and discussions will start for 2029, supply can only be increased at ceiling every year, hard to build out these fabs quickly, I would still expect the pricing power to be extremely strong going into 2028 as well
@KrisPatel99 I heard you some weeks ago on @Futurenvesting channel about the Coreweave’s renegotiated contracts till 2029 might not mean anything, although no one knows what happens in the future, I would say that is not true entirely, there is an entire tail end of GPU cases that are there, you do not need Vera Ruben for everything and also many companies can’t even afford the latest chips especially in a supply crunch, Amazon sagemaker still supports a ton of old Nvidia GPUs that are being used everyday by Data Scientist and Machine learning engineers
Everyone who is worried about the entire GPU depreciation thing needs to read this, also I can tell you the notebook instances I use everyday for machine learning work still has v100 attached to them, that was GPU what was launched in 2017, and they are still being used for training and building models. AWS clusters still support them. There are huge tail end cases for GPUs that people might not know off
Ever wonder why truly from a technical standpoint $NBIS keeps raising Hopper H100/H200 prices and robinhood:0x5f10a1c971b69e47e059e1dc91901b59b3fb49c3 lands massive multi year A100 deals while wall street is still obsessed about GPU depreciation? FinX and Wall Street analysts seems to have 0 understanding on the fact maybe because they focus on the bigger picture rather than the engineering granularities.
Look at the attached https://t.co/KX6nYQiL4j chart: These are 4 year old RTX 3090s surging ~25%+ in rental price over the last month. Why? Because this consumer GPU offer some of the highest VRAM per dollar in the entire $NVDA ecosystem.
Here is the 3 core technical reasons Wall Street and FinX keeps missing:
1/ Inference is memory bound, not compute bound. It’s VRAM and HBM density per dollar that dictates real world deployment cost, not peak TFLOPS. TFLOPS matter, but in this demand cycle, fitting the model weights and KV Cache to simply get the model running matters far more.
2/ The enterprise agentic boom runs on sub 70B models. If you’re running a 30B model in FP8 for non-tech automation, agentic networks, or internal tools with ~50 concurrent users, the total KV cache and weights footprint only requires 30GB to 60GB of VRAM.
3/ Serving engines (vLLM / SGLang) favor modular GPU blocks. Because engines like vLLM and SGLang serve one base model per process, deploying diverse agentic pipelines (separate routing, coding, and extraction models) requires isolated container nodes. Slicing up a massive Blackwell cluster for multiple separate engine instances introduces useless orchestration overhead and wastes Blackwell’s NVLink scale out capabilities.
Spinning up ultra high bandwidth Blackwell (B200/GB200) clusters for these workloads is extreme hardware overprovisioning in the above cases. Enterprises don't need raw compute flex they need cost per token efficiency.
A100s and H100s deliver the exact granular, segmented VRAM blocks required to serve real world models without burning cash.
"Legacy" GPUs aren't aging out. The are and will be the workhorses of everyday enterprise AI. Until structural supply constraints across the broader compute economy end, pricing power for these segmented GPUs isn't going anywhere.
The reason you wouldn’t see this type of analysis elsewhere because the crowd that’s aware of this is not conjectured with the financial markets or the finance guys the other way around.
End of the day as someone in the conjunction I will end with this. Bullish $NBIS $CRWV $NVDA and $MU NFA
@pequityresearch I just heard that their CEO said if their was a new diamond tier in ClusterMaX rankings , their principal architect would work himself to death to achieve that status, I would be bullish for a company like that🤣
@Futurenvesting@Dr_Crossroads
very bullish $AVGO, the efficiency gains from custom asics might not be looking good today compared to Nvidia, but every frontier lab is going to try to get away from Nvidia premium, also shifting to asics does not mean accepting lower performance, hardware and software co-design will deliver a lot of gains
@ScrooogeUncle@amitisinvesting@stevenfiorillo Fair point, but less memory per query doesn’t necessarily mean less memory demand overall. Cheaper inference could mean way more usage. For $MU, the question is whether demand grows faster than the efficiency gains.
Okay so memory is cyclical but then HPE, DELL and SMCI are not? How is mr market giving different multiples to these if the underlying driver is just capex? explain me like I am five
@stevenfiorillo@amitisinvesting