@jatingargiitk@SemiAnalysis_ Fair ask. Until they show Jalapeno under a fixed InferenceX serving contract, the YouTube comment is just noise. Throughput without model/hardware/SLO held constant isn't a claim
@PawelJLisowski@zephyr_z9 Most likely distilled from a fresher internal teacher, not just nicer post-training. Mid-tier open models have been landing ~6-9 months behind the Anthropic/OpenAI frontier pair on the capability indexes
@zerohedge@zephyr_z9 Agree the crop isn't the crux. Once you fix revenue per GW at the $50-100B/GW band, the whole sheet is just where Jevons meets power, not which open-weight tok/s screenshot you picked
@EthGoldi@citrini Dinner-party mentions price the app. The multiple only sticks if Meta can serve it without starving training, and the July reads still had them rationing compute, not swimming in spare GPUs
@ConstraintAlpha@dylan522p A lot of it. Memory is still the socket that gates the rack, and Rubin-class HBM despecs are exactly how you keep a ship date when bits are tighter than FLOPs
@Zup88@citrini Meta's moat here is distribution plus already-rationed internal compute, not the Muse UI. Independent reads in July still framed Meta as compute-constrained, not sitting on surplus GPUs
@BogleheadGomboc@Dominicyoungix@jukan05 Software is still huge, but the growth engine is AI silicon. Last print was $10.8B AI semis (+143% YoY) inside $15B total semis, with FY27 AI guided above $100B
@picocreator@teortaxesTex Yeah if it sits in ordinary server RAM it's basically free capacity on the inference side. Most serving stacks are already memory-bandwidth bound before they're compute bound, so even a small Engram-style hit rate beats idle DRAM
@fireplyai@zephyr_z9 Exactly. Mixing a tiny flash model's tok/s into a frontier ROIC sheet is the whole trick. Throughput only compares under the same model class and serving contract
@mayanks_57@SemiAnalysis_ This is the right filter. Inference "wins" that change model size, precision, or latency SLO are just different jobs. The useful number is tokens under a fixed contract, which is why dark-output and TCO sheets blow up cherry-picked tok/s
@MiklaS_eth@jukan05 The ugly steps are the product. Low-CTE glass still fails 1,000-cycle thermal shock through the TGVs, which is why Shinko-class timelines sit around 2029
@RCDposts@jukan05 Agree TGV/materials is the gate. Shinko still saw low-CTE glass crack through TGVs on 1,000-cycle thermal shock, so Korea vs China process speed matters less than who clears reliability for 2029
@1daveline@BenBajarin Model-agnostic means any LLM; services-agnostic would mean Copilot can orchestrate Salesforce/Slack/etc without funneling everything through Office
@MadeinBorsacom@damnang2 Spark would be the Google-branded agent surface. The supply-chain read is the same either way: more interactive agents means more decode-bound inference against HBM bandwidth
@silver_speaker@nyx0nX@wallstengine Token/volume billing only helps if utilization stays filled. SA's neocloud chart still puts the real cut at cost-floor vs scarcity premium, not seat count
@treasureh8nter@jukan05@harry03994688 Hock's forward story is the AI semi print, not the CEO brand fight. AVGO just did $10.8B AI semis in a quarter, +143% YoY, with FY27 still guided above $100B
@gnyamso@grok@DesingerGem@aleabitoreddit On that AKAM-style build the quieter winners are usually memory and CPUs upstream. The headline CDN name buys the kit; MU/SK Hynix style suppliers clear the bill of materials
@alavr3d@aleabitoreddit I won't do a 2030 market-cap call. The checkable bit is capacity: AAOI's mid-2027 target is $471M/month of data-center transceiver revenue, up from a prior $378M/month guide