Offline Base is excited to announce that we have been accepted into the @nvidia Inception program. We’re excited to continue building with edge devices and open source models from Nvidia to help build local AI products for all. Thank you @NVIDIAAI!
We look forward to building with models from @googlegemma@PrismML@liquidai Nemotron Models and more!!
AI is driving US inflation higher:
Consumer prices for computer software and accessories surged +14.5% YoY in May, the biggest annual increase on record in data going back to 2000.
Producer prices for electronic components soared +27% YoY, also the biggest increase on record.
To put this into perspective, before 2026, prices for software and electronic components fell in almost every year since 2000.
Memory prices alone have more than doubled, with DDR5 and DDR4 RAM prices up +290% YoY, as AI data centers absorb the vast majority of global chip supply.
RAM price shocks are will likely keep inflation elevated well into 2027, adding to existing pressures from the Iran War.
The AI boom is fueling technology inflation.
In light of the @GroqInc and @nvidia news, I am going to start posting more publicly about a personal project I am working on: Augment - https://t.co/S3y4XSTT22
Here's a demo of me making a website on Augment using @Kimi_Moonshot running on @GroqInc .
Less than 10 seconds to generate.
I made a very simple website for the sake of the example.
I think current vibe coding tools are a bit too slow to iterate comfortably given the fact that it may or may not do what you want, but you have to wait ~ 2 minutes to find out (and potentially spend money on tokens that wound up making it worse).
I also think deep research queries take a little too long for most folks to casually use. Shoutout to @perplexity_ai for how well they do around this.
You can use open source models on @GroqInc , @cerebras , and LLMs from the major labs as well.
Working on a small project to demonstrate the value of faster inference https://t.co/S3y4XSTT22
Thesis:
Frontier models have pretty fast inference, though not fast enough for test-time compute/reasoning and 10+ tool calls.
If we are to augment with AI and save time, let's not create a new type of down time... waiting for responses.
We used Cerebras as our inference provider to create websites, analysis, and textbook chapters within 2 seconds.
We typically hit around 2000 tokens/s.
Details:
Using Cerebras and Groq as inference providers.
Models - GLM 4.6, GPT-OSS-120B, Qwen...
Horrible demo video below. Only the typing is sped up.
I am still working on this.
I just wanted to start posting updates for posterity because @GroqInc and @IBM partnered after I started this and I don't want to miss anymore news regarding the benefit of fast inference when it comes to tool-calling, Reasoning, Enterprise, and healthcare need for speed.
In 24 hours if:
Sol hits 180
Eth hits 3,700
Bitcoin hits 125,000
I’ll send one of my followers that likes this post one of each.
Total: $128,880 - Sent right to your wallet.
I swear to god.