@iamgingertrash Is it inaccurate to say financial institutions will buy the dip faster than during 1929/2000 bursts, expecting an earlier/greater rebound ?
(Deeper markets + stronger conviction?!)
@ctjlewis i'll try to remember this moment in time in like 2-3 years when OAI signs the deal to pay spacexai billions of dollars per month to use their orbital compute
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching.
Better Defaults (not Usage Caps) โ Engineers can choose any model they want, but defaults matter. Weโre experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work.
Better Routing โ In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task.
Better Caching โ Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so weโre reusing a warm cache wherever possible. For example, our cache hit rate went from 5% โ 60% in LibreChat once properly implemented.
Keep Context Lean โ Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted.
Better Visibility โ Our engineers can use as many tokens as they want, from whatever model they want, but weโve made usage visible โ and the more you spend on AI, the more impact we expect.
The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable.
Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.
@julien_c Big tech corps announce 10x the same fundraise & partnerships: โจ๐ฅนโจ
France hosts the only global investors national summit in Europe: โ๐คฌโ
@thewhiteboxAI@Chrisgpt@DavidSacks Reliability is a system design issue.
You can create lots of checking loops, harnesses, or fast human in the loop reviews that fix it for the most part
@Element_82@justalexoki@VerbumEng Finite number of vulnerabilities at given level of abstraction/connectors.
Software tends to get more complex over time no?
@huskies20001 @CommodMkt@HedgieMarkets Twice a week one tweet blows up then 95% of TPOT masturbates to it.
Nonetheless on this instance there is truth to the fact that pricing bouta blow up