@OneTriangleAI (YC S26) is building the fastest lightweight, cheap inference and hosting @deepseek_ai V4 Flash today at $0.15/M input and $0.35/M output tokens. Live at https://t.co/3xk6ys5T3W.
Our group of @MIT grads and olympiad medalists from @GoogleDeepMind, @JaneStreetGroup, and @NeurIPSConf /ICML publishers are setting a new standard for cost and latency efficiency using KV cache optimizations.
We beat out @nvidia's KV cache transfer for KL divergence, and we're using it to build the fastest inference @TryTrustAI at @ycombinator.
17.8% less TTFT, $641 less per 1M requests, and 82.5% held-out top-1 agreement on Qwen 32B, prefilled from Qwen 8B.
We beat out @nvidia's KV cache transfer for KL divergence, and we're using it to build the fastest inference @TryTrustAI at @ycombinator.
17.8% less TTFT, $641 less per 1M requests, and 82.5% held-out top-1 agreement on Qwen 32B, prefilled from Qwen 8B.
We’re in YC!
Here’s what we’ve been up to:
> Dropped out of MIT
> Rejected our NYC quant and banking offers
> Moved to SF
> Fixed broken AI adoption with personalized one-click automations @TryTrustAI
Excited to put everything into our product, great start to a new chapter.