After a recent price reduction by OpenAI, GPT-4o tokens now cost $4 per million tokens (using a blended rate that assumes 80% input and 20% output tokens). GPT-4 cost $36 per million tokens at its initial release in March 2023. This price reduction over 17 months corresponds to about a 79% drop in price per year. (4/36 = (1 - p)^{17/12})
As you can see, token prices are falling rapidly! One force that’s driving prices down is the release of open weights models such as Llama 3.1. If API providers, including startups Anyscale, Fireworks, Together AI, and some large cloud companies, do not have to worry about recouping the cost of developing a model, they can compete directly on price and a few other factors such as speed.
Further, hardware innovations by companies such as Groq (a leading player in fast token generation), Samba Nova (which serves Llama 3.1 405B tokens at an impressive 114 tokens per second), and wafer-scale computation startup Cerebras (which just announced a new offering this week), as well as the semiconductor giants NVIDIA, AMD, Intel, and Qualcomm, will drive further price cuts.
When building applications, I find it useful to design to where the technology is going rather than only where it has been. Based on the technology roadmaps of multiple software and hardware companies — which include improved semiconductors, smaller models, and algorithmic innovation in inference architectures — I’m confident that token prices will continue to fall rapidly.
This means that even if you build an agentic workload that isn’t entirely economical, falling token prices might make it economical at some point. As I wrote previously, being able to process many tokens is particularly important for agentic workloads, which must call a model many times before generating a result. Further, even agentic workloads are already quite affordable for many applications. Let's say you build an application to assist a human worker, and it uses 100 tokens per second continuously: At $4/million tokens, you'd be spending only $1.44/hour – which is significantly lower than the minimum wage in the U.S. and many other countries.
So how can AI companies prepare?
- First, I continue to hear from teams that are surprised to find out how cheap LLM usage is when they actually work through cost calculations. For many applications, it isn’t worth too much effort to optimize the cost. So first and foremost, I advise teams to focus on building a useful application rather than on optimizing LLM costs.
- Second, even if an application is marginally too expensive to run today, it may be worth deploying in anticipation of lower prices.
- Finally, as new models get released, it might be worthwhile to periodically examine an application to decide whether to switch to a new model either from the same provider (such as switching from GPT-4 to the latest GPT-4o-2024-08-06) or a different provider, to take advantage of falling prices and/or increased capabilities.
Because multiple providers now host Llama 3.1 and other open-weight models, if you use one of these models, it might be possible to switch between providers without too much testing (though implementation details — specifically quantization, does mean that different offerings of the model do differ in performance). When switching between models, unfortunately, a major barrier is still the difficulty of implementing evals, so carrying out regression testing to make sure your application will still perform after you swap in a new model can be challenging. However, as the science of carrying out evals improves, I’m optimistic that this will become easier.
[Original text (with links): https://t.co/txk7q32EXn ]
@rseroter@googlecloud They say that the only one that doesn't make mistakes is the one who does nothing. Great example of how R&D is fueling the market needs and vice-versa.
The global ocean is at record high surface ocean temperatures
Large parts of the tropical ocean are the warmest ever and the global extra-polar ocean has spent the last SIX weeks above 21ºC
Check out our new global temperature tool: https://t.co/znApXolt0A
#ClimateAction
🔥 TypeScript Tip #3 🔥
TypeScript's string interpolation powers are incredible, especially since 4.1. Add some utilities from ts-toolbelt, and you've got a stew going.
Here, we decode some URL search params AT THE TYPE LEVEL.
Today Apple, Bocoup, Google, Igalia, Microsoft, & Mozilla are announcing Interop 2022, a collaboration to improve interoperability of specific web technologies. Progress is scored from the automated testing of 15 focus areas, plus 3 investigation projects. https://t.co/8tzahaz1T3
Od 21 na terenie 🇵🇱 obowiązuje stopień alarmowy CHARLIE-CRP. Ma to związek z sytuacją na Ukrainie i działaniami, które obserwujemy w polskiej cyberprzestrzeni. Do 4 marca obowiązywać będą podwyższone wymogi dla instytucji utrzymujących kluczowe systemy IT.
#cdCon 2021 (June 23-24) will gather thousands of #ContinuousDelivery and #DevOps leaders, industry icons, practitioners, and open source developers from all over the globe. Join us there! https://t.co/MQsnDpuBI0
#MF przygotowuje wprowadzenie nowego sposobu poboru opłat za przejazd #autostrada.
Przejazd odcinkami Konin-Stryków (A2) i Bielany Wrocławskie-Sośnica (A4) bez zatrzymywania się przed szlabanem będzie możliwy od IV kw. 2021 r.
Szczegóły w komunikacie⤵️
https://t.co/FAwxFCWe9p
Join the CD Foundation for a two-day virtual event focused on improving the world’s capacity to deliver software with security & speed. Register now for #cdCON 2020 Virtual, Oct. 7 & 8. @CDeliveryFdn https://t.co/UES2jrsrZP
cdCON is on! Registration is officially opened. We hope you join us from the comfort of your fav spot on Oct. 7 & 8, 2020! Register now: https://t.co/wnci1zt3Hd
We won! Organizations including @Mozilla and @EFF took our signatures to @Zoom_us, calling on it to make #e2ee available to all, and it listened: https://t.co/L9Qg0uUfth
We won! Organizations including @Mozilla and @EFF took our signatures to @Zoom_us, calling on it to make #e2ee available to all, and it listened: https://t.co/M0wFJjoEge
We've just fixed Jenkins+Jira configuration feature that can help in better configuration management for organisations, feel free to test it out in latest Jenkins jira-plugin!
https://t.co/lgnKD7Z3jP
#jenkins#jira