GPUs are built with more memory bandwidth, but higher latency and lower capacity than CPUs. AI accelerators could usefully make a different memory trade — the bandwidth of a GPU (or more), but the capacity of a CPU, in exchange for even higher latency.
For inference, all the weights, potentially a couple TB, can be accessed in a completely linear manner. It could even be a single transaction, streaming at a constant speed into an L2 ring buffer for a GPU core to chase calculations in, akin to racing the beam on old CRT game architectures. You could build a memory system out of masses of dirt cheap RAM, fully in parallel.
Even for training, the memory access patterns can be just a forward read of the weights, a reverse read , and a staggered reverse write of gradients and weights. You could have minimum transaction sizes in the megabytes, and first byte latencies in the many microseconds.
At the new Free Play Dallas arcade, I saw a cabinet I had never seen before: Donkey Kong II: Jumpman Returns. It was super hard, so I assumed it just never got widely released, but it turns out it was a ROM hack for the original hardware: https://t.co/RAVsEJ0Pnd Well done!
I started a tar with bzip command on a big directory, and it has been running for two days. Of course, it is only using 1.07 cores out of the 128 available. The Unix pipeline tool philosophy often isn’t aligned with parallel performance.
I don’t feel it, and I don’t think it is rational, but I do see it. The zeitgeist of the thought leaders does tend negative. I remember waving my flag of technological optimism in the 90s and getting smiling chuckles from people, while now I can count on at least some angry disapproval.
On net, the world is better than it has ever been in history, and the derivative is still positive — it is going to get even better! Miraculous and clever things are all around, and I try to appreciate them all.
I want to tell people to snap out of it and get on board the progress train; there is room for everyone! To a large degree, it will happen regardless of the negative feels, but not completely. Culture does make a difference, and I would like it to be positive.
If it’s not clear yet, this is what I think will happen soon.
Every major tech company except Apple has announced their own LLM.
Apple have spent years perfecting their on-device neural engine. Capable of some absolute insane operations. Loads of compute in a small and energy efficient form factor.
With M1, M2 & soon M3 the neural engine is even more powerful than their A series mobile chipsets.
While we currently need the cloud to run ChatGPT and it’s clunky, I think Apple is going to blow everyone out of the water here. Both on desktop class hardware and mobile.
I think Apple will be launching their own secure and private LLM that runs on device (edge compute). And when necessary it offloads more heavy workloads to a cloud based LLM that’s optimized for heavier tasks. So we will initially have some hybrid.
Personal, with tight hardware and software integration this AI will be omnipresent. Apple will probably use this to sell a lot of new hardware that they claim is needed to run this. They will make a lot of moneys.
For me the LLM’s will form the new protocol level technology upon which most new software will be built. We will have to re-wire our core understanding about what an application is.
Single-use apps will be a huge thing. If you need to solve a unique problem, and nobody has ever done software for that because not enough market. With an LLM even a problem with only one user, will be doable, enter your ask, and code gets written, problem gets solved. Runtime ends, app dies. Done. Single use apps are born.
It’s hard to predict or try to understand how the world will look just 10 years from today. It will be very different, we have passed the inflection point, the rocket engines have been lit. We’ve taken off.
Add to all the above that every single field, category and market will be disrupted at the same time. And not only with text/coding but with any multi-media we have. Images, video & audio. Anything we can come up with can and will be enhanced or disrupted by AI.
Once we got more people that will have their AI A-ha moment the rate of change and adoption will continue to increase. This will continue until we have global access and coverage.
People will get left behind, and this will be one of the most important things to try to combat. Having a 0% left behind policy. We need to make sure AI benefits all.
We’re living through a paradigm shift, and we’re witness a new protocol level technology. We’re seeing it arrive in real-time and most people have no clue about what’s about to happen.
I’m not an AI alarmist, I’m an AI gardener, and optimist.
We will have time to adapt. Not as long as we had during the Industrial Revolution, but enough time to make sure we have a chance at a positive outcome.
We’re moving away from the Information Age into the Age of Intelligence. With unlimited access to intelligence anywhere, anytime.
18th Mars 2023 - Linus Ekenstam
Ok, so here is my first ever #rust post with my first ever Rust program. I hope it can give some people a nice overview of the language and how things are written and put together.
https://t.co/2JJhSxfE2j
Hi Fox Rider, When we did the Epistory plushie giveaway with Makeship, we promised you that we'll also do one on Steam too and here it is! https://t.co/FD8EZ3qcwC #gaming_news via @steam
It was great joining @reneritchie to discuss the state of silicon!
And yes, if you catch it, I meant Taiwan not China related to TSMC. Sometimes my mouth gets ahead of my brain lol.
https://t.co/jNgcuH2j6k
Sam Bankman-Fried: I don’t know where $10 billion dollars went
The Pentagon: We don’t know where $2.2 trillion dollars went
The IRS: You just sent $601.37 don’t forget to report it.
@thekangminlee I’m a fellow Korean American and Christian. I can relate to the fear but God has to be our rock. I will pray for you and I hope you keep praying.