We are releasing our first step in validating and independently confirming the claims of the Bitnet paper, a 1B model trained on the first 60B tokens of the Dolma dataset.
Comparisons made on the @weights_biases charts below are between the Bitnet implementation and a full FP16 run (all hyperparameters equivalent).
Model: https://t.co/3WQPrdBbP8
Weights & Biases: https://t.co/crm5JZqHVo
We made a craigslist for gpu clusters: https://t.co/79pYb2TeHy
Check it out if you're looking for gpus. And if you have any extra hardware you want to rent out, posting is free.
Follow @gpulist to be notified of new posts.
Tenstorrent plan update
1. open source software (Buda and Metalium) - January
2. sell Gen 1 Greyskull devkits, enable
online purchase - today
3. coming soon Gen 1 networked AI
4. power on Gen 2 in April
5. design best possible Gen 3 low cost AI - next
https://t.co/IaFvekQd6N
Probably the first operation cost analysis of owning @GroqInc hardware to run Llama2-70b.
First of all, let me say I am a big fan of Groq. Great performance, great potential. The below is just a showcase how challenging things might be when rivaling the industry lead, but given time I look forward to it.
1. Each Groq card has a memory of 230MB. For the LLaMA 70b model, assuming int8 quantization and completely disregarding the memory consumption for inference, the minimum number of cards needed is 305. In reality, more are needed, with reports indicating 572 cards, so we will calculate based on 572 cards.
2. The price of each Groq card is $20,000, therefore, the cost for purchasing 572 cards is $11.44 million. Of course, due to sales strategies and benefit of scale, the price per card might be much lower, but let's calculate using the list price for now.
3. For 572 cards, and the average power consumption per card of 185W, the total power consumption is 105.8kW excluding peripherals. (Note, actual consumption will be higher)
4. Currently, the average price per kW per month in data centers is about $200, meaning the annual electricity cost is 105.8 * 200 * 12 = $254,000.
5. Basically, using 4 H100 cards can achieve half the performance of Groq, meaning an 8-card H100 box is roughly equivalent in capability to the above. The nominal maximum power of an 8-card H100 is 10kW (actually about 8-9 kW), so the annual electricity cost is $24,000 or slightly lower.
6. Today, the price for an 8-card H100 box is about $300,000.
7. Therefore, if operated for three years, Groq's hardware purchase cost is $11.44 million, and operational cost is $762,000. For an 8-card H100 box, the hardware purchase cost is $300,000, and operational cost is $72,000 or slightly lower.
These are all rough numbers. Please kindly point out any outrageous errors I made above.
ollama run llava:34b
>>> What's in this picture? ./img_0064.jpg
This is an image of a graffiti on a wall. The graffiti features a blue whale with white fins, and the words "BUILD SHIP RUN LOVE" are written in white above it.
Project #2: LLM Visualization
So I created a web-page to visualize a small LLM, of the sort that's behind ChatGPT. Rendered in 3D, it shows all the steps to run a single token inference. (link in bio)
We’ve brought the new M3 family of chips to iMac and the new MacBook Pro lineup, and they’re now available! There’s never been a better time to experience a Mac.
[알고리즘 코딩 테스트 책 추천]
이 책에는 수많은 면접자들을 대상으로 코딩 테스트를 출제한 경험, 코딩 인터뷰를 수행한 경험, 그리고 면접을 더 잘하기 위해 수많은 회사의 코딩 면접, 기술 면접 과정을 면밀히 살펴본 경험을 담았습니다.
https://t.co/Fk4Q9HBMxX알고리즘-코딩-테스트-책-추천