CUDA Tile has shipped! You can now `pip install cuda-tile`. I'm excited to see what y'all will build with it!
Docs & resources:
https://t.co/COaPkL7ZIV
GitHub:
https://t.co/XWmNmrUi8F
introducing gpuup: you no longer have to put any effort into setting up CUDA toolkit + drivers on a node (single or multi gpu). just copy paste a short command (in replies)
Perplexity has definitely changed the way I conduct search related tasks as part of my day to day research work AND my personal life. It has fundamentally altered how I leverage the internet. Definitely looking forward to Comet! @perplexity_ai@AravSrinivas
@Zuby_Tech Hmm… Nvidia always has a one to one ratio of RT cores to SM cores. I wonder if this is yield related but I don’t see how current Nvidia RT paradigms adapt to this unless it’s a typo. RT cores would have to be 12 to since their RT architecture has one RT core per SM.
@highyieldYT@Dee_Batch Also, Nvidia always has a one to one ratio of RT cores to SM cores (since Turing). I wonder if this is yield related but I don’t see how current Nvidia RT paradigms adapt to this unless it’s a typo.
@nholzschuch@carnets_jupyter That’s a bummer. I was hoping that was a plausible injection point. Oh well… Once again, thanks for putting Carnets together! It’s a life saver for on the go Python usage on iPad!
@carnets_jupyter@nholzschuch Hey! Out of intellectual curiosity, would the recent relaxation by Apple allowing emulators allow the execution of JIT compiled code such as the one generated by Numba?
🎉 Great news! Pylustrator, my open-source tool for creating publication-ready @matplotlib plots, works seamlessly with the latest Matplotlib version! 🖌️✨ Curious how it can transform your figures and save your changes as reproducible Python code? Let me show you! 🧵👇
@gr_roman1770 API Efficiency matters a ton. When you own the Software and the Hardware you can do a ton more. Just look at Nvidia and CUDA. If Qualcomm had a more efficient API for its hardware it is very likely the gap would be smaller or even tilt in Qualcomm’s favor.
@gr_roman1770@jordanstonenz It’s exactly the same to compare OpenCL to CUDA to Metal to DirectX to Vulcan. The APIs are different but if the timing metrics or throughput metrics are the same, they are comparable. API efficiency, ease of use by developers or level of abstractions influence performance.
Open NotebookLM with Llama 3.1 405B in this @huggingface space.
Convert your PDFs into podcasts with open-source AI models (Llama 3.1 405B, MeloTTS, Bark).
Only the text content of the PDFs and the max total content is 100,000 characters due to the context length of Llama 3.1 405B.
@ThePbzwriter@Dachsjaeger There was a patch released today:
tl;dr:“Texture Quality now uses fixed-sized VRAM caches, reducing VRAM usage and performance degradation.”
@ThePbzwriter@Dachsjaeger I haven’t experienced similar issues but I tested it on a GPU with more VRAM. I read through Alex’s article and didn’t see it mentioned but I am still waiting to see the video. Hopefully it gets sorted out soon. It’s been such a great experience (story and gameplay) thus far.