AI chip startup Taalas @taalas_inc is showing off a chip that can do 16,000 tps/user on Llama3.1-8B, many multiples of its nearest competitor. The catch? The chip ONLY runs Llama3.1-8B, and a model like DeepSeekR1-671B would need 30 separate tapeouts:
https://t.co/IJuprQZqaE
Language models do just keep scaling!
Today we’re announcing Megatron-Turing NLG: a 530B parameter language model.
Joint work from @NVIDIA and @Microsoft. Trained using Megatron and DeepSpeed on DGX SuperPod.
https://t.co/xFN7c8xYQR
Megatron running on Selene sustains 140 TFlops per GPU on 1024 GPUs while training GPT-3.
That's 6X faster than the numbers reported in this paper, at about the same energy consumption.
Give Megatron a try and lower your CO2 emissions.
https://t.co/6Ht56c71lB
MLPerf Inference 1.0 results prove the incredible efficiency of A100 MIG in boosting infrastructure utilization for AI. A single A100 ran all seven MLPerf Inference 1.0 tests in offline mode simultaneously while providing 98% of the performance of a single MIG running alone.
Nvidia is now worth more than Intel + AMD combined.
That's the market's way of saying: x86 is a legacy architecture.
The workload of the future is SIMULATION.
Light -> raytracing
Physics -> n-body, CFD
AI -> neural nets
..and these will run on GPUs.
Congrats Microsoft team on great results achieved using V100 GPUs. Reducing time-to-train and cost-to-train continues to be vital for progress of AI. Cannot wait to see what A100 GPUs will enable for these workloads at scale.
https://t.co/fkomPIznHE
Now available our whitepaper providing detailed architecture information on our newly launched NVIDIA Ampere Architecture and A100 GPUs
https://t.co/4OEOtMPwRm
PC Gamers, let’s put those GPUs to work.
Join us and our friends at @OfficialPCMR in supporting folding@home and donating unused GPU computing power to fight against COVID-19!
Learn more → https://t.co/EQE4u7xTZT
Eagerly awaited MLPerf Inference results are out. NVIDIA Turing GPUs, Xavier SoC emerge toppers in their respective categories. Read our blog to learn more.
https://t.co/DU1YptEcLd
NVIDIA research team trained the largest ever language model based on transformers with a novel model parallel approach using Pytorch. The code that implements this approach is published on GitHub.
https://t.co/lKvRyUhysV
A huge milestone in natural language processing. Software optimizations used to accomplish these breakthroughs in conversational AI available to developers.
https://t.co/9PxEFqLMJv
Our Safety Force Field (SFF) vehicle software is designed specifically for collision avoidance. Watch how it double-checks vehicle controls and vetoes unsafe actions. #DRIVELabs https://t.co/0EDS9vm9eH
Important milestone in GPU accelerated Inference. NVIDIA TensorRT plugins, parsers, & samples are now open source. Easy to extend now with your custom models & layers. https://t.co/GrOxw3U7ZP