Expectation: the age of the IDE is over
Reality: we’re going to need a bigger IDE
(imo).
It just looks very different because humans now move upwards and program at a higher level - the basic unit of interest is not one file but one agent. It’s still programming.
I'm pretty familiar with Groq's technology and have led research in similar directions.
Serving from SRAM is recognized as a frontier for achieving high Tokens/sec/user, becoming increasingly critical as more complex apps (agentic loops, test-time compute) emerge.
Nvidia DID NOT acquire Groq, contrary to the CNBC report
It's inference technology licensing and a few Groq execs will be joining Nvidia. Groq will still be a separate company.
https://t.co/94bq7PdNeS
Nice.
I would mention "improved inference systems" as an enabler behind many of the highlights.
A wide combination of domains enables wrapping AI as a readily available product (with an affordable price, reasonable response times, and quality) behind our favorite tailored GUIs.
TPUv7: Google Takes a Swing at the King
Potential end of the CUDA Moat?
Anthropic’s 1GW+ TPU Purchase,
The more (TPU) Meta/SSI/xAI/OAI/Anthro buy the more (GPU capex) you save,
Next Generation TPUv8AX and TPUv8X versus Vera Rubin,
All the 3D Torus Cube Diagrams you could possibly want and more...
https://t.co/Xh1ohGxVjB
How does @deepseek_ai Sparse Attention (DSA) work?
It has 2 components: the Lightning Indexer and Sparse Multi-Latent Attention (MLA). The indexer keeps a small key cache of 128 per token (vs. 512 for MLA). It scores incoming queries. The top-2048 tokens to pass to Sparse MLA.
🚀 Introducing DeepSeek-V3.2-Exp — our latest experimental model!
✨ Built on V3.1-Terminus, it debuts DeepSeek Sparse Attention(DSA) for faster, more efficient training & inference on long context.
👉 Now live on App, Web, and API.
💰 API prices cut by 50%+!
1/n
People who took a swing at LLM Dissagregation in 2023, probably invested all their savings in Nvidia during 2021-2022 (like I did). I guess that dissagregation will be key in mitigating Nvidia's dominance, and mass-tailored disaggregated AI systems will come.
A very nice article from SemiAnalysis https://t.co/NOKDgegh8u . Nice to see disaggregated serving and specialized hardware for it arriving >2.5 years after my first PPT discussing disaggregated serving for LLMs
I found that I can't join the X "Machine Learning" community, as I am blocked by the moderators. Guess who...
I probably blocked him randomly as a practice of avoiding hate. Let's talk tech, not politics
Pytorch and vLLM are under the Linux Foundation. Looking forward to a "Huggingface" replacement or similar pushed forward by LF. This will surely advance AI infrastructure and competition.
Besides clear technical reasons, another reason might be public hate speech and antisemitic propaganda by HF organization leaders. I do not feel safe promoting this by referring to a model on HF, or a space on Gradio. Time has come for a change
Introducing GitHub Copilot agent mode (preview): the next evolution in AI-assisted coding.
Copilot in agent mode is capable of running terminal commands, iterating on its own code to fix errors, and so much more.
Available in VS Code Insiders today. Learn more:
https://t.co/vY4JttoVLs