I wanted to map out the "picks and shovels"—the next Micron, SanDisk, or Vertiv—for a potential space-based AI data center economy.
Medium: https://t.co/XVMXGEuorA
Paper: https://t.co/9dT9v5ffRw
Website: https://t.co/fPlgtlleWz
::Personal hobbyist project not investment advice::
Core principle behind @elonmusk's companies seem to be "If it is theoretically possible, I'll conquer the engineering". I ended up going down a rabbit hole of looking at the technical challenges to building Orbital compute: https://t.co/QxzF78fdti
1/ today we're releasing muse spark, the first model from MSL. nine months ago we rebuilt our ai stack from scratch. new infrastructure, new architecture, new data pipelines. muse spark is the result of that work, and now it powers meta ai. 🧵
Understanding reinforcement learning for model training from scratch. This took me a lot longer to write than anticipated, partly because the RLMT literature is not an easy read: https://t.co/DysLNLl85p
Our CRAG-MM Challenge (KDD Cup 2025) invites you to develop innovative multi-modal, multi-turn question-answering systems with a focus on RAG, using agentic tools to retrieve information. The goal is to improve visual reasoning: https://t.co/yiXanHpoRZ
I think the market may be drawing the wrong conclusion about how AI will evolve going forward based on the Deepseek paper – looking at Nvidia stock (NVDA) today.
The paper drives improvements using Reinforcement Learning. In this case it is done in the areas of math and ..
learn using reinforcement – much like a child does. This is a compute intensive process as the rewards are generally going to be much more sparse than math or reasoning. I think in the future the compute requirements will grow dramatically and I won’t be surprised if ..
@AIreviewbro @Ahmad_Al_Dahle We do release all our eval data publicly to make reproducibility easier, this contains information such as prompts used and raw and parsed responses: https://t.co/psZZBs0lks
We are releasing Llama 3.3 today. An updated Lllama 70B open source instruct model which is comparable in performance to the 405B model. Happy holidays!!! 🥳 #llm#ai#llama
One Meta: https://t.co/fvPjQtVDmu
Oh Huggingface: https://t.co/IqTqELFSSk
We’re releasing 1B & 3B quantized Llama models with same quality as the original, while achieving 2-4x speedup. We used two techniques: Quantization-Aware Training with LoRA adaptors, and SpinQuant
https://t.co/NLaUlj9IOz
Curious about Large Language Models but don't know where to start? @_Rohit_Patel_'s latest article breaks it all down from the basics, requiring only your ability to add and multiply.
#LLM#ML
https://t.co/1ZPqmReoQI
Today we're open source releasing the latest versions of our Llama models, Llama 3.2. We have 1B/3B models for text and 11B/90B multimodal models: https://t.co/AcqJKkDhXh
Due to strong community interest, we've collaborated with @AIatMeta to compare the bf16 and fp8 versions of Llama-3.1-405b in Chatbot Arena!
With over 5K community votes, both versions show similar performance across the board:
- Overall: 1266 vs 1266
- Hard prompts: 1267 vs 1271
- Instruction following: 1269 vs 1266
In coding/longer queries, bf16 gets slightly higher score, but remains within the confidence intervals.
This is great news for the community: fp8 version could closely match bf16 performance while significantly reducing costs.
Leaderboard link in the below post👇