Came across the 'Translate Gemma' paper a few days ago and was particularly intrigued by the token-level RM scores in the RL stage. Fed the PDF and my queries into Kimi 2.5 via CLI for a deep dive.
It helped me implement a MetricX-style evaluation suite supporting both API and local vLLM/sglang deployments. I’ve even tasked it to monitor vLLM/sglang source code for real-time updates. Will be pushing the code to GitHub soon to continue my research on token-level reward models! 🚀
🚀 Day 0: Warming up for #OpenSourceWeek!
We're a tiny team @deepseek_ai exploring AGI. Starting next week, we'll be open-sourcing 5 repos, sharing our small but sincere progress with full transparency.
These humble building blocks in our online service have been documented, deployed and battle-tested in production.
As part of the open-source community, we believe that every line shared becomes collective momentum that accelerates the journey.
Daily unlocks are coming soon. No ivory towers - just pure garage-energy and community-driven innovation.
🚀 Day 0: Warming up for #OpenSourceWeek!
We're a tiny team @deepseek_ai exploring AGI. Starting next week, we'll be open-sourcing 5 repos, sharing our small but sincere progress with full transparency.
These humble building blocks in our online service have been documented, deployed and battle-tested in production.
As part of the open-source community, we believe that every line shared becomes collective momentum that accelerates the journey.
Daily unlocks are coming soon. No ivory towers - just pure garage-energy and community-driven innovation.
🤔Now most LLMs have >= 128K context sizes, but are they good at generating long outputs, such as writing 8K token chain-of-thought for a planning problem?
🔔Introducing LongProc (Long Procedural Generation), a new benchmark with 6 diverse tasks that challenge LLMs to synthesize highly dispersed information and generate long, structured outputs.
1/3 Today, an anecdote shared by an invited speaker at #NeurIPS2024 left many Chinese scholars, myself included, feeling uncomfortable. As a community, I believe we should take a moment to reflect on why such remarks in public discourse can be offensive and harmful.
(1/7) Physics of LM, Part 2.1 with 8 results for LLM reasoning is out: https://t.co/fxJeO1nCdE. Probing reveals that LLMs secretly develop some "level-2" reasoning skill beyond Humans. Although I recommend watching my ICML tutorial first... Come in this thread to see the slides.
I’m happy to share the published version of our ConVIRT algorithm, appearing in #MLHC2022 (PMLR 182). In 2020, this was a pioneering work in contrastive learning of perception by using naturally occurring paired text. Unfortunately, things took a winding path from there. 🧵👇
@GeoCalhoun520@GFPhilosophy@Forbes Depending on the number you calculate, if there are 1.6 million deaths, at least multiplying that number by nearly 20 to 50 is the total number of people with the COVID-19 in China. I think this calculation is absurd.
@GeoCalhoun520@GFPhilosophy@Forbes Well, the total number of deaths according to your article is even more than the total number of infections in China. Of course, you can also assume that all these statistics are not credible.