Va a dejar a los jóvenes andaluces sin acceso a la universidad a no ser que se hipotequen por 20 años. Pero ojo, que se puso botas de agua y se metió en un charco durante las inundaciones...
@BrioEnfurecida Mañana hay recogida en El Puerto!! Me consta que Grazalema ha recibido muchas donaciones y están bastante cubiertos, pero los que organizan la recogida van a intentar mover para otros pueblos.
🚨 La Universidad de Almería reconoce que debe dinero a su personal, pero que no puede pagarlo porque la Junta del PP no ha puesto los fondos.
Es salario ya reconocido. No es un extra.
El Rector ha tenido que dictar una resolución solo para que no prescriba la deuda.
Así “apoya” Moreno Bonilla a la universidad pública:
asfixiándola para que crezca lo privado. 🎓✊
Researchers: Quality work takes time. That’s a strength, not a setback. ✍️
Great projects don’t come from rushing; they come from giving yourself the margin to think, revise, and get it right. Build extra time into your process so you can do your best work without burning out. 💭
Reinforcement Learning for LLMs just got a blueprint — and it starts with a first-order approximation.
Reinforcement learning is powering the next generation of reasoning-capable LLMs… but anyone who has trained RL on large models knows one truth:
👉 Stability is the real bottleneck.
Collapse, oscillation, reward hacking, and training–inference mismatch have become recurring nightmares.
The Qwen team’s new paper, “Stabilizing Reinforcement Learning with LLMs: Formulation and Practices,” delivers one of the cleanest explanations yet of why LLM RL breaks and more importantly, how to fix it at scale.
Here’s the big idea:
💡 The Token-Level Objective Isn’t Wrong, It’s an Approximation
The paper proposes a first-order approximation showing that token-level objectives like REINFORCE can optimize sequence-level rewards… but only when two gaps remain small:
- Training–Inference Discrepancy
The numerical mismatch between the training engine (Megatron/FSDP) and inference engine (vLLM/SGLang).
This mismatch alone is enough to destabilize RL, especially in FP8 inference settings.
- Policy Staleness
When rollouts and updates are out of sync (off-policy updates, batching, async rollout, etc.), the approximation breaks.
When either gap grows too large → gradients explode, KL spikes, and RL collapses.
🧠 Why MoE Models Are Especially Fragile
- For Mixture-of-Experts models, expert routing magnifies both problems:
- Training & inference may activate different experts
- Parameter updates shift routing, worsening policy staleness
This explains why RL on MoE models is notoriously unstable compared to dense models.
🔧 The Fix: Importance Sampling + Clipping + Routing Replay
The team shows, across hundreds of thousands of GPU hours, that three techniques consistently stabilize RL:
✔ Importance Sampling (IS) correction
✔ Clipping to restrain aggressive updates
✔ Routing Replay (R2/R3) for MoE to freeze experts during updates
Their findings:
On-policy RL?
Simple IS-corrected REINFORCE is surprisingly the most stable.
Off-policy RL?
Routing Replay becomes essential —
R2 works better at small off-policiness,
R3 wins at large off-policiness.
Initializations matter less than stability.
Once RL is stable, different cold-starts converge to nearly identical final performance.
If you’re building reasoning models, MoE models, or scaling RL pipelines, this is the paper you should read this week.
#ReinforcementLearning #LLMTraining #LargeLanguageModels #AIResearch #RLHF #MoE #DeepLearning #MachineLearning #Qwen #AIStability #AIGovernance #ModelOptimization #AIEngineering
Popper used to begin his lecture course on the philosophy of science by asking the students simply to ‘observe’. Then he would wait in silence for one of them to ask what they were supposed to observe. This was his way of demonstrating one of many flaws in the empiricism that is still part of common sense today. So he would explain to them that scientific observation is impossible without pre-existing knowledge about what to look at, what to look for, how to look, and how to interpret what one sees. And he would explain that, therefore, theory has to come first. It has to be conjectured, not derived.
@DavidDeutschOxf
Google acaba de lanzar esta BOMBA atómica de paper, quizás comparable con el que comenzó la revolución de IA.
No exagero: explica por qué ChatGPT, Gemini,etc tienen el mismo problema: no pueden aprender después de su entrenamiento. La solución que proponen es MUY elegante: 🧵
🚨 ASAMBLEA PRECARIUS: ELECCIONES RECTORADO
1️⃣ Analizaremos las propuestas de los candidatos
2️⃣ Presentaremos el informe de nuestras reuniones
3️⃣ Definiremos nuestras reclamaciones como PDI precarizado
💪 El futuro de la @unisevilla no puede decidirse sin nuestra voz.
(2/2) Implica clases impartidas por otros profesores de gratis o, si no se da a basto, se dejan sin cubrir. Se avisa a personal docente, pero no parece importar. Pero, oh milagro, esta vez no ha sido así. Contratados ipso-facto. Cómo se notan los tiempos de campaña, ¿eh?
(1/2) Siempre pasa igual, echan a PSI para recontratarlos por motivo de cambio de plaza adjudicada. Aunque en personal docente estén alertados y los papeles estén en su carpeta de entrada, demoran el proceso de contratación entre una semana... ¡e incluso meses!
Los políticos, de todos los colores, mintiendo en sus currículums, y yo aterrorizado cada vez que presento un proyecto por si pongo mal el número de páginas de los artículos científicos y piensan que miento. Qué hartura de ser pobre.
Are you a student or early-career researcher in BPM? Join the *Mentoring Lunch* hosted by the *DEI Committee* for a relaxed, inspiring chat with experienced professors!
📅 Wed, 3 Sept
🕐 13:00–14:30
📍 Magnolia Room
🔗 Register by 10 August
👉 https://t.co/RnMJCeChxi
📲📃Desde FPU Investiga compartimos esta carta abierta dirigida al @cienciagob para denunciar la situación límite que sufren las personas FPU que solicitan estancias.
👥 ¡Exigimos una solución ya!
📢 ¡Comparte!