Tried Gemma 4 ran locally on my iPhone today
I thought it'd be useful in case the apocalypse happens and I need to ask it for survival tips
Like how to make a fire ๐ฅ
I guess I'll freeze to death instead ๐ซ
Is it possible I used Claude Code so much somehow my USB-C connectors burned out on my MacBook Pro? Two of mine are dead and wonโt charge and now the third maxes out at 15 watts
Now I am afraid my code tamagotchi is about to die
Solve this in under 5 minutes and Iโll offer you $500k/year in cash plus several million in equity
I'm building a Computer-Use team, goal is to use computers better than humans
No experience or PhD needed
Instructions:
1. Solve all 30 challenges on this website in under 5 minutes: https://t.co/7DmmS5OQNl
2. Feel free to use any tools or vibe code it. Provide us a zip folder with instructions on how to run the agent and reproduce your results, as well your run statistics
3. The agent should be able to solve all the challenges, use browser, and provide overall metrics around time taken, token usage and token cost. Your agent must solve this challenge in under 5 minutes
Email your response: [email protected]
If you have any questions about this challenge, feel free to email us
@deepseek_ai just found a surprisingly clean way to make โmulti-laneโ residuals actually train efficiently, and the benchmark jumps are hard to ignore. Here is the breakdown:
๐ Primer concept: the residual โhighwayโ is the main path that carries what the model already knows forward through every layer while each layer adds a small tweak on top.
๐ฃ๏ธ Primer idea: split the residual โhighwayโ into multiple parallel lanes so different lanes carry different info, then mix them across layers.
โ ๏ธ Problem: if each layer can freely โamplifyโ lanes, gains compound โ some lanes explode, others go dead โ unstable training.
โ Fix: a simple conservation rule โ layers can redistribute/mix across lanes but canโt increase total magnitude (like moving sliders on a mixer board without raising the master volume). This is the gist for mHC.
๐ฏ Why it matters: activation stays stable even with many stacked layers โ predictable training + better quality
This is impressive research. Weโre excited to see the next generation of models built on top of this architecture.
๐ Eval Protocol is Open Sourced!
Reinforcement fine-tuning is complicated, because there are hundreds of environments and tens of trainers you can pick and choose and integrate with. Even worse, in production, agents donโt live in clean โgyms.โ They operate in messy, async environments - flaky APIs, partial observability, conflicting objectives, long feedback loops.
We solve that problem by open sourcing Eval Protocol. The goal is for you to build your production RFT flow without reinventing the wheel of managing such complex integration.
๐ Day 0 support for trainers and environments like TRL (@huggingface), @rllm_project , OpenEnv (@PyTorch),ย as well as support for proprietary trainers like @OpenAI RFT and @thinkymachines Tinker. More to come.
๐ Instrument agents in production instead of toy or simulated environments
๐ Move from offline benchmarks to live, continuous improvement