Netflixより。LLM-as-a-Judge の継続的改善
The Lifecycle of LLM-as-a-Judge: Building, Aligning, and Monitoring at scale
https://t.co/g7HO6O076j
LLM-as-a-Judgeに使っているLLMも当然劣化する。この記事ではこの部分の品質をどうやって維持するかについて書いている。
端的に言えば、Human-in-the-loopで取り組む。例えば、おすすめ作品の推薦文というタスクのLLM-as-a-Judgeの場合、推薦文の良し悪しだけでなく、推薦文がどう良いか悪いかを説明できる必要がある。同じタスクを人間にやらせて、人間の結果と合うかどうかでLLM-as-a-Judge自体の品質評価をしている。
DeepSeek just released DSpark for V4 Flash & Pro, a new speculative decoding method boosting throughput by 51% to 400%!
DS also showed DSpark works well for other models like Gemma & Qwen
Github: https://t.co/EGVYpc1kcK
Paper: https://t.co/TaBMRVlaW9
HF: https://t.co/289jVU2pxh
New NVIDIA physical AI agent-ready skills are changing how robotics researchers work. 🤖
With NVIDIA robotics skills, researchers can automate the most common development steps from scene preparation to simulation and learning with Omniverse libraries, Isaac frameworks, and physical AI open datasets.
Specialized skills like Isaac Mobility extend that even further.
Learn more ➡️ https://t.co/WwGv2XYpEi