I have open-sourced TTS model implementation based on LatentLM supercharged with Transformer engine!!
I believe the next frontier of spoken language modeling is based on continuous tokens.
Check this out!
https://t.co/zXduhxWOqc
Muso Actionとして、エクイティ+政策金融公庫の融資を合わせて総額1.0億円の資金調達を実施しました。
East Ventures、GMO AI&ロボティクス商事、田中渓さんにご出資いただきました。
ロボット基盤モデルを軸に、「現場で本当に使われるロボットワーカー」の開発を加速させていきます。
そして、創業期を一緒に創っていくAIロボティクスエンジニア/ロボット制御エンジニアを本気で採用中です👇
このタイミングでぜひ一緒にやりましょう。
✨ Meet our new open family of models: @NVIDIA Nemotron 3
Open in weights, data, tools, and training, Nemotron 3 is built for multi-agent apps and features:
• An efficient hybrid Mamba‑Transformer MoE architecture
• 1M token context for long-term memory and improved reasoning
• Multi‑environment reinforcement learning via NeMo Gym for advanced skill adaptation
Plus NVFP4 pre-training, latent MoE, 1T tokens of data, and more.
📗Read the details in our tech blog: https://t.co/9FD1jGp3zc
🤗 Try the model on @huggingface: https://t.co/n7an9b7hN9
🎶 Meet Audio-Flamingo 3 – a fully open LALM trained on sound, speech, and music datasets. 🎶
Handles 10-min audio, long-form text, and voice conversations. Perfect for audio QA, dialog, and reasoning.
On @huggingface ➡️ https://t.co/ubf3G3jdvj
From #NVIDIAResearch.
NVIDIA Japan で私と同じチームの求人がオープンしています。生成 AI の学習や推論に GPU を使うのをサポートする仕事です。いわゆる SA 業務がメインにはなりますが、最新技術や論文などのキャッチアップも推奨される環境で、エンジニアとしても楽しい職場です。
ご興味ある方、ぜひご応募ください!