Microsoft just released VibeVoice - 1.5B SoTA Text to Speech model - MIT Licensed 🔥
> It can generate up 90 minutes of audio
> Supports simultaneous generation of > 4 speakers
> Streaming and larger 7B model in-coming
> Capable of cross-lingual and singing synthesis
Love the expressiveness and the emotion control on the model! Kudos to Microsoft 🤗
🎙 VibeVoice da Microsoft: TTS open-source revolucionário! 🗣 Até 4 vozes, 90 min de fala, super expressivo com LLM Qwen2.5-1.5B. 🔓 Ideal para podcasts e dublagens. 💡 Como você usaria essa tech? #VibeVoice#TTS#OpenSourc#VibeVoice#TTS#OpenSource
🚀 Grok 2.5 is now open-source! A powerful and accessible AI, ready to drive innovation. Finally, a true @OpenAI AI for everyone! 🙌 #Grok25#OpenSource#AI@xai
Vlw @elonmusk
The @xAI Grok 2.5 model, which was our best model last year, is now open source.
Grok 3 will be made open source in about 6 months.
https://t.co/TXM0wyJKOh
The @xAI Grok 2.5 model, which was our best model last year, is now open source.
Grok 3 will be made open source in about 6 months.
https://t.co/TXM0wyJKOh
🚀 Chatbot Developer seeking new opportunities! Skilled in conversational AI, Gen AI, and API integration. Delivering high-quality, innovative projects. 📩 DMs open for gigs! #JobSearch#Chatbots#JobSearch, #Chatbots, #AI, #DevJobs
🚀 Desenvolvedor de chatbots em busca de novas oportunidades! Expertise em IA conversacional, NLP e integração de APIs. Projetos entregues com qualidade e inovação. 📩 DM aberta para propostas! #VagaDev
state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce Megalodon, a neural architecture for efficient sequence modeling with unlimited context length. Megalodon inherits the architecture of Mega
Google presents Mixture-of-Depths
Dynamically allocating compute in transformer-based language models
Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate
In a first, AI was able to reconstruct images from brain activity with over 75% accuracy.
Japanese researchers have achieved a significant breakthrough in AI-generated imaging, reaching a record 75% accuracy in reconstructing images from brain activity.
This advancement marks a considerable improvement over previous methods which only achieved 50.4% accuracy. The process involves recording brain activity while subjects view images and later recalling these images.
Utilizing a neural signal translator and generative AI, researchers were able to reconstruct these images with high accuracy. This technology opens new possibilities in understanding the human mind and could lead to novel forms of non-verbal communication.
Here is my selection of papers for today (4 Dec):
paper pages: https://t.co/QzUWybzXRg
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
PyNeRF: Pyramidal Neural Radiance Fields
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
MoMask: Generative Masked Modeling of 3D Human Motions
Text-Guided 3D Face Synthesis -- From Generation to Editing
DREAM: Diffusion Rectification and Estimation-Adaptive Models
X-Dreamer: Creating High-quality 3D Content by Bridging the Domain Gap Between Text-to-2D and Text-to-3D Generation
VideoBooth: Diffusion-based Video Generation with Image Prompts
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
HiFi Tuner: High-Fidelity Subject-Driven Fine-Tuning for Diffusion Models
Beyond ChatBots: ExploreLLM for Structured Thoughts and Personalized Model Responses
Merlin:Empowering Multimodal LLMs with Foresight Minds
Instruction-tuning Aligns LLMs to the Human Brain
Dolphins: Multimodal Language Model for Driving
Towards Accurate Differential Diagnosis with Large Language Models
SeaLLMs -- Large Language Models for Southeast Asia
GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs
Cerca de 20.000 dos cristais identificados experimentalmente no banco de dados ICSD são computacionalmente estáveis. Abordagens computacionais extraídas do Materials Project, do Open Quantum Materials Database e do banco de dados WBM aumentaram esse número para 48.000 cristais estáveis. O GNoME expande o número de materiais estáveis conhecidos pela humanidade para 421.000.
Introducing GNoME: an AI tool that helped discover 2.2 million new crystals. 💎
Crystals are found in everything from the chips powering our phones to solar cells creating clean energy.
The model also better predicts the stability of new materials. 🧵 https://t.co/O3YdnVcJt1
Cerca de 20.000 dos cristais identificados experimentalmente no banco de dados ICSD são computacionalmente estáveis. Abordagens computacionais extraídas do Materials Project, do Open Quantum Materials Database e do banco de dados WBM aumentaram esse número para 48.000 cristais estáveis. O GNoME expande o número de materiais estáveis conhecidos pela humanidade para 421.000.
Introducing GNoME: an AI tool that helped discover 2.2 million new crystals. 💎
Crystals are found in everything from the chips powering our phones to solar cells creating clean energy.
The model also better predicts the stability of new materials. 🧵 https://t.co/O3YdnVcJt1
Introducing GNoME: an AI tool that helped discover 2.2 million new crystals. 💎
Crystals are found in everything from the chips powering our phones to solar cells creating clean energy.
The model also better predicts the stability of new materials. 🧵 https://t.co/O3YdnVcJt1
@ylecun@NandoDF Congratulations on the excellent work! Now, to enhance it even further, it would be amazing to acquire more specific data for Brazilian Portuguese. Let's enrich this experience even more
@yoachlacombe@metaai Congratulations on the excellent work! Now, to enhance it even further, it would be amazing to acquire more specific data for Brazilian Portuguese. Let's enrich this experience even more
Model summary
Developed by: Argilla (based on HuggingFace H4 and MistralAI previous efforts and amazing work)
Shared by: Argilla
Model type: GPT-like 7B model DPO fine-tuned
Language(s) (NLP): Mainly English
License: MIT (same as Zephyr 7B-beta)
Finetuned from model: alignment-handbook/zephyr-7b-sft-full
Repository: https://t.co/Imrm6ZjJvr
Paper: N/A
Play with Notus on HuggingChat: https://t.co/pXhgcXCTRk
@argilla_io@huggingface
Model summary
Developed by: Argilla (based on HuggingFace H4 and MistralAI previous efforts and amazing work)
Shared by: Argilla
Model type: GPT-like 7B model DPO fine-tuned
Language(s) (NLP): Mainly English
License: MIT (same as Zephyr 7B-beta)
Finetuned from model: alignment-handbook/zephyr-7b-sft-full
Repository: https://t.co/Imrm6ZjJvr
Paper: N/A
Play with Notus on HuggingChat: https://t.co/pXhgcXCTRk
@argilla_io@huggingface