We are SoundsRight (SN105) - Bittensor's SOTA Open-Source Speech Enhancement Subnet!
https://t.co/sD2X0ST9tb
As more and more of our daily lives revolve around consuming content online, more emphasis is placed on making it as high-quality as possible. The demand for better quality audio thus grows by the day!
That being said, speech enhancement is a rather complex field of study. For instance, with the task of denoising audio, separating the speech we want from background noise requires training models capable of distinguishing the two from each other under many different circumstances.
However, what we find is that much of this technology is hidden behind a paywall, when in reality all the components necessary to open the floodgates for open-source innovation are already in place.
SoundsRight is dedicated to the research and development of non-proprietary speech enhancement models through daily fine-tuning competitions, powered by the decentralized Bittensor ecosystem.
The subnet’s miners upload models to HuggingFace, and validators benchmark them on a fresh dataset generated every two days to find the miner with the best-performing model.
The subnet's models are already outperforming the SOTA open-source alternative on our benchmarks. We are excited to see what the future will bring! 🚀
$TAO #Bittensor #Subnet #OTF #DecentralizedAI #Audio #SpeechEnhancement #OpenSource
Unfortunately, we have deregistered. We'd like to take this moment to thank each and every person who contributed to SoundsRight. It's been a pleasure working with you all.
The body of open-source work created for speech enhancement is magnitudes larger than it was before the subnet's inception, and the world is that much better off because of it!
We will take some time off to learn from our mistakes and develop what is necessary for the subnet to succeed.
We wish you all the best in your future endeavors!
Best,
The SoundsRight Team
Today, we'd like to take a look at the benchmarks we are using for assessing miner model performance on the subnet!
We have five metrics, whose usage varies depending on the task/sample rate:
📊 PESQ (Perceptual Evaluation of Speech Quality)
This metric’s aim is to quanitify a person’s percieved quality of speech, and is useful as a holistic determination of the quality of speech enhancement performed.
It is important to note that PESQ only works for 8kHz and 16kHz audio.
📊ESTOI (Extended Short-Time Objective Intelligibility)
This metric’s aim is to quantify the intelligibility of speech–how easy it is to understand the speech itself.
📊SI-SDR (Scale-Invariant Signal-to-Distortion Ratio)
This metric determines how much distortion is present in the audio. Distortion can be thought of as unwanted changes to the speech signal as a result of the enhancement operation.
📊SI-SAR (Scale-Invariant Signal-to-Artifacts Ratio)
This metric determines the level of artifacts present in the audio. Artifacts can be thought of as new, unwanted components introduced as a result of the speech enhancement operation.
📊SI-SIR (Scale-Invariant Signal-to-Interference Ratio)
This metric determines the level of interference present in the audio. Interference can be thought of as unwanted audio from outside sources still present in the recording, such as the noise from a crowded room.
More information is available in our docs!
https://t.co/NxOQsRUum1
$TAO #Bittensor #Subnet #OTF #DecentralizedAI #Audio #SpeechEnhancement #OpenSource
SoundsRight is now on Medium! We think it's important for us to start writing down some of our thoughts.
We've just published our first article and would appreciate it if you'd give it a read!♥️
https://t.co/VG4eR592z5
$TAO #Bittensor#Subnet#OTF#DecentralizedAI#Audio #SpeechEnhancement #OpenSource
SN105 has been on quite a journey so far, and we'd like to take this time to reflect on what the subnet has accomplished, and the direction we'd like to go in for the future!
Speech enhancement has a lot of potential, as demand grows for higher-quality audio in all the content we view in our daily lives. The market is projected to reach $70bn in value by 2030 as higher-quality speech becomes more and more important in applications such as teleconferencing, online education and content creation.
https://t.co/8FiHbcUhmC
We chose to build on Bittensor because while simple speech enhancement (removing the noise from a fan from the background, and other such repetitive noise) has been done already, the problem becomes exponentially harder when you have more complex and dynamic noise (people talking, for example). This creates a need for an intelligence marketplace that fosters the creation of higher-quality models.
Our initial goal was to create a large body of open-source work for future miners to build off of once we create a monetized product, since not a lot existed before the subnet.
We started with fine-tuning competitions for 16 kHz open-source models, and subnet models quickly became the majority of the open-source speech enhancement models available on HuggingFace.
Since the models started to outperform the SOTA open-source comparison, we have now switched to hosting competitions for 48 kHz models (high-quality audio). 48 kHz models have also started performing above the SOTA open-source comparison, and you can track their progress on our leaderboard here:
https://t.co/L037NpC3dE
So, what is there to look forward to? Here is our plan:
1. Create an interface and API so that other teams can incorporate the subnet's open-source models into their workflows (almost done!)
2. Once we have a large enough body of open-source work, revamp the subnet so that models are now closed-source, and release the product at the same time.
The product will be:
a. A web app for retail (for example, content creators)
b. A DAW plugin for professionals
c. Licensing the models to companies for use (for example, live speech enhancement for teleconferencing)
We are been incredibly happy with the results so far! Bittensor is the perfect ecosystem for us, and we are incredibly grateful to the community for the opportunities they have provided us!♥️
We cannot wait to see what the future holds for our subnet!🚀
48 kHz competitions have been live for a bit over a week now on SN105, and subnet models are already outperforming the SOTA open-source comparison on our benchmarks! 🔥🔥
Check it out for yourself!
https://t.co/L037NpC3dE
We're excited to see how things progress! 🚀
$TAO #Bittensor #Subnet #OTF #DecentralizedAI #Audio #SpeechEnhancement #OpenSource
Happy to be a part of this!🎉
We strongly believe that transparency is necessary for the Bittensor ecosystem to reach its full potential!
Much love to the @TaoPortal team for this initiative, it's been a pleasure working with you all♥️
✅ 20% of Subnets Verified on Tao Revenue ✅
We’ve now verified on-chain metrics for 20% of all active Bittensor subnets, a strong step towards full transparency across the $TAO ecosystem.
View subnet on-chain inflows, outflows and profitability, fully transparent and verifiable on-chain.
https://t.co/7gfljsmo32