In May I handed in my final year dissertation at Queen Mary University of London titled "Personalised Audio Deepfake Detection". I am proud to share that I received a score of 82.4%, a first class result.
The idea for the project came from my job at Brightside AI. For the past 8 months, at the same time as finishing my degree at QMUL, I have also been working hard to build the best vishing (voice phishing) simulation platform from scratch. Seeing how common, easy, and convincing these attacks have become, one idea stuck with me: a fake voice is far easier to catch when you actually know the real person and what their voice sounds like. Most current audio deepfake detection systems completely ignore that. They simply ask "is this audio fake?" in isolation, while I wanted to ask a different question: "is this really X's voice?"
So, I built a two-stream detector (WavLM + AASIST) with a personalisation layer that conditions on a few seconds of a known person's real audio. On the standard benchmark it reached a 1.07% Equal Error Rate, beating both models on their own and competitive with the published AASIST baseline. This was enough to confirm what the field already knows: in-domain detection is essentially solved. The hard, unsolved problem is generalising to real-world audio, which is where I thought my personalisation hypothesis could make a difference. I deployed and self-hosted the whole thing end to end, using FastAPI to expose my detection architecture as an API paired with a Next.js frontend to conduct, an admittedly very flawed, listening study.
The findings were, of course, the most interesting part. Personalisation did help - a small but statistically significant improvement. Though it was uneven, helping some speakers and hurting others. The results provide a real signal, just not the breakthrough I'd secretly hoped for, and a big part of why comes down to my own setup: every model was trained on my five year old Mac with a mere 16GB of RAM, a comically tiny amount for this kind of work, though one look at current memory prices makes me want to treasure it. With frozen backbones and a single training dataset, the model simply never saw enough variety to generalise. If I pick this back up, more training data and more diverse training data is the first and biggest lever - likely the difference between a detector that shines on a benchmark and one that actually holds up in the real world.
I am grateful to my supervisor, Soumya, and to everyone who sat through 20 audio clips for the study. Plenty here I want to keep working on and I look forward to watching this space in the future.
You can find my whole project and report on my GitHub: https://t.co/sfRo0S7ZVZ
@tripl3z I think features like Telegram's rewrite with AI can be very helpful for the type of individuals you describe, acting as a translation layer between their normal communication style and corporatese. Human partners are an interesting idea but can be hit or miss
@bcherny Hi Boris, it would be great if we were able to let Claude work on a plan it generated in plan mode with completely fresh context. Something like a fourth "Yes, clear context and auto-accept edits" option
@bcherny Hi Boris, it would be great if we were able to let Claude work on a plan it generated in plan mode with completely fresh context. Something like a fourth "Yes, clear context and auto-accept edits" option
@AmirMushich Damn sometimes all it takes is a single line prompt, incredible results! I turned your prompt into a web Ui so anyone can try it with their own photo and tune the settings really easily
https://t.co/wwSwceXPdr
If you're an anxious traveler like me this is a must
Get every travel requirement for your trip in one checklist
Visa rules, health docs, customs forms, registration deadlines, all personalized to your passport and destination
Use it for free:
https://t.co/7BQ7btg5G7