Oruk wins at emotion recognition. By a lot.
64 systems on our benchmark. Nine commercial APIs on one we had no hand in.
Same winner all three times and every number is checkable.
Oruk wins at emotion recognition. By a lot.
64 systems on our benchmark. Nine commercial APIs on one we had no hand in.
Same winner all three times and every number is checkable.
Sarcasm is where sentiment analysis goes to die. The words say the opposite of the meaning.
Same probe, same folds: oruk beats Deepgram on all three splits and stays above the trivial baseline even on a sitcom it never saw. Sentiment ends below it.
The voice knows when the words are lying.
Not our benchmark. Same winner.
IEMOCAP, identical audio through nine commercial APIs: oruk 71.5%, next best 63.6, and that vendor’s free tier quit after 118 clips.
The transcript-sentiment products cluster near chance. The signal is in the voice.
oruk Spectra: 77.6%. The next best thing anyone has: open model, audio LLM, commercial API, anything: 68.7.
64 systems, one label mapping, one scorer. Text-only LLMs reading transcripts cap at 41.2 — the words alone can’t do this job.
Six shapes a voice can trace through a sentence.
Each one makes several readings more likely. Not one of them picks a single answer.
There is structure in a tone of voice.
Same filters, same layer, before and after training. It does not scramble them, it tunes them.
Every figure in these posts is real measured data, and every one of them is live on the site: rotate the geometry, scrub through a model's layers, morph the filters yourself.
All three write-ups, free to read:
https://t.co/agWAlAXtIu
A speech model starts life with random noise for ears.
Thirty epochs later it has grown a bank of wavelets, each one tuned to a pitch, each one sharp in time.
Nobody designed these shapes. Watch them arrive.
A grid of six pitch-contour families down the side and listener labels across the top, with dots sized by how much more often each label was chosen.
No family lines up with a single label.
8,000 colored dots representing tiny slivers of speech
Each dot is a fraction of a second of someone talking.
Clustered into distinct neighborhoods by sound type:
vowels, plosives, fricatives, nasals, and more.
The model was never told what a vowel or an S sound is.
They sorted themselves anyway.
Here's the cool part...
Four AIs, built by different labs, trained on different data. They've never seen each other.
Same voices in → same shape out. All four agree on where anger lives.
We asked 91 people to say the same sentence — "It's eleven o'clock" — with anger, joy, fear, disgust, and sadness.
Then we watched every voice travel through an AI's brain.
This is what a sentence looks like inside a machine that's listening.
Zoom out. These are 8,000 tiny slivers of speech — each dot is a fraction of a second of someone talking.
The model was never told what a vowel or an S sound is. They sorted themselves anyway.
In 1952, phoneticians drew a map of human vowels by hand. It's in every linguistics textbook.
An AI that was never taught phonetics just… rebuilt it. Same map. Nobody showed it the textbook.