DemoHumane IntelligenceNov 2025
Happy or Angry? An Emotion-AI Fairness Audit
I won the Data Track of the Bias Bounty Challenge, run by Humane Intelligence with Valence AI and CoNA Lab at Virginia State University. I audited Valence's emotion AI by labeling 55 voice clips myself, comparing my labels with its labels, and tracing where we differed. Step through how I did it.
Stage 1 of 5
I labeled every clip twice
I labeled all 55 clips twice, 3 days apart. At first I hid the filenames, because they held an emotion hint.
For each clip I coded five things.
- Emotion label
- The basic 4, plus 5 I added.
- Voice
- Pitch, pacing, pauses, volume and how clear the pronunciation was.
- Words
- Word choice, sentence type and any explicit emotional language.
- Demographics
- My estimate of gender, age and neurotype.
- Confidence
- Emotion AI's confidence score. I called it high above 0.5 and low below.
I was the only labeler. I am a Korean woman in my 30s, English is my second language, and I had little prior contact with neurodivergent speech. That shaped my labels too.
Stage 2 of 5
Emotion AI's labels next to mine
Emotion AI gives each clip one of four labels. Pick an attempt to see how my labels compare.
Else is any label outside the basic 4. Match rates compare Emotion AI with my 2nd attempt. For sad, 1 of Emotion AI's 7 clips matched; my write-up's table lists 28.57%, which is a typo.
Emotion AI's mean confidence was 0.43, under the 0.5 line. On average it was unsure, yet it still gave each clip one label.
Stage 3 of 5
Where the 14 mismatches went
Pick one of Emotion AI's labels to see what I called those clips in my 2nd attempt.
Emotion AI said
Sad6 of its 7 clips differEmotion AI's sad split into anxious, upset, scared and annoyed. The feeling was more specific than four labels allow.
| Emotion AI | My label, 2nd attempt | Clips |
|---|---|---|
| Total | 14 |
Stage 4 of 5
Four bias patterns in the mismatches
The 14 mismatches showed four patterns. Pick one.
Emotion AI
Me
Stage 5 of 5
What I proposed
A higher match rate alone does not make emotion AI fair. The goal is respecting different ways of communicating. I proposed four changes so emotion AI can serve neurodivergent people in video calls.
Multimodal weighting you can see
Show a label and confidence for voice and for words separately, and let people weight them. For autistic users, voice might count less than content.
A wider emotion taxonomy
There is no single objective label for emotion. Add intensity, more specific labels, more than one label per clip, and an explicit "multiple interpretations possible".
Labeling you can check
Bias lives in the labels too. Publish who labeled the data and why they chose each label, and read the sentences around a clip before labeling it.
Inclusive data standards
Build with neurodivergent people, not only for them. At least 30% neurodivergent speakers in training data, neurodivergent people on every labeling team, and testing on neurotypical and neurodivergent clips separately before launch. Like the curb cut, this makes video calls work better for everyone.
Takeaway
The 74.6% match turned out to be shared bias, not accuracy.
Emotion AI and I both stumbled on flat tone, were more confident with male and younger voices, and struggled with atypical prosody. We agreed because we learned the same assumptions about how emotions should sound.