c-emolia-VN1

Corpus emolia in isolation, rule VN1.

Rule. VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C
Source. trajectories_v5.parquet  |  Family. one corpus in isolation
Sampled from 2,304,000 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
BRGT — brightness of timbrec-emolia-VN1 · #1

This is a VoiceNet dimension, not an emotion: brightness of timbre (BRGT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with brightness of timbre (BRGT) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.57, higher than 57 % of clips in this corpus. That is a total rise of 0.56.

It takes 5 clips to get there. Clip to clip the moves are +0.01, then +0.24, then +0.22, then +0.10 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 24 s · en · emolia

hear it un-normalised (raw levels, max seam 2.9 dB)
k 5d_a 0.565d_b 0.565step_a 0.237step_b 0.237min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_yyUqGQ0dJX0track EN_yyUqGQ0dJX0total 24.2slevel spread 2.9 dBmax seam 2.9 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · thin
(helplessness, fatigue exhaustion, distress · measured, normally alert, fully relaxed, storytelling) I call my dad. I'm leaving this stupid school.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, moderately variable; timbre is slightly cool, dark, very rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is deeply negative, slightly dominant, guarded; reads as helplessness, fatigue exhaustion, distress; style: storytelling, casual; poor recording, noisy background; mildly explicit content; genuineness 2.5/6; vocal-burst blend 2.0/10; 4.0s, EN.
EN_yyUqGQ0dJX0_W000161 · in -22.5 dBFS · gain +2.5 dB · emolia-00863
(confusion, doubt, impatience and irritability · normal-paced, energised, tense, casual) This little engine mission was to take some achy-chicka-chicka-chicka-chicka-chicka.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, tense, volatile; timbre is cool, dark, very rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, guarded; reads as confusion, doubt, impatience and irritability; style: casual, cartoonish; poor recording, some background noise; mildly explicit content; genuineness 4.7/6; vocal-burst blend 3.8/10; 4.8s, EN.
EN_yyUqGQ0dJX0_W000164 · in -20.3 dBFS · gain +0.3 dB · emolia-00863
(astonishment surprise, intoxication altered states of consciousness, amusement · normal-paced, energised, tense, casual) Aww, look at her. Giving those bedtime eyes again.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, tense, moderately variable; timbre is slightly cool, dark, very rough, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly dominant, guarded; reads as astonishment surprise, intoxication altered states of consciousness, amusement; style: casual, playful; below-average recording, noisy background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.5/10; 3.5s, EN.
EN_yyUqGQ0dJX0_W000165 · in -20.4 dBFS · gain +0.4 dB · emolia-00863
(disgust, intoxication altered states of consciousness, malevolence malice · measured, energised, neutral tension, storytelling) Not even when they climbed aboard the train and popped in his way across the dresser.
full caption & clip details
An adult masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as disgust, intoxication altered states of consciousness, malevolence malice; style: storytelling, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.1/10; 4.6s, EN.
EN_yyUqGQ0dJX0_W000166 · in -22.8 dBFS · gain +2.8 dB · emolia-00863
(pain, astonishment surprise, amusement · normal-paced, energised, tense, casual) Oh yea, she crazy. Okay. She crazy. She really in the major pain. In major pain.
full caption & clip details
A child somewhat masculine voice; delivery is energised, normal-paced, tense, volatile; timbre is slightly cool, dark, slightly rough, thin; slurred, some disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, very vulnerable; reads as pain, astonishment surprise, amusement; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 4.4/6; vocal-burst blend 3.2/10; 6.7s, EN.
EN_yyUqGQ0dJX0_W000174 · in -19.9 dBFS · gain -0.1 dB · emolia-00863
S_DRAM — style: dramaticc-emolia-VN1 · #2

This is a VoiceNet dimension, not an emotion: style: dramatic (S_DRAM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: dramatic (S_DRAM) around average — 0.54, higher than 54 % of clips in this corpus — and ends with it high at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.24.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.72 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.72 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · en · emolia

hear it un-normalised (raw levels, max seam 3.4 dB)
k 2d_a 0.243d_b 0.243step_a 0.243step_b 0.243min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_Ma2xaFh1jDctrack EN_Ma2xaFh1jDctotal 12.0slevel spread 3.4 dBmax seam 3.4 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · slightly cool, dark, thin, below-average recording, some background noise, normal-paced, normally alert, moderately variable
(confusion, helplessness, impatience and irritability · fully relaxed, audible breath, casual, conversational) I can't find anything green on this room.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is slightly cool, dark, very rough, thin; slurred, some disfluency, wide pitch range, audible breath; affect is mildly negative, slightly submissive, neutral openness; reads as confusion, helplessness, impatience and irritability; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.4/6; vocal-burst blend 1.5/10; 3.5s, EN.
EN_Ma2xaFh1jDc_W000151 · in -21.1 dBFS · gain +1.1 dB · emolia-02600
(teasing, amusement, affection · relaxed, heavy breath, casual, conversational) (childlike giggle) Ah, a book, yeaah. Look, how much things you have. I have just a book. Okay, I'll take this book.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, affection; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.3/6; vocal-burst blend 1.7/10; 8.3s, EN.
EN_Ma2xaFh1jDc_W000152 · in -17.6 dBFS · gain -2.4 dB · emolia-02600
R_THRT — resonance: throatc-emolia-VN1 · #3

This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: throat (R_THRT) below average — 0.27, lower than 73 % of clips in this corpus — and ends with it above average at 0.64, higher than 64 % of clips in this corpus. That is a total rise of 0.36.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 45 s · en · emolia

hear it un-normalised (raw levels, max seam 5.7 dB)
k 3d_a 0.364d_b 0.364step_a 0.238step_b 0.238min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00029_S05686track EN_B00029_S05686total 44.7slevel spread 5.7 dBmax seam 5.7 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, slightly relaxed, light breath
(infatuation, helplessness, emotional numbness · normal-paced, normally alert, steady, monologue) Vicky's voice changed completely as she said to Phil.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, dark, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, helplessness, emotional numbness; style: monologue, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 5.3/10; 3.1s, EN.
EN_B00029_S05686_W000052 · in -28.2 dBFS · gain +8.2 dB · emolia-00813
(triumph, contentment, shame · measured, subdued, fairly steady, narration) (ahem) Cut that in before the shots of Daniel singing under the tree. Right, said Phil. There was a moment of quiet when they finished. Then all the young people began to clap. Vicki looked surprised. That's enough, she said, and picked up her bag and papers. I'll see you at the grave about 4.30, Daniel. Bring Lily Anne and Blessing. See you in the car, Phil.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as triumph, contentment, shame; style: narration, storytelling; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.6/10; 27.7s, EN.
EN_B00029_S05686_W000053 · in -22.5 dBFS · gain +2.5 dB · emolia-00813
(malevolence malice, fear, affection · measured, normally alert, steady, narration) Those kids are going to need help very soon and you have nothing to offer. And that girl's bad news. And she lives a thousand kilometers away. Forget her. Vicky had disappeared.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as malevolence malice, fear, affection; style: narration, storytelling; good recording, no background noise; mildly explicit content; genuineness 0.9/6; vocal-burst blend 1.3/10; 13.7s, EN.
EN_B00029_S05686_W000054 · in -26.9 dBFS · gain +6.9 dB · emolia-00813
VALN — valencec-emolia-VN1 · #4

This is a VoiceNet dimension, not an emotion: valence (VALN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with valence (VALN) low — 0.21, lower than 79 % of clips in this corpus — and ends with it high at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.60.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.19, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 40 s · zh · emolia

hear it un-normalised (raw levels, max seam 3.6 dB)
k 4d_a 0.597d_b 0.597step_a 0.242step_b 0.242min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00004_S09748track ZH_B00004_S09748total 39.9slevel spread 3.6 dBmax seam 3.6 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · balanced body, average recording, slightly relaxed, fairly steady, moderate pitch range
(contempt, malevolence malice, anger · fast, normally alert, some disfluency, storytelling) 原来俩人出去的时候,带多少人呢?一百多人嘛,哎在外边足足,过了十三年,就剩他们俩回来了。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, malevolence malice, anger; style: storytelling, authoritative; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 5.8/10; 7.8s, ZH.
ZH_B00004_S09748_W000018 · in -17.6 dBFS · gain -2.4 dB · emolia-03320
(intoxication altered states of consciousness, malevolence malice, contempt · measured, very low-energy, some disfluency, storytelling) 好言好语的慰劳啊,说你们俩辛苦啦,我要重用你们。
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is warm, slightly dark, slightly rough, balanced body; slurred, some disfluency, moderate pitch range, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, malevolence malice, contempt; style: storytelling, narration; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 5.9/10; 5.1s, ZH.
ZH_B00004_S09748_W000019 · in -21.2 dBFS · gain +1.2 dB · emolia-03320
(pride, malevolence malice, astonishment surprise · normal-paced, normally alert, some disfluency, monologue) (low mumble) (ahem) 哎,咱们啊算是把张骞这个前因这个故事讲完了,然后还得回到说张骞随大将军卫青出征这段历史里去啊,因为他很熟悉匈奴地理和情况,他就随着大将军卫青一块儿打仗。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, malevolence malice, astonishment surprise; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 8.0/10; 15.6s, ZH.
ZH_B00004_S09748_W000020 · in -20.2 dBFS · gain +0.2 dB · emolia-03320
(measured, normally alert, frequent disfluency, monologue) (ahem) 哪儿啊有水哪儿有马能吃的草,卫青都特地要去问张谦张骞,就给他指大军得到了不少好处啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.7/10; 10.9s, ZH.
ZH_B00004_S09748_W000021 · in -20.8 dBFS · gain +0.8 dB · emolia-03320
TEMP — tempoc-emolia-VN1 · #5

This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with tempo (TEMP) above average — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the range at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 9 s · zh · emolia

hear it un-normalised (raw levels, max seam 0.9 dB)
k 2d_a 0.250d_b 0.250step_a 0.250step_b 0.250min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00011_S03448track ZH_B00011_S03448total 9.1slevel spread 0.9 dBmax seam 0.9 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
(astonishment surprise, impatience and irritability · measured, authoritative, formal) 投的票也是,我就是你每一轮你说的话。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as astonishment surprise, impatience and irritability; style: authoritative, formal; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 2.5/10; 4.3s, ZH.
ZH_B00011_S03448_W000020 · in -17.3 dBFS · gain -2.7 dB · emolia-03389
(normal-paced, authoritative, conversational) 和你言行是完全不一样的。你第一轮踩j歪,你投j歪了吗?你没有投吧。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.9/10; 4.7s, ZH.
ZH_B00011_S03448_W000021 · in -18.2 dBFS · gain -1.8 dB · emolia-03389
COGL — cognitive loadc-emolia-VN1 · #6

This is a VoiceNet dimension, not an emotion: cognitive load (COGL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with cognitive load (COGL) high — 0.87, higher than 87 % of clips in this corpus — and works its way down to around average at 0.55, higher than 55 % of clips in this corpus. That is a total fall of 0.32.

It takes 5 clips to get there. Clip to clip the moves are +0.07, then -0.24, then +0.07, then -0.21 — not a clean run: step 1 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of -0.01 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.01 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 57 s · en · emolia

k 5d_a -0.320d_b -0.320step_a 0.243step_b 0.243min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_kb6O5YkvGW0track EN_kb6O5YkvGW0total 57.1slevel spread 4.5 dBmax seam 4.5 dB
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, quiet background
(contemplation · normal-paced, normally alert, slightly relaxed, casual) Like expecting from you that within half a day you will make some intervention. Sometimes, you know, you have so many different things to do that you are not able, and people are that dissatisfied when you are not intervening just immediately to some (low mumble) situations, but you know, you must really feel the pulse of the country.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: casual, monologue; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 7.6/10; 17.8s, EN.
EN_kb6O5YkvGW0_W000374 · in -19.2 dBFS · gain -0.8 dB · emolia-01034
(doubt · measured, normally alert, slightly relaxed, casual) In my opinion, not that to, (low mumble) uh, to get what is, what is going, (ahem) what is going on.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: casual, monologue; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 2.2/10; 4.7s, EN.
EN_kb6O5YkvGW0_W000375 · in -16.7 dBFS · gain -3.3 dB · emolia-01034
(thankfulness gratitude, fatigue exhaustion, relief · normal-paced, very low-energy, relaxed, conversational) Thank you. We are slowly running out of time, but (ahem) I was wondering whether (low mumble) (ahem) you would be (ahem) so kind and stay with us like five (low mumble) minutes more to take...
full caption & clip details
An elderly feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, fatigue exhaustion, relief; style: conversational, playful; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.2/10; 11.9s, EN.
EN_kb6O5YkvGW0_W000376 · in -16.1 dBFS · gain -3.9 dB · emolia-01034
(embarrassment, amusement · normal-paced, normally alert, slightly relaxed, casual) I have just five minutes more. Sorry for this, but you know, that's how it looks.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as embarrassment, amusement; style: casual, conversational; below-average recording, quiet background; genuineness 6.0/6; vocal-burst blend 5.2/10; 6.6s, EN.
EN_kb6O5YkvGW0_W000377 · in -20.5 dBFS · gain +0.5 dB · emolia-01034
(thankfulness gratitude, interest · normal-paced, normally alert, slightly relaxed, casual) (ahem) Uh, thank you so much for taking the time to answer all our questions and for a speech. So my question would be the following. Uh, (ahem) what have been the measures taken under so-called antixenophobic document that you signed in Warsaw last year with the Ukrainian ombudsman Denisova?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, interest; style: casual, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.8/10; 15.6s, EN.
EN_kb6O5YkvGW0_W000381 · in -18.5 dBFS · gain -1.5 dB · emolia-01034
ARSH — harshness of articulationc-emolia-VN1 · #7

This is a VoiceNet dimension, not an emotion: harshness of articulation (ARSH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with harshness of articulation (ARSH) high — 0.79, higher than 79 % of clips in this corpus — and works its way down to around average at 0.44, lower than 56 % of clips in this corpus. That is a total fall of 0.35.

It takes 4 clips to get there. Clip to clip the moves are -0.14, then -0.03, then -0.17 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 48 s · en · emolia

k 4d_a -0.345d_b -0.345step_a 0.172step_b 0.172min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_PWaeOS4irmEtrack EN_PWaeOS4irmEtotal 48.3slevel spread 1.2 dBmax seam 1.1 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(emotional numbness, concentration · measured, fairly narrow pitch, light breath, formal) The background is described by par. A flat initio means, from first principles, or, from the beginning.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.3s, EN.
EN_PWaeOS4irmE_W000002 · in -16.2 dBFS · gain -3.8 dB · emolia-02237
(pain, emotional numbness · normal-paced, moderate pitch range, light breath, formal) Implying that the only inputs into an ab initio calculation are physical constants
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.9/10; 5.6s, EN.
EN_PWaeOS4irmE_W000003 · in -15.1 dBFS · gain -4.9 dB · emolia-02237
(concentration, emotional numbness · normal-paced, moderate pitch range, minimal breath, newsreading) A-flat initio quantum chemistry methods attempt to solve the electronic Schrödinger equation given the positions of the nuclei and the number of electrons in order to yield useful information such as electron densities, energies and other properties of the system.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, emotional numbness; style: newsreading, authoritative; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.0/10; 16.5s, EN.
EN_PWaeOS4irmE_W000004 · in -15.2 dBFS · gain -4.8 dB · emolia-02237
(concentration · normal-paced, moderate pitch range, minimal breath, newsreading) A-flat initio electronic structure methods have the advantage that they can be made to converge to the exact solution, when all approximations are sufficiently small in magnitude and when the finite set of basis functions tends toward the limit of a complete set.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 16.5s, EN.
EN_PWaeOS4irmE_W000007 · in -15.0 dBFS · gain -5.0 dB · emolia-02237
S_STRY — style: storytellingc-emolia-VN1 · #8

This is a VoiceNet dimension, not an emotion: style: storytelling (S_STRY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: storytelling (S_STRY) above average — 0.69, higher than 69 % of clips in this corpus — and works its way down to below average at 0.32, lower than 68 % of clips in this corpus. That is a total fall of 0.37.

It takes 5 clips to get there. Clip to clip the moves are -0.21, then +0.02, then -0.16, then -0.03 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.80 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.80 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 30 s · en · emolia

k 5d_a -0.373d_b -0.373step_a 0.208step_b 0.208min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00005_S09778track EN_B00005_S09778total 29.9slevel spread 1.0 dBmax seam 1.0 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, clear
(fear, distress, helplessness · slow, steady, almost no disfluency, narration) They worry that the Earth's climate is changing and that this may be harmful. The greenhouse effect. The Earth's climate.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, distress, helplessness; style: narration, whispered; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.1s, EN.
EN_B00005_S09778_W000005 · in -19.1 dBFS · gain -0.9 dB · emolia-00370
(awe · normal-paced, fairly steady, almost no disfluency, narration) Of all of the planets in our solar system, Earth has the most hospitable climate for human life.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: narration, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 6.1s, EN.
EN_B00005_S09778_W000006 · in -18.1 dBFS · gain -1.9 dB · emolia-00370
(normal-paced, fairly steady, no disfluency, formal) Earth's climate has changed dramatically over time.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.2/10; 3.4s, EN.
EN_B00005_S09778_W000007 · in -18.4 dBFS · gain -1.6 dB · emolia-00370
(normal-paced, fairly steady, no disfluency, formal) But these natural changes came gradually.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.6/10; 3.2s, EN.
EN_B00005_S09778_W000008 · in -18.7 dBFS · gain -1.3 dB · emolia-00370
(normal-paced, fairly steady, no disfluency, formal) Scientists worry today because the climate seems to have changed so quickly in the last hundred years.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.7/10; 6.5s, EN.
EN_B00005_S09778_W000009 · in -18.4 dBFS · gain -1.6 dB · emolia-00370
TENS — tensionc-emolia-VN1 · #9

This is a VoiceNet dimension, not an emotion: tension (TENS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with tension (TENS) low — 0.25, lower than 75 % of clips in this corpus — and ends with it around average at 0.49, lower than 51 % of clips in this corpus. That is a total rise of 0.24.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then -0.01, then +0.19, then -0.06 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of -0.18 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.18 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 37 s · de · emolia

k 5d_a 0.239d_b 0.239step_a 0.185step_b 0.185min_cos_consec min_cos_anchor dataset emolialang despeaker DE_Kjh9Gx6t3TQtrack DE_Kjh9Gx6t3TQtotal 37.2slevel spread 2.5 dBmax seam 2.5 dB
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, slightly relaxed, moderate pitch range
(measured, normally alert, fairly steady, monologue) und, äh, (low mumble) bei den Aufgaben, die Bibliotheken haben, methodische Fortentwicklung.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 5.4s, DE.
DE_Kjh9Gx6t3TQ_W000025 · in -18.1 dBFS · gain -1.9 dB · emolia-00118
(slow, subdued, fairly steady, monologue) (low mumble) euh, nämlich, (low mumble) euh, als verlässlicher Ansprechpartner für belastbares Wissen, (ahem) für evidenzbasierungen.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, submissive, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.0/10; 9.8s, DE.
DE_Kjh9Gx6t3TQ_W000026 · in -19.9 dBFS · gain -0.1 dB · emolia-00118
(longing · normal-paced, normally alert, moderately variable, casual) Ich übernehme jetzt mal für Frau Albers, die leider gerade wieder rausgefallen ist aus dem (ahem) Zoom. Möchte sonst noch jemand was zu Citizen Science sagen?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.0/10; 9.3s, DE.
DE_Kjh9Gx6t3TQ_W000027 · in -20.2 dBFS · gain +0.2 dB · emolia-00118
(concentration · normal-paced, normally alert, fairly steady, didactic) Wissenschaftliche Bibliotheken und Privatwirtschaft notwendig, sinnvoll oder verwerflich? Ich denke, wir fangen einfach wieder vorne an, Herr Nelle.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as concentration; style: didactic, formal; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.2/10; 9.1s, DE.
DE_Kjh9Gx6t3TQ_W000028 · in -17.7 dBFS · gain -2.3 dB · emolia-00118
(emotional numbness, doubt, confusion · measured, normally alert, fairly steady, casual) (low mumble) sich einseitig abhängig zu machen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, doubt, confusion; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 0.0/10; 3.0s, DE.
DE_Kjh9Gx6t3TQ_W000029 · in -18.0 dBFS · gain -2.0 dB · emolia-00118
ROUG — roughnessc-emolia-VN1 · #10

This is a VoiceNet dimension, not an emotion: roughness (ROUG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with roughness (ROUG) high — 0.82, higher than 82 % of clips in this corpus — and works its way down to below average at 0.32, lower than 68 % of clips in this corpus. That is a total fall of 0.51.

It takes 4 clips to get there. Clip to clip the moves are -0.05, then -0.21, then -0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 36 s · en · emolia

k 4d_a -0.506d_b -0.506step_a 0.239step_b 0.239min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_6gx9fz6XzeEtrack EN_6gx9fz6XzeEtotal 36.5slevel spread 2.6 dBmax seam 2.6 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, moderate pitch range
(concentration · fairly steady, some disfluency, average clarity, didactic) Okay, the last couple minutes here, I want to talk about some of the ways that (low mumble) uhm, that an active nucleus can change the star formation rates within galaxies.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.7/10; 10.3s, EN.
EN_6gx9fz6XzeE_W000059 · in -17.9 dBFS · gain -2.1 dB · emolia-02575
(interest, relief, concentration · steady, some disfluency, clear, monologue) So you've got all kinds of stuff happening in your active galactic nucleus. You've got your jets, your particles, radiation from the accretion disk, and all these things can impact star formation in both ways. They can increase it and decrease it.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, relief, concentration; style: monologue, casual; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.7/10; 15.0s, EN.
EN_6gx9fz6XzeE_W000060 · in -19.7 dBFS · gain -0.3 dB · emolia-02575
(fairly steady, little disfluency, clear, monologue) If you have the situation where you have cool grass clouds that are compressed by some of the (low mumble) winds from your galactic nucleus,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.2/10; 7.4s, EN.
EN_6gx9fz6XzeE_W000061 · in -17.0 dBFS · gain -3.0 dB · emolia-02575
(fairly steady, no disfluency, clear, casual) Then those gas clouds can collapse and form new stars.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.1/10; 3.3s, EN.
EN_6gx9fz6XzeE_W000062 · in -18.4 dBFS · gain -1.6 dB · emolia-02575
BKGN — background noise levelc-emolia-VN1 · #11

This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with background noise level (BKGN) around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the range at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.40.

It takes 4 clips to get there. Clip to clip the moves are +0.08, then +0.20, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 53 s · en · emolia

k 4d_a 0.397d_b 0.397step_a 0.204step_b 0.204min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00039_S05624track EN_B00039_S05624total 52.8slevel spread 0.6 dBmax seam 0.6 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, normal-paced, normally alert, slightly relaxed, fairly steady, moderate pitch range
(relief, astonishment surprise, fear · some disfluency, average clarity, light breath, casual) There are situations in life where slamming the gas throttle would actually be the safest thing to do. Here is a video of a Toyota FJ rolling over because the driver hesitated and braked at the wrong time in the wrong place.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as relief, astonishment surprise, fear; style: casual, whispered; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 5.0/10; 11.8s, EN.
EN_B00039_S05624_W000017 · in -21.4 dBFS · gain +1.4 dB · emolia-00999
(fear, teasing, interest · almost no disfluency, clear, minimal breath, newsreading) Driving in a straight line is easy, right? Well, if you're three, four, or even five times the factory horsepower in a dedicated drag racing car, things can get hairy really quick. At the Indy 2022 event, the driver of this Ford Mustang almost crashes into Cletus McFarland after hitting a wet patch of asphalt right after the finish line.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as fear, teasing, interest; style: newsreading, formal; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 2.4/10; 18.5s, EN.
EN_B00039_S05624_W000018 · in -20.8 dBFS · gain +0.8 dB · emolia-00999
(emotional numbness, interest, concentration · almost no disfluency, clear, minimal breath, formal) You've probably heard about the catalytic converter apocalypse, where thieves check up your car and simply cut it out. Just like with this Toyota Prius. And there is a reason why Priuses are mainly targeted. Hybrid cars use catalytic converters with a higher concentration of precious metal compared to normal cars. In other words, they are much more expensive.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, minimal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as emotional numbness, interest, concentration; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.9/10; 18.5s, EN.
EN_B00039_S05624_W000019 · in -21.0 dBFS · gain +1.0 dB · emolia-00999
(astonishment surprise, longing, shame · little disfluency, clear, light breath, storytelling) By the way, the same thing happened to my father's car a few years ago.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as astonishment surprise, longing, shame; style: storytelling, casual; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.5/10; 3.6s, EN.
EN_B00039_S05624_W000020 · in -21.0 dBFS · gain +1.0 dB · emolia-00999
VALS — valence stabilityc-emolia-VN1 · #12

This is a VoiceNet dimension, not an emotion: valence stability (VALS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with valence stability (VALS) low — 0.13, lower than 87 % of clips in this corpus — and ends with it above average at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.62.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.18, then +0.14, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 37 s · en · emolia

k 5d_a 0.623d_b 0.623step_a 0.177step_b 0.177min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00017_S00649track EN_B00017_S00649total 37.3slevel spread 1.6 dBmax seam 1.6 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(teasing, contempt, sourness · normal-paced, little disfluency, storytelling, formal) Gandhi's attitude was not that of most western pacifists.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as teasing, contempt, sourness; style: storytelling, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.4/10; 3.6s, EN.
EN_B00017_S00649_W000078 · in -17.1 dBFS · gain -2.9 dB · emolia-00591
(emotional numbness, sadness, fear · measured, almost no disfluency, narration, formal) Satyagraha, first evolved in South Africa, was a sort of non-violent warfare, a way of defeating the enemy without hurting him, and without feeling or arousing hatred.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, sadness, fear; style: narration, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 12.0s, EN.
EN_B00017_S00649_W000079 · in -17.1 dBFS · gain -2.9 dB · emolia-00591
(fear, emotional numbness, sadness · measured, little disfluency, narration, formal) It entails such things as civil disobedience, strikes, lying down in front of railway trains, enduring police charges without running away and without hitting back, and the like.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, emotional numbness, sadness; style: narration, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.4/10; 11.8s, EN.
EN_B00017_S00649_W000080 · in -17.7 dBFS · gain -2.3 dB · emolia-00591
(intoxication altered states of consciousness, teasing · measured, almost no disfluency, conversational, narration) Gandhi objected to passive resistance as a translation of Satyagraha.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, teasing; style: conversational, narration; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.6/10; 5.5s, EN.
EN_B00017_S00649_W000081 · in -18.3 dBFS · gain -1.7 dB · emolia-00591
(normal-paced, almost no disfluency, monologue, formal) In Gujarati, it seems the word means firmness in the truth.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.5/10; 3.7s, EN.
EN_B00017_S00649_W000082 · in -16.7 dBFS · gain -3.3 dB · emolia-00591
S_NARR — style: narrationc-emolia-VN1 · #13

This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: narration (S_NARR) around average — 0.47, lower than 53 % of clips in this corpus — and works its way down to low at 0.17, lower than 83 % of clips in this corpus. That is a total fall of 0.30.

It takes 5 clips to get there. Clip to clip the moves are -0.15, then +0.07, then -0.20, then -0.02 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.63 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.63 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 40 s · en · emolia

k 5d_a -0.299d_b -0.299step_a 0.198step_b 0.198min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_ntrLxf5KiAytrack EN_ntrLxf5KiAytotal 39.9slevel spread 6.7 dBmax seam 3.9 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth
(normal-paced, normally alert, slightly relaxed, monologue) The real time execution of, of this model on the Opal RT simulator.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 1.4/10; 4.5s, EN.
EN_ntrLxf5KiAy_W000384 · in -19.4 dBFS · gain -0.6 dB · emolia-01516
(normal-paced, normally alert, slightly relaxed, casual) I can just mention that this is the, the, actually the first page view of the RTLab. So when you build the model on the Simulink, then you should...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.6/10; 8.1s, EN.
EN_ntrLxf5KiAy_W000388 · in -19.1 dBFS · gain -0.9 dB · emolia-01516
(brisk, normally alert, slightly relaxed, monologue) (low mumble) And then you need this RT lab to be able to communicate with the Opal RT platform.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, formal; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.9/10; 7.5s, EN.
EN_ntrLxf5KiAy_W000389 · in -18.9 dBFS · gain -1.1 dB · emolia-01516
(intoxication altered states of consciousness, pain, concentration · measured, subdued, relaxed, monologue) Yes, this is the model file, uh, (ahem) (low mumble) (low mumble) uh, in that model it contains, (low mumble) uh, eevee module and (low mumble) battery. I will show you that file.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is neutral, submissive, slightly guarded; reads as intoxication altered states of consciousness, pain, concentration; style: monologue, didactic; below-average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.8/10; 12.0s, EN.
EN_ntrLxf5KiAy_W000390 · in -15.0 dBFS · gain -5.0 dB · emolia-01516
(pain, fear · measured, subdued, relaxed, monologue) And (low mumble) yeah, this is the real time simulation file here in this in main.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, steady; timbre is slightly cool, dark, fairly smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is neutral, submissive, neutral openness; reads as pain, fear; style: monologue, casual; below-average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.3/10; 7.2s, EN.
EN_ntrLxf5KiAy_W000391 · in -12.7 dBFS · gain -7.3 dB · emolia-01516
TENS — tensionc-emolia-VN1 · #14

This is a VoiceNet dimension, not an emotion: tension (TENS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with tension (TENS) low — 0.16, lower than 84 % of clips in this corpus — and ends with it high at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.64.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.18, then +0.01, then +0.21 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 76 s · fr · emolia

k 5d_a 0.638d_b 0.638step_a 0.239step_b 0.239min_cos_consec min_cos_anchor dataset emolialang frspeaker FR_YnSvmbWx_YQtrack FR_YnSvmbWx_YQtotal 75.6slevel spread 2.0 dBmax seam 2.0 dB
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, some disfluency, light breath
(sourness, jealousy and envy, impatience and irritability · brisk, energised, neutral tension, dramatic) ça te retardera pas ton raid, mais ça va pas te porter ton raid non plus. Ce qui explique qu'il est en A. C'est-à-dire que tu vas pouvoir quand même faire des trucs assez sexy, vu que t'es un voleur, t'as une mobilité qui est correcte, t'as des bons dégâts, qui sont bien. Mais, bah, tja, c'est très dur à jouer, donc c'est quand même réservé (chuckle) à une niche, hein, le voleur ornadoire. Ça demande de beaucoup de spam et de spam intelligemment. Et c'est juste que, voilà, ça brise sur aucun fight, mais c'est nul sur aucun fight. Donc, voilà, c'est un A.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, jealousy and envy, impatience and irritability; style: dramatic, conversational; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 7.9/10; 30.0s, FR.
FR_YnSvmbWx_YQ_W000012 · in -22.8 dBFS · gain +2.8 dB · emolia-02836
(confusion · brisk, normally alert, slightly relaxed, casual) par contre, ouais, ça reste en sim, ça reste dans l'optique où t'arrives à être dans le dos du monstre, et ça reste également dans l'optique où t'arrives à avoir un uptime permanent sur le monstre. Et dès que tu vas perdre de l'uptime en finesse, tu vas vite, euh, (low mumble) tu vas vite en perdre beaucoup.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: casual, dramatic; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 5.6/10; 11.3s, FR.
FR_YnSvmbWx_YQ_W000013 · in -20.8 dBFS · gain +0.8 dB · emolia-02836
(fear, disgust · fast, energised, slightly relaxed, dramatic) Tu vas vite perdre beaucoup de dégâts. D'autant plus que tu as des bursts qui sont fréquents. C'est-à-dire qu'en gros, toutes les 30 secondes, tu vas à burst.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as fear, disgust; style: dramatic, storytelling; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 3.9/10; 6.3s, FR.
FR_YnSvmbWx_YQ_W000014 · in -22.7 dBFS · gain +2.7 dB · emolia-02836
(jealousy and envy, sourness, disgust · brisk, energised, neutral tension, dramatic) Et si tu burstes toutes les 30 secondes, ça implique de, euh, (low mumble) pouvoir toujours être là, être contre le contact du boss et pouvoir taper. C'est-à-dire que si tu vas faire des mécaniques sur des timings où par exemple, tes symboles sont revenus, tu vas vite perdre énormément en DPS au niveau de ton raid.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, sourness, disgust; style: dramatic, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.3/10; 13.2s, FR.
FR_YnSvmbWx_YQ_W000015 · in -20.8 dBFS · gain +0.8 dB · emolia-02836
(disgust, sourness, astonishment surprise · brisk, energised, neutral tension, casual) Donc, euh, (low mumble) voilà, voleur finesse peut vite perdre. Par contre, dès que tu vas avoir des adds, dès que t'as, genre, 5 adds ou plus, là, le DPS, il est intestable. C'est monstrueux. Un, t'as aucun cooldown sur tes dégâts de zone. Ça, c'est, c'est trop fort, ça.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disgust, sourness, astonishment surprise; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.2/10; 14.2s, FR.
FR_YnSvmbWx_YQ_W000016 · in -21.7 dBFS · gain +1.7 dB · emolia-02836
VOLT — loudness / volumec-emolia-VN1 · #15

This is a VoiceNet dimension, not an emotion: loudness / volume (VOLT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with loudness / volume (VOLT) around average — 0.56, higher than 56 % of clips in this corpus — and works its way down to low at 0.25, lower than 75 % of clips in this corpus. That is a total fall of 0.31.

It takes 3 clips to get there. Clip to clip the moves are -0.16, then -0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 54 s · en · emolia

k 3d_a -0.311d_b -0.311step_a 0.161step_b 0.161min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00048_S00998track EN_B00048_S00998total 54.0slevel spread 1.2 dBmax seam 1.2 dB
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, slow, very low-energy, relaxed
(sadness, teasing, jealousy and envy · fairly narrow pitch, whispered, monologue) And he's always wearing the coat until he is murdered and the coat gets stolen and things in the plot get very difficult for him. (low mumble) Uhm, Old Bailey lives on the top.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; reads as sadness, teasing, jealousy and envy; style: whispered, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_B00048_S00998_W000121 · in -20.1 dBFS · gain +0.1 dB · emolia-01182
(interest, contemplation, doubt · narrow pitch range, casual, ASMR) Of roofs and up on the rooftops, he's dressed in a cape of feathers. He's described as looking like what Robinson Crusoe in the popular imagination would have looked like if he were cast away on the rooftops of London with nothing but birds to create his world out of. (low mumble) Uhm, so you take these characters and
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, slightly guarded; reads as interest, contemplation, doubt; style: casual, ASMR; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.0/10; 23.8s, EN.
EN_B00048_S00998_W000122 · in -20.7 dBFS · gain +0.7 dB · emolia-01182
(contemplation, relief, pain · fairly narrow pitch, monologue, whispered) Securing the knowledge that none of them could be confused for any of the others. None of them talk the same. And when they turn up, (low mumble) uhm, they, you know who they are. The reader does not have to work at that thing.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as contemplation, relief, pain; style: monologue, whispered; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.8/10; 18.4s, EN.
EN_B00048_S00998_W000123 · in -19.5 dBFS · gain -0.5 dB · emolia-01182
R_NASL — resonance: nasalc-emolia-VN1 · #16

This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: nasal (R_NASL) around average — 0.50, right about the corpus median — and works its way down to at the very bottom of the range at 0.07, lower than 93 % of clips in this corpus. That is a total fall of 0.44.

It takes 5 clips to get there. Clip to clip the moves are -0.21, then -0.16, then +0.09, then -0.15 — not a clean run: step 3 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 45 s · en · emolia

k 5d_a -0.436d_b -0.436step_a 0.214step_b 0.214min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_vZgy7NN7z2utrack EN_vZgy7NN7z2utotal 44.9slevel spread 1.9 dBmax seam 1.7 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, slightly relaxed, no disfluency, clear
(awe · measured, normally alert, steady, formal) Iron can have various redox and spin states, and it can be held in many stereochemistries
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.5s, EN.
EN_vZgy7NN7z2u_W000091 · in -13.9 dBFS · gain -6.1 dB · emolia-01845
(emotional numbness · normal-paced, normally alert, steady, formal) Around 4–3 Ga, anaerobic prokaryotes began developing metal and organic cofactors for light absorption
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 7.8s, EN.
EN_vZgy7NN7z2u_W000093 · in -14.5 dBFS · gain -5.5 dB · emolia-01845
(normal-paced, normally alert, steady, formal) They ultimately ended up making chlorophyll from Mg2, as is found in cyanobacteria and plants, leading to modern photosynthesis
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.8s, EN.
EN_vZgy7NN7z2u_W000094 · in -13.8 dBFS · gain -6.2 dB · emolia-01845
(normal-paced, energised, fairly steady, formal) However, chlorophyll synthesis requires numerous steps
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 3.9s, EN.
EN_vZgy7NN7z2u_W000095 · in -14.1 dBFS · gain -5.9 dB · emolia-01845
(emotional numbness · normal-paced, normally alert, steady, formal) The process starts with uroporphyrin, a primitive precursor to the porphyrin ring which may be biotic or abiotic in origin, which is then modified in cells differently to make Mg, Fe, Ni, Ni, and Cobalt-Co complexes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.3s, EN.
EN_vZgy7NN7z2u_W000096 · in -15.8 dBFS · gain -4.2 dB · emolia-01845
TENS — tensionc-emolia-VN1 · #17

This is a VoiceNet dimension, not an emotion: tension (TENS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with tension (TENS) around average — 0.54, higher than 54 % of clips in this corpus — and works its way down to low at 0.24, lower than 76 % of clips in this corpus. That is a total fall of 0.30.

It takes 3 clips to get there. Clip to clip the moves are -0.08, then -0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · zh · emolia

k 3d_a -0.303d_b -0.303step_a 0.223step_b 0.223min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00061_S03037track ZH_B00061_S03037total 17.9slevel spread 0.8 dBmax seam 0.8 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(normal-paced, fairly steady, formal, narration) 政府希望人们开办企业来创造大量的工作机会。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.8/10; 4.1s, ZH.
ZH_B00061_S03037_W000039 · in -19.0 dBFS · gain -1.0 dB · emolia-03889
(confusion, doubt · measured, fairly steady, narration, monologue) 因此为企业主提供了税收优惠刺激,在投资方面也是一样。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, doubt; style: narration, monologue; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.6/10; 6.1s, ZH.
ZH_B00061_S03037_W000040 · in -18.9 dBFS · gain -1.1 dB · emolia-03889
(measured, steady, monologue, narration) 政府为了吸引人们投资在政府发展的项目,例如房屋项目上而提供税收优惠。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.1/10; 7.4s, ZH.
ZH_B00061_S03037_W000041 · in -19.7 dBFS · gain -0.3 dB · emolia-03889
METL — metallic qualityc-emolia-VN1 · #18

This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with metallic quality (METL) at the very top of the range — 0.91, higher than 91 % of clips in this corpus — and works its way down to around average at 0.55, higher than 55 % of clips in this corpus. That is a total fall of 0.35.

It takes 3 clips to get there. Clip to clip the moves are -0.20, then -0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 35 s · zh · emolia

k 3d_a -0.352d_b -0.352step_a 0.205step_b 0.205min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00041_S08295track ZH_B00041_S08295total 34.8slevel spread 2.1 dBmax seam 2.1 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, fast, some disfluency, average clarity
(pride · energised, neutral tension, moderately variable, casual) 用一个类似于公交车的场景去错位到咱们这个机场的场景里来。虽然说庞博遇到了这个事儿呃。 (low mumble)
full caption & clip details
A child masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as pride; style: casual, dramatic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.1/10; 9.2s, ZH.
ZH_B00041_S08295_W000017 · in -14.0 dBFS · gain -6.0 dB · emolia-03687
(normally alert, slightly relaxed, fairly steady, storytelling) 到这儿都还是好笑的,然后最后到最后要升华了。不过在那个时刻,我真的听到了一个声音,说我要去上海,我仔细听了一下,是十八岁的我自己。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, dramatic; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 7.0/10; 9.3s, ZH.
ZH_B00041_S08295_W000018 · in -16.1 dBFS · gain -3.9 dB · emolia-03687
(interest, elation, hope enthusiasm optimism · normally alert, neutral tension, moderately variable, casual) 这个文本说完了,就是我会在那个各个网站上啊,包括知乎啊、微博呀、小红书啊,看一看这个大家对这个脱口秀大会的反响。我发现很多人都在反映一个问题,说这一季呢不是这一季啊,就这两季,因为只有这两期是主题赛啊。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, elation, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 8.0/10; 16.0s, ZH.
ZH_B00041_S08295_W000019 · in -15.4 dBFS · gain -4.6 dB · emolia-03687
EMPH — emphasis / stress strengthc-emolia-VN1 · #19

This is a VoiceNet dimension, not an emotion: emphasis / stress strength (EMPH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with emphasis / stress strength (EMPH) high — 0.80, higher than 80 % of clips in this corpus — and works its way down to below average at 0.30, lower than 70 % of clips in this corpus. That is a total fall of 0.50.

It takes 5 clips to get there. Clip to clip the moves are -0.19, then -0.14, then +0.00, then -0.17 — not a clean run: step 3 moves back the other way by 0.00 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · en · emolia

k 5d_a -0.503d_b -0.503step_a 0.191step_b 0.191min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_dm541P-nrWEtrack EN_dm541P-nrWEtotal 46.6slevel spread 3.6 dBmax seam 2.8 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, quiet background, slightly relaxed, fairly steady, moderate pitch range, light breath
(thankfulness gratitude, contentment, elation · normal-paced, normally alert, some disfluency, casual) Yeah, and thanks to our audience for sharing your thoughts on this. (low mumble) So yeah, so back to you, Curtis. We're about to dive into the journey of your research, so can you, (low mumble) uh,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment, elation; style: casual, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.6/10; 9.8s, EN.
EN_dm541P-nrWE_W000159 · in -17.7 dBFS · gain -2.3 dB · emolia-02173
(normal-paced, normally alert, some disfluency, casual) Tell us a little bit about, (low mumble) uh, more of the specifics about the work that you do.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.6/10; 3.1s, EN.
EN_dm541P-nrWE_W000160 · in -16.4 dBFS · gain -3.6 dB · emolia-02173
(shame, sadness, fatigue exhaustion · measured, normally alert, frequent disfluency, casual) With the slide I left, this concept of weather severity and disease. And so,
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, sadness, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.5/10; 5.4s, EN.
EN_dm541P-nrWE_W000162 · in -18.2 dBFS · gain -1.8 dB · emolia-02173
(normal-paced, normally alert, some disfluency, didactic) Generally summed up, weather severity indices are a way in which we can quantify the weather conditions, their severity, their impact.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.6/10; 7.8s, EN.
EN_dm541P-nrWE_W000163 · in -17.3 dBFS · gain -2.7 dB · emolia-02173
(concentration, interest, contemplation · normal-paced, subdued, some disfluency, monologue) For a variety of applications, this can be for risk communication. It could be for decision support. It can be for emergency planning or recovery and resilience. And so let's look at some of the weather severity indices that folks might be familiar with already. And so really when we look at this severity index space,
full caption & clip details
A middle-aged masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest, contemplation; style: monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.0/10; 19.8s, EN.
EN_dm541P-nrWE_W000164 · in -20.1 dBFS · gain +0.1 dB · emolia-02173
R_CHST — resonance: chestc-emolia-VN1 · #20

This is a VoiceNet dimension, not an emotion: resonance: chest (R_CHST) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.

The chain starts with resonance: chest (R_CHST) at the very bottom of the range — 0.01, lower than 99 % of clips in this corpus — and ends with it above average at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.65.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.22, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 28 s · en · emolia

k 4d_a 0.645d_b 0.645step_a 0.232step_b 0.232min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00014_S08482track EN_B00014_S08482total 28.3slevel spread 1.0 dBmax seam 0.7 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · neutral-toned, neutral-bright, balanced body, good recording, no background noise, measured, clear, light breath
(confusion, doubt, affection · energised, neutral tension, moderately variable, storytelling) But be very careful. Remember, she's a wild animal. She might be so upset and confused that she could attack you by mistake. We'll remember. Come on, Jack, said Annie. Oh, wait.
full caption & clip details
A child feminine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is negative, slightly submissive, neutral openness; reads as confusion, doubt, affection; style: storytelling, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.8/10; 13.1s, EN.
EN_B00014_S08482_W000268 · in -17.1 dBFS · gain -2.9 dB · emolia-00522
(malevolence malice, teasing, pain · normally alert, slightly relaxed, moderately variable, narration) Said Dr. Ling. So everyone will know you are here to help. Please put your volunteer clothes back on.
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as malevolence malice, teasing, pain; style: narration, storytelling; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 1.6/10; 6.0s, EN.
EN_B00014_S08482_W000269 · in -17.3 dBFS · gain -2.7 dB · emolia-00522
(pain, emotional numbness · very low-energy, slightly relaxed, fairly steady, narration) She pointed to the bin, then she and Master Li hurried off.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pain, emotional numbness; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 6.5/10; 3.7s, EN.
EN_B00014_S08482_W000270 · in -17.4 dBFS · gain -2.6 dB · emolia-00522
(normally alert, slightly relaxed, steady, narration) Jack and Annie grab their coveralls from the bin and pull them on over their wet clothes.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.9/10; 5.1s, EN.
EN_B00014_S08482_W000271 · in -18.1 dBFS · gain -1.9 dB · emolia-00522