VN1 at chain length k=5, all corpora, at the mining floor.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. rule x chain length Sampled from 1,334,387 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
R_THRT — resonance: throat ↓k-VN1-k5 · #1
This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: throat (R_THRT) above average — 0.74, higher than 74 % of clips in this corpus — and works its way down to around average at 0.44, lower than 56 % of clips in this corpus. That is a total fall of 0.30.
It takes 5 clips to get there. Clip to clip the moves are -0.11, then +0.13, then -0.12, then -0.20 — not a clean run: step 2 moves back the other way by 0.13 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 41 s · en · emolia
hear it un-normalised (raw levels, max seam 1.1 dB)
k 5d_a -0.303d_b -0.303step_a 0.203step_b 0.203min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_ew79xrISCpctrack EN_ew79xrISCpctotal 41.2slevel spread 1.4 dBmax seam 1.1 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, clear
(normal-paced, fairly steady, no disfluency, formal)Given a vector V in Euclidean space N, the formula for the reflection in the hyperplane through the origin, orthogonal to A, is given by
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.7/10; 10.7s, EN.
EN_ew79xrISCpc_W000034 · in -18.2 dBFS · gain -1.8 dB · emolia-01686
(concentration· normal-paced, fairly steady, no disfluency, formal)Note that the second term in the above equation is just twice
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.8/10; 7.5s, EN.
EN_ew79xrISCpc_W000037 · in -18.1 dBFS · gain -1.9 dB · emolia-01686
(confusion· normal-paced, fairly steady, no disfluency, formal)Rifa' v' equals −v, if v is parallel to a, and
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.8/10; 4.2s, EN.
EN_ew79xrISCpc_W000038 · in -17.7 dBFS · gain -2.3 dB · emolia-01686
(awe, emotional numbness· normal-paced, fairly steady, no disfluency, formal)Since these reflections are isometries of Euclidean space fixing the origin they
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.2/10; 7.3s, EN.
EN_ew79xrISCpc_W000042 · in -18.0 dBFS · gain -2.0 dB · emolia-01686
(concentration·measured, steady, frequent disfluency, formal)The orthogonal matrix corresponding to the above reflection is
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, authoritative; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.2/10; 10.9s, EN.
EN_ew79xrISCpc_W000043 · in -19.1 dBFS · gain -0.9 dB · emolia-01686
This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with vocal flexibility / inflection (VFLX) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to around average at 0.48, lower than 52 % of clips in this corpus. That is a total fall of 0.46.
It takes 5 clips to get there. Clip to clip the moves are -0.07, then -0.04, then -0.20, then -0.16 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · en · emolia
hear it un-normalised (raw levels, max seam 2.0 dB)
k 5d_a -0.459d_b -0.459step_a 0.197step_b 0.197min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00006_S07307track EN_B00006_S07307total 39.8slevel spread 2.3 dBmax seam 2.0 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, moderately variable, some disfluency, average clarity, wide pitch range, light breath
(disgust, sourness, elation · brisk, energised, slightly relaxed, casual)It's more earth-friendly too, by the way. It's sourced from renewable forest, consumes ten times less water to grow, and it's an ultra-smooth, waste-free production process.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disgust, sourness, elation; style: casual, playful; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.8/10; 6.9s, EN.
EN_B00006_S07307_W000085 · in -19.1 dBFS · gain -0.8 dB · emolia-00386
(elation, hope enthusiasm optimism, pleasure ecstasy· brisk, energised, neutral tension, casual)20,000 five star reviews, right Bob? It's soft, it's like angel skin, man. Angel skin, that should be the new tag. I love it, yeaah, buddy. Buffy, it's like angel skin. They're offering a free trial, free shipping, and free returns every day. We joke around a lot, but we mean it. This is, it's so phenomenal. It's the best comforter. You have to do it. You can try it for free, alright, before you commit to buying. If you don't love it,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 4.7/6; vocal-burst blend 9.7/10; 17.3s, EN.
EN_B00006_S07307_W000086 · in -18.3 dBFS · gain -1.7 dB · emolia-00386
(elation, pleasure ecstasy, teasing·normal-paced, normally alert, neutral tension, casual)Yeah, I'm excited. The clips by the way look great. Of you eating the bug was very funny. There's some really good stuff.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as elation, pleasure ecstasy, teasing; style: casual, conversational; good recording, quiet background; genuineness 5.3/6; vocal-burst blend 1.4/10; 6.3s, EN.
EN_B00006_S07307_W000087 · in -18.6 dBFS · gain -1.4 dB · emolia-00386
(pain, sadness, bitterness· normal-paced, normally alert, neutral tension, casual)Money. He did it because of money. Money.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; style: casual, conversational; average recording, no background noise; reads as pain, sadness, bitterness; genuineness 3.9/6; vocal-burst blend 4.5/10; 3.4s, EN.
EN_B00006_S07307_W000088 · in -20.6 dBFS · gain +0.6 dB · emolia-00386
(teasing, doubt, sexual lust· normal-paced, normally alert, neutral tension, conversational)You're gonna be funny on it, Bob. You know you are. I don't think so. Have you seen any of it? (ahem)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as teasing, doubt, sexual lust; style: conversational, casual; average recording, no background noise; mildly explicit content; genuineness 3.4/6; vocal-burst blend 0.6/10; 5.3s, EN.
EN_B00006_S07307_W000089 · in -19.7 dBFS · gain -0.3 dB · emolia-00386
DFLU — disfluency ↓k-VN1-k5 · #3
This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with disfluency (DFLU) around average — 0.45, lower than 55 % of clips in this corpus — and works its way down to low at 0.13, lower than 87 % of clips in this corpus. That is a total fall of 0.32.
It takes 5 clips to get there. Clip to clip the moves are -0.23, then +0.15, then -0.09, then -0.15 — not a clean run: step 2 moves back the other way by 0.15 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · zh · emolia
hear it un-normalised (raw levels, max seam 1.4 dB)
k 5d_a -0.323d_b -0.323step_a 0.233step_b 0.233min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00072_S09201track ZH_B00072_S09201total 40.3slevel spread 1.4 dBmax seam 1.4 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, some disfluency, average clarity, conversational)这人咽气了,但是也不能让他就这么坐着呀。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, storytelling; average recording, no background noise; genuineness 4.3/6; vocal-burst blend 2.5/10; 3.9s, ZH.
ZH_B00072_S09201_W000011 · in -16.8 dBFS · gain -3.2 dB · emolia-03993
(awe, relief, longing·measured, some disfluency, average clarity, storytelling)当我们试图把它抬出来的时候,那吸管好像有生命力一样,一下子就缩了回去。几乎同时,这里响起了哀乐。
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as awe, relief, longing; style: storytelling, narration; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.7/10; 10.5s, ZH.
ZH_B00072_S09201_W000012 · in -15.4 dBFS · gain -4.6 dB · emolia-03993
This is a VoiceNet dimension, not an emotion: audible breath / respiration (RESP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with audible breath / respiration (RESP) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it around average at 0.52, higher than 52 % of clips in this corpus. That is a total rise of 0.50.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.07, then +0.10, then +0.15 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 66 s · en · podcast
hear it un-normalised (raw levels, max seam 2.5 dB)
k 5d_a 0.502d_b 0.502step_a 0.173step_b 0.173min_cos_consec 0.7153min_cos_anchor 0.5345dataset podcastlang enspeaker 33775track 33775total 66.0slevel spread 2.5 dBmax seam 2.5 dBcos from recomputed from spkemb_traj
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, average clarity
(emotional numbness, relief, embarrassment · normal-paced, neutral tension, moderately variable, casual)I have a No Effects record on my list. However, going through my chronological.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness, relief, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 3.8/10; 4.6s, EN.
33775_00120032 · in -17.6 dBFS · gain -2.4 dB · podcast-04456
(pride, triumph· normal-paced, neutral tension, moderately variable, conversational)it will be. So my second record is Nirvana's Nevermind. What?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as pride, triumph; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 0.0/10; 8.8s, EN.
33775_00120687 · in -20.1 dBFS · gain +0.1 dB · podcast-04466
(contempt, teasing, impatience and irritability·slow, slightly relaxed, fairly steady, conversational)Okay, because I'm the youngest of three boys.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contempt, teasing, impatience and irritability; style: conversational, casual; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 0.0/10; 5.2s, EN.
33775_00121792 · in -19.4 dBFS · gain -0.6 dB · podcast-04467
(infatuation, sexual lust, longing·normal-paced, neutral tension, moderately variable, casual)my older brothers loved Guns N' Roses and all of like the hair poison, all the hair metal at the time. Then they leaked into like Snoop Dogg and Dr. Dre, which I liked as well. Then I heard Nirvana. And that was like mine. That like put me in my own little thing.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as infatuation, sexual lust, longing; style: casual, storytelling; average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 5.9/10; 22.4s, EN.
33775_00122456 · in -18.3 dBFS · gain -1.7 dB · podcast-06120
(disappointment, sourness, intoxication altered states of consciousness· normal-paced, neutral tension, moderately variable, casual)And every song was just different. My parents hated it, which was great. (exhausted groan) Because it just made it it it gave me my pocket away from now. Everybody heard smells like Teen Spirit, right? But uh (low mumble) at the time, that whole record just changed everything. It changed the music industry. It killed hair metal. It killed poison and quiet riot.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as disappointment, sourness, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 4.8/10; 24.4s, EN.
33775_00124692 · in -18.3 dBFS · gain -1.7 dB · podcast-00789
S_NARR — style: narration ↓k-VN1-k5 · #5
This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: narration (S_NARR) above average — 0.71, higher than 71 % of clips in this corpus — and works its way down to below average at 0.38, lower than 62 % of clips in this corpus. That is a total fall of 0.33.
It takes 5 clips to get there. Clip to clip the moves are -0.15, then -0.08, then -0.05, then -0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 48 s · fr · emolia
hear it un-normalised (raw levels, max seam 3.0 dB)
k 5d_a -0.326d_b -0.326step_a 0.154step_b 0.154min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_Jsclp5CNqwItrack FR_Jsclp5CNqwItotal 47.9slevel spread 3.0 dBmax seam 3.0 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · fairly smooth, quiet background, normally alert, wide pitch range
(fear, distress · normal-paced, slightly relaxed, moderately variable, dramatic)et réactiver chaque fois qu'on a une situation qui va toucher cette blessure, ben, la blessure se réactive.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as fear, distress; style: dramatic, playful; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.2/10; 7.8s, FR.
FR_Jsclp5CNqwI_W000053 · in -20.0 dBFS · gain +0.0 dB · emolia-02726
(pride, distress, relief·brisk, slightly relaxed, moderately variable, dramatic)forcément, je vais vivre, cette blessure réactivée fait que moi, en tant qu'adulte, je vais être confronté à cette situation. Je vous donne un exemple. J'ai grandi avec un père qui était alcoolique et j'étais une intolérante.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, guarded; reads as pride, distress, relief; style: dramatic, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.8/10; 15.5s, FR.
FR_Jsclp5CNqwI_W000054 · in -20.6 dBFS · gain +0.6 dB · emolia-02726
(emotional numbness, disappointment, bitterness· brisk, neutral tension, moderately variable, dramatic)J'étais intolérante. Faut dire que c'est comme ça. C'est-à-dire que dès que je voyais quelqu'un boire, la limite pour moi était très fine. Dès lors que je voyais, (surprised gasp) rien n'était proportionnel ni logique, en fait. Voilà.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as emotional numbness, disappointment, bitterness; style: dramatic, cartoonish; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.3/10; 13.2s, FR.
FR_Jsclp5CNqwI_W000055 · in -19.7 dBFS · gain -0.3 dB · emolia-02726
(contemplation· brisk, slightly relaxed, moderately variable, dramatic)l'intensité de, l'intensité à la manière à laquelle j'ai vécu les choses, bah, dès que je vois quelqu'un me voir, ça y est, j'y suis.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contemplation; style: dramatic, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.3/10; 6.5s, FR.
FR_Jsclp5CNqwI_W000056 · in -18.1 dBFS · gain -1.9 dB · emolia-02726
(impatience and irritability, contempt, helplessness· brisk, slightly relaxed, fairly steady, dramatic)J'y suis, quoi. Donc moi, il faut que je parte de la salle. C'était hyper compliqué à gérer.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, contempt, helplessness; style: dramatic, authoritative; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.6/10; 4.3s, FR.
FR_Jsclp5CNqwI_W000057 · in -21.1 dBFS · gain +1.1 dB · emolia-02726
S_NARR — style: narration ↑k-VN1-k5 · #6
This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: narration (S_NARR) low — 0.25, lower than 75 % of clips in this corpus — and ends with it around average at 0.56, higher than 56 % of clips in this corpus. That is a total rise of 0.32.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.12, then +0.06, then -0.00 — not a clean run: step 4 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.86 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.86 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 46 s · en · emolia
k 5d_a 0.317d_b 0.317step_a 0.137step_b 0.137min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_4yOVRgU2ND4track EN_4yOVRgU2ND4total 46.4slevel spread 1.5 dBmax seam 1.5 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, some disfluency
(pleasure ecstasy, affection, elation · slightly relaxed, fairly steady, conversational, casual)The King 2 is really, (low mumble) um, I know we talked about this previously is, (ahem) um, I had a session with a family, (low mumble) um, at the start of the year and.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pleasure ecstasy, affection, elation; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.7/10; 8.1s, EN.
EN_4yOVRgU2ND4_W000175 · in -16.2 dBFS · gain -3.8 dB · emolia-00367
(fatigue exhaustion, contentment, disappointment·neutral tension, moderately variable, casual, conversational)There was a lot of, she was feeling quite overwhelmed. And so we just started, got a piece of paper and I just started writing on the table and using her words. So just writing out all the kind of things happening and drawing on what Jackie was saying. We just started to prioritize those really key things and then writing out the steps.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, contentment, disappointment; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 10.0/10; 16.0s, EN.
EN_4yOVRgU2ND4_W000176 · in -16.8 dBFS · gain -3.2 dB · emolia-00367
(affection, triumph, pleasure ecstasy·slightly relaxed, fairly steady, conversational, casual)And anyway, at the end, we had a really clear goal between the visits and it was often about three months until I was gonna see them again.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, triumph, pleasure ecstasy; style: conversational, casual; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.4/10; 5.8s, EN.
EN_4yOVRgU2ND4_W000177 · in -17.6 dBFS · gain -2.4 dB · emolia-00367
(affection, amusement, infatuation· slightly relaxed, fairly steady, conversational, authoritative)So I said, I'll, I'll send you an email with the information. (low mumble) Um, and she's like, no, do you mind if I take that piece of paper with me?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, amusement, infatuation; style: conversational, authoritative; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.7/10; 5.5s, EN.
EN_4yOVRgU2ND4_W000178 · in -16.1 dBFS · gain -4.0 dB · emolia-00367
(astonishment surprise, embarrassment· slightly relaxed, fairly steady, casual, conversational)So I went off and wrote my, my report, you know, it's been an hour and a half. And, you know, it was the kind of words in the report. And next time I saw her, I said, Oh, how was, how was everything and how was the report?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise, embarrassment; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 8.5/10; 10.4s, EN.
EN_4yOVRgU2ND4_W000179 · in -16.9 dBFS · gain -3.0 dB · emolia-00367
EXPL — expressiveness ↑k-VN1-k5 · #7
This is a VoiceNet dimension, not an emotion: expressiveness (EXPL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with expressiveness (EXPL) low — 0.18, lower than 82 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.51.
It takes 5 clips to get there. Clip to clip the moves are +0.07, then +0.10, then +0.19, then +0.15 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · en · emolia
k 5d_a 0.515d_b 0.515step_a 0.192step_b 0.192min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_q2fBrZuFT5ytrack EN_q2fBrZuFT5ytotal 51.2slevel spread 1.7 dBmax seam 1.6 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · fairly steady, moderate pitch range, light breath
(contemplation, longing · measured, normally alert, slightly relaxed, formal)And a sentence starter here is a place I want to focus my energy is.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, longing; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.8/10; 5.2s, EN.
EN_q2fBrZuFT5y_W000458 · in -17.1 dBFS · gain -2.9 dB · emolia-01935
(contentment, affection, shame·normal-paced, normally alert, slightly relaxed, didactic)So I'm introducing a practice. It's a way of having a conversation with ourselves. Can also be a conversation that happens between people where every day we say, okay, for supporting me to live, I give thanks to.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, affection, shame; style: didactic, monologue; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.3/10; 12.6s, EN.
EN_q2fBrZuFT5y_W000459 · in -17.0 dBFS · gain -3.0 dB · emolia-01935
(contemplation, awe, relief·slow, very low-energy, relaxed, conversational)When I look out at the world, what concerns me is, and feelings that come up about this are, what happens through me is,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, awe, relief; style: conversational, casual; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.9/10; 8.3s, EN.
EN_q2fBrZuFT5y_W000460 · in -15.4 dBFS · gain -4.6 dB · emolia-01935
(concentration, contemplation, contentment·normal-paced, very low-energy, slightly relaxed, didactic)And in order to influence and shape what happened through me is a place I want to focus my energy is. And through that what happens is we give our gift of active hope. Our gift of active hope is our contribution to this larger story of change.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, contemplation, contentment; style: didactic, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.0/10; 15.2s, EN.
EN_q2fBrZuFT5y_W000461 · in -16.0 dBFS · gain -4.0 dB · emolia-01935
(thankfulness gratitude, hope enthusiasm optimism, fatigue exhaustion·measured, normally alert, slightly relaxed, whispered)We don't have time to wait. We do have time to act. The story of ActiveHope, the practice of ActiveHope can happen through us all. Thank you.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, full; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, hope enthusiasm optimism, fatigue exhaustion; style: whispered, formal; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.1/10; 9.4s, EN.
EN_q2fBrZuFT5y_W000462 · in -15.9 dBFS · gain -4.1 dB · emolia-01935
EXPL — expressiveness ↑k-VN1-k5 · #8
This is a VoiceNet dimension, not an emotion: expressiveness (EXPL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with expressiveness (EXPL) low — 0.15, lower than 85 % of clips in this corpus — and ends with it around average at 0.43, lower than 57 % of clips in this corpus. That is a total rise of 0.28.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then -0.14, then +0.03, then +0.24 — not a clean run: step 2 moves back the other way by 0.14 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 43 s · zh · emolia
k 5d_a 0.276d_b 0.276step_a 0.242step_b 0.242min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00023_S08790track ZH_B00023_S08790total 42.7slevel spread 1.1 dBmax seam 0.7 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, no background noise, measured, normally alert, slightly relaxed, light breath
(steady, no disfluency, clear, monologue)李逵不好意思的交出了银子,嘴里小声嘀咕道。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.7/10; 5.3s, ZH.
ZH_B00023_S08790_W000012 · in -20.1 dBFS · gain +0.1 dB · emolia-03502
(fairly steady, some disfluency, average clarity, casual)我才不稀罕这几个钱,只是这银子是大哥的,怎么能给他呀?
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, storytelling; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 5.1/10; 7.1s, ZH.
ZH_B00023_S08790_W000013 · in -20.7 dBFS · gain +0.7 dB · emolia-03502
(steady, no disfluency, crisply articulate, ASMR)李逵跟着宋江戴宗到琵琶亭继续喝酒。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: ASMR, whispered; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.4/10; 5.3s, ZH.
ZH_B00023_S08790_W000014 · in -20.2 dBFS · gain +0.2 dB · emolia-03502
(steady, no disfluency, clear, formal)李逵把三份鱼汤三斤牛肉都吃完后,这才发现宋江没怎么动筷子。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.6/10; 8.0s, ZH.
ZH_B00023_S08790_W000015 · in -19.5 dBFS · gain -0.5 dB · emolia-03502
(triumph, pride· steady, no disfluency, clear, didactic)他想,一定是宋江咸鱼不够新鲜,就自告奋勇下楼去。渔船上要些新鲜的鱼来江边,停了十多条渔船,李逵向渔夫要鱼。
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, pride; style: didactic, monologue; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.0/10; 16.4s, ZH.
ZH_B00023_S08790_W000016 · in -19.7 dBFS · gain -0.3 dB · emolia-03502
TEMP — tempo ↑k-VN1-k5 · #9
This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with tempo (TEMP) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it low at 0.25, lower than 75 % of clips in this corpus. That is a total rise of 0.24.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.08, then +0.04, then -0.08 — not a clean run: step 4 moves back the other way by 0.08 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.75 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.75 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 44 s · en · emolia
k 5d_a 0.241d_b 0.241step_a 0.203step_b 0.203min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00038_S08897track EN_B00038_S08897total 43.9slevel spread 4.2 dBmax seam 3.5 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult somewhat feminine voice · good recording, no background noise
(fear, sexual lust, malevolence malice · slow, very low-energy, relaxed, narration)Her prominent eyes swam with tears as she gasped for breath, staring at Ron.
full caption & clip details
An adult somewhat feminine voice; delivery is very low-energy, slow, relaxed, variable; timbre is slightly warm, slightly dark, gravelly, very thin; slurred, almost no disfluency, narrow pitch range, audible breath; affect is negative, submissive, vulnerable; reads as fear, sexual lust, malevolence malice; style: narration, monologue; good recording, no background noise; mildly explicit content; genuineness 0.7/6; vocal-burst blend 4.2/10; 6.3s, EN.
EN_B00038_S08897_W000135 · in -24.2 dBFS · gain +4.2 dB · emolia-00985
(infatuation, malevolence malice, contentment· slow, energised, slightly relaxed, narration)Attalin nonplussed, he looked around at the others, who were now laughing at the expression on Ron's face, and at the ludicrously prolonged laughter of Luna Lovegood, who was rocking backwards and forwards, clutching her sides.
full caption & clip details
A middle-aged feminine voice; delivery is energised, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, very wide pitch range, light breath; affect is negative, neutral stance, neutral openness; reads as infatuation, malevolence malice, contentment; style: narration, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.5/10; 14.3s, EN.
EN_B00038_S08897_W000136 · in -20.7 dBFS · gain +0.7 dB · emolia-00985
(teasing, amusement·measured, normally alert, slightly relaxed, storytelling)Are you taking the mickey? said Ron, frowning at her.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing, amusement; style: storytelling, whispered; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 5.7/10; 3.3s, EN.
EN_B00038_S08897_W000137 · in -20.0 dBFS · gain -0.0 dB · emolia-00985
(emotional numbness, amusement, malevolence malice· measured, normally alert, slightly relaxed, narration)Everyone else was watching Luna laughing, but Harry, glancing at the magazine on the floor, noticed something that made him dive for it.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness, amusement, malevolence malice; style: narration, whispered; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.5/10; 8.4s, EN.
EN_B00038_S08897_W000138 · in -22.2 dBFS · gain +2.2 dB · emolia-00985
(confusion, disappointment, sadness· measured, energised, slightly relaxed, narration)Upside down, it had been hard to tell what the picture on the front was, but Harry now realized it was a fairly bad cartoon of Cornelius Funch.
full caption & clip details
An adult feminine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, disappointment, sadness; style: narration, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.1/10; 10.9s, EN.
EN_B00038_S08897_W000139 · in -22.1 dBFS · gain +2.1 dB · emolia-00985
BKGN — background noise level ↑k-VN1-k5 · #10
This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with background noise level (BKGN) low — 0.10, lower than 90 % of clips in this corpus — and ends with it above average at 0.60, higher than 60 % of clips in this corpus. That is a total rise of 0.50.
It takes 5 clips to get there. Clip to clip the moves are -0.04, then +0.10, then +0.20, then +0.24 — not a clean run: step 1 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 61 s · en · podcast
k 5d_a 0.505d_b 0.505step_a 0.238step_b 0.238min_cos_consec 0.0438min_cos_anchor 0.1364dataset podcastlang enspeaker 155595track 155595total 60.8slevel spread 4.2 dBmax seam 2.5 dBcos from recomputed from spkemb_traj
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, quiet background
(amusement, pleasure ecstasy, teasing · normal-paced, normally alert, neutral tension, casual)I think you're talented writer, we should you should come join us. And he was like, I only played Daisy, and I was like that's his line. That's his line, it's like his bumper sticker. And (low mumble) uh when when our Daisy player at the time decided that they might be dropping it, I was like, yes, get out of here, drop him, drop 'em. So again, I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, pleasure ecstasy, teasing; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 10.0/10; 15.6s, EN.
155595_00010600 · in -26.8 dBFS · gain +6.8 dB · podcast-02634
(amusement, embarrassment, pleasure ecstasy · normal-paced, normally alert, neutral tension, casual)got it. And it was funny because I tried to play it cool. It's like I'm gonna think about it and then like five minutes later I already
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, embarrassment, pleasure ecstasy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 8.6/10; 7.5s, EN.
155595_00012160 · in -29.3 dBFS · gain +9.3 dB · podcast-02625
(embarrassment, intoxication altered states of consciousness, amusement · normal-paced, normally alert, fully relaxed, casual)later. Yeah, I had registered and had the app like on the works.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, intoxication altered states of consciousness, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 2.1/10; 4.2s, EN.
155595_00013040 · in -28.6 dBFS · gain +8.6 dB · podcast-02123
(hope enthusiasm optimism, elation, contentment·measured, very low-energy, relaxed, casual)Well we'll certainly ha happy to have you on the podcast, buddy. (low mumble) Um Yeah, so this week we are uh (ahem) we we rewatched (low mumble) um The Incredible Hulk and at the behest of Obi we also had to (low mumble) um endure uh (low mumble) two thousand three Amy's Hulk movie. (low mumble) Um yeah. So let's see. So I have the Incredible Hulk's IMDB pulled up here, (low mumble) uh and it was directed by
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, slightly guarded; reads as hope enthusiasm optimism, elation, contentment; style: casual, conversational; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.3/10; 29.2s, EN.
155595_00014072 · in -30.3 dBFS · gain +10.3 dB · podcast-02614
(sexual lust, amusement, intoxication altered states of consciousness·normal-paced, normally alert, slightly relaxed, casual)directed by Lewis Letier Off the top of my head.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, amusement, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.9/10; 3.6s, EN.
155595_00017128 · in -30.9 dBFS · gain +10.9 dB · podcast-02129
AROU — arousal / activation ↑k-VN1-k5 · #11
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with arousal / activation (AROU) below average — 0.40, lower than 60 % of clips in this corpus — and ends with it high at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.43.
It takes 5 clips to get there. Clip to clip the moves are -0.01, then +0.03, then +0.24, then +0.18 — not a clean run: step 1 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 43 s · en · emolia
k 5d_a 0.426d_b 0.426step_a 0.236step_b 0.236min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_h-Fz9XDOU24track EN_h-Fz9XDOU24total 43.1slevel spread 6.5 dBmax seam 5.1 dB
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, quiet background, normally alert
(normal-paced, slightly relaxed, fairly steady, casual)You (chuckle) know, so, uh, (low mumble) so if people are registered or if people can register for the serious play conference, they'll, uh, (low mumble) they'll get more for you.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 2.6/10; 7.2s, EN.
EN_h-Fz9XDOU24_W000368 · in -20.2 dBFS · gain +0.2 dB · emolia-01794
(thankfulness gratitude, affection, contentment·measured, relaxed, moderately variable, casual)Thank you. (low mumble) (low mumble) These were some really interesting ways of assessing people in a much more fun way than (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, affection, contentment; style: casual, conversational; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.7/10; 9.9s, EN.
EN_h-Fz9XDOU24_W000371 · in -22.7 dBFS · gain +2.7 dB · emolia-01794
(contentment ·normal-paced, neutral tension, fairly steady, casual)I think we're, we've all, we've all been used to. And I guess some people have tried a couple of these in the, in the past, but, (low mumble) uh, the way you presented them was, was really quite, (low mumble) um, quite interesting. Thank you.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contentment; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 3.3/10; 11.7s, EN.
EN_h-Fz9XDOU24_W000372 · in -24.8 dBFS · gain +4.8 dB · emolia-01794
(normal-paced, slightly relaxed, fairly steady, casual)Okay, okay, so (low mumble) uhm, I'll just sign off then. This is Mitch Weisberg. (ahem) Uh, actually before I do that, Marcia, can you stop sharing your screens for a second? Sure.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.1/10; 9.6s, EN.
EN_h-Fz9XDOU24_W000374 · in -19.7 dBFS · gain -0.3 dB · emolia-01794
(thankfulness gratitude, relief, affection· normal-paced, slightly relaxed, fairly steady, casual)And that way we can, (ahem) and I'll sign off. Thank you everybody for coming. Marcia.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, relief, affection; style: casual, conversational; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.5/10; 3.9s, EN.
EN_h-Fz9XDOU24_W000375 · in -18.3 dBFS · gain -1.7 dB · emolia-01794
R_THRT — resonance: throat ↑k-VN1-k5 · #12
This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: throat (R_THRT) low — 0.22, lower than 78 % of clips in this corpus — and ends with it above average at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.53.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.04, then +0.21, then +0.19 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 83 s · dutch · mls
k 5d_a 0.531d_b 0.531step_a 0.206step_b 0.206min_cos_consec —min_cos_anchor —dataset mlslang dutchspeaker 496track 496|cameraobscura_57_hildetotal 82.9slevel spread 4.9 dBmax seam 4.9 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly somewhat feminine voice · quiet background, slow, very low-energy, moderately variable, frequent disfluency, audible breath
(intoxication altered states of consciousness · slightly relaxed, clear, narrow pitch range, whispered)met een wensch in de oogen beurtelings den jager en de weitasch en het geweer heeft aangekeken gisteren avend ging er temet ien tusschen me bienen deur een dikke
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is warm, slightly dark, slightly rough, balanced body; clear, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as intoxication altered states of consciousness; style: whispered, ASMR; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.5/10; 17.8s, DUTCH.
496_1747_000019 · in -29.5 dBFS · gain +9.5 dB · mls-00097
(intoxication altered states of consciousness, jealousy and envy, infatuation·relaxed, slurred, fairly narrow pitch, storytelling)mag de jongen rais meeloopen vraagt arie aan krelis oom nou ja antwoordt deze t zel wel lukken
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as intoxication altered states of consciousness, jealousy and envy, infatuation; style: storytelling, whispered; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.8/10; 11.2s, DUTCH.
496_1747_000042 · in -26.8 dBFS · gain +6.8 dB · mls-00097
(emotional numbness, fatigue exhaustion, disgust·slightly relaxed, slurred, narrow pitch range, whispered)piet verslikt zich haast aan de laatste korst van zijn roggebrood met kaas een taaie sliet wordt uit den dorsch te voorschijn gehaald en pols en polsdrager zijn geïmproviseerd
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, rough, thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is negative, neutral stance, neutral openness; reads as emotional numbness, fatigue exhaustion, disgust; style: whispered, narration; poor recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 16.7s, DUTCH.
496_1747_000115 · in -31.7 dBFS · gain +11.7 dB · mls-00097
(fatigue exhaustion, awe, intoxication altered states of consciousness· slightly relaxed, slurred, narrow pitch range, whispered)zoodanig is de wording van den polsdrager maar nooit was een schepsel ter wereld dankbaarder voor zijn bestaan geen begunstigde slaaf kleeft zijn meester getrouwer aan dan de polsdrager den jager
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is neutral-toned, dark, rough, thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is negative, neutral stance, neutral openness; reads as fatigue exhaustion, awe, intoxication altered states of consciousness; style: whispered, storytelling; below-average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.2/10; 17.1s, DUTCH.
496_1747_000237 · in -29.7 dBFS · gain +9.7 dB · mls-00097
(fatigue exhaustion, emotional numbness, affection· slightly relaxed, slurred, very wide pitch range, cartoonish)hij verlaat zijn zijde niet hij springt den jager vr over alle slooten en klimt hem over honderd dijkjes na hij wandelt met hem het jachtveld met vermoeiende ziegezagen af
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, frequent disfluency, very wide pitch range, audible breath; affect is positive, neutral stance, neutral openness; reads as fatigue exhaustion, emotional numbness, affection; style: cartoonish, whispered; poor recording, quiet background; genuineness 0.8/6; vocal-burst blend 0.4/10; 19.5s, DUTCH.
496_1747_000111 · in -27.8 dBFS · gain +7.8 dB · mls-00097
S_STRY — style: storytelling ↑k-VN1-k5 · #13
This is a VoiceNet dimension, not an emotion: style: storytelling (S_STRY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: storytelling (S_STRY) below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the range at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.54.
It takes 5 clips to get there. Clip to clip the moves are +0.01, then +0.22, then +0.14, then +0.17 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.07 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 37 s · en · podcast
k 5d_a 0.544d_b 0.544step_a 0.223step_b 0.223min_cos_consec 0.0680min_cos_anchor 0.0704dataset podcastlang enspeaker 551961track 551961total 36.8slevel spread 5.2 dBmax seam 4.0 dBcos from recomputed from spkemb_traj
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice
(sadness, helplessness, contentment · measured, very low-energy, relaxed, casual)But it is one of the most satisfying things. Like we sat outside last night.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, dark, smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, slightly vulnerable; reads as sadness, helplessness, contentment; style: casual, whispered; below-average recording, some background noise; genuineness 4.6/6; vocal-burst blend 3.6/10; 5.1s, EN.
551961_00067680 · in -41.1 dBFS · gain +21.1 dB · podcast-05185
(fatigue exhaustion, embarrassment, amusement·normal-paced, very low-energy, relaxed, casual)We're like, furniture's clean. You know, like I could cook on the outdoor kitchen without like, let me s get some of that pollen off the counter, you know.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is slightly cool, very dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, embarrassment, amusement; style: casual, whispered; below-average recording, some background noise; mildly explicit content; genuineness 3.9/6; vocal-burst blend 3.4/10; 8.3s, EN.
551961_00068240 · in -39.9 dBFS · gain +19.9 dB · podcast-05193
(doubt, helplessness, longing·measured, normally alert, relaxed, casual)Growing up, I don't remember if there was ever that. I don't uh (wistful sigh) we we have too much patio furniture because we got a big patio.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is warm, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, helplessness, longing; style: casual, conversational; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 5.0/10; 6.4s, EN.
551961_00069312 · in -35.9 dBFS · gain +15.9 dB · podcast-05180
(longing, affection, infatuation·normal-paced, normally alert, neutral tension, casual)And we didn't have that much when I grew up and you guys had more 'cause you guys had a pretty big backyard there.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as longing, affection, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.8/10; 4.7s, EN.
551961_00070072 · in -36.1 dBFS · gain +16.1 dB · podcast-05208
(amusement, teasing, pleasure ecstasy· normal-paced, very low-energy, relaxed, casual)there being that much pollen covering everything when I grew up and having like that. We need to clean up the deck and the patio area. (contented sigh) Yeah.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is slightly warm, very dark, slightly rough, slightly thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is positive, slightly submissive, neutral openness; reads as amusement, teasing, pleasure ecstasy; style: casual, storytelling; poor recording, some background noise; genuineness 3.4/6; vocal-burst blend 4.1/10; 11.8s, EN.
551961_00070720 · in -38.1 dBFS · gain +18.1 dB · podcast-05175
METL — metallic quality ↓k-VN1-k5 · #14
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with metallic quality (METL) high — 0.78, higher than 78 % of clips in this corpus — and works its way down to around average at 0.51, right about the corpus median. That is a total fall of 0.28.
It takes 5 clips to get there. Clip to clip the moves are -0.10, then -0.00, then -0.20, then +0.02 — not a clean run: step 4 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 70 s · hr · eurospeech
k 5d_a -0.277d_b -0.277step_a 0.195step_b 0.195min_cos_consec —min_cos_anchor —dataset eurospeechlang hrspeaker croatia_20121204142431-778track croatia_20121204142431-778total 69.8slevel spread 0.6 dBmax seam 0.6 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly cool, fairly smooth, thin, average recording, quiet background, neutral tension, moderately variable, some disfluency
(affection, thankfulness gratitude, malevolence malice · brisk, normally alert, average clarity, dramatic)Hvala lijepa. Pa kolega Burić ovo što ste sada rekli kako Vlada obrazlaže zašto ne treba davati vjerodostojno to bi se moglo nazvati jednostavno jedan nespretni juristeraj.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection, thankfulness gratitude, malevolence malice; style: dramatic, authoritative; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.1/10; 10.9s, HR.
croatia_20121204142431-7787_5682464_5693392 · in -17.5 dBFS · gain -2.5 dB · eurospeech-01363
(confusion·fast, normally alert, average clarity, dramatic)gledajte ako oni uopće se upuštaju u to obrazlaganje onda je to dokaz da je potrebno vjerodostojno tumačenje gdje bi oni u vjerodostojnom tumačenju posegnuli povijesno
full caption & clip details
A child feminine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion; style: dramatic, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.9/10; 13.1s, HR.
croatia_20121204142431-7787_5693392_5706512 · in -17.2 dBFS · gain -2.8 dB · eurospeech-01363
(impatience and irritability, sourness, bitterness·brisk, energised, clear, cartoonish)povijesno dakle u ono kad je zakonodavac donosio zakon što je mislio, a to što je mislio to bi onda iščitavali iz toga što su oni predlagali amandmane, a predlagatelj nije prihvatio pa se to. Ali to je već tumačenje, onda bi trebali reći da tumačenje je potrebno,
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as impatience and irritability, sourness, bitterness; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 6.2/10; 16.7s, HR.
croatia_20121204142431-7787_5706512_5723200 · in -17.8 dBFS · gain -2.2 dB · eurospeech-01363
(shame, bitterness, contempt· brisk, energised, clear, dramatic)i istumačit ćemo ga tako, tako i tako, prema namjeri zakonodavca iz vremena donošenja zakona, ali oni sad dakle tumačenje neće dati, ali tvrde zašto je jasan, potpuno jasan zakon u smjeru kako je
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, bitterness, contempt; style: dramatic, cartoonish; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.1/10; 13.2s, HR.
croatia_20121204142431-7787_5723200_5736352 · in -17.7 dBFS · gain -2.3 dB · eurospeech-01363
(triumph, pride· brisk, normally alert, average clarity, cartoonish)njima jasan, jer vjerojatno (ahem) njihov kvocijent, odnosno premijerov kvocijent inteligencije je takav da je njemu uvijek sve jasno. A istovremeno ne podržava recimo, kažu oni mi ne podržavamo stajalište pravobraniteljice za
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, pride; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.3/10; 15.4s, HR.
croatia_20121204142431-7787_5736352_5751712 · in -17.3 dBFS · gain -2.7 dB · eurospeech-01363
RCQL — recording quality ↑k-VN1-k5 · #15
This is a VoiceNet dimension, not an emotion: recording quality (RCQL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with recording quality (RCQL) low — 0.11, lower than 89 % of clips in this corpus — and ends with it above average at 0.64, higher than 64 % of clips in this corpus. That is a total rise of 0.53.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.20, then +0.12, then +0.03 — a plateau around step 4, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 54 s · en · emolia
k 5d_a 0.525d_b 0.525step_a 0.196step_b 0.196min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_Xa7psxjs3aktrack EN_Xa7psxjs3aktotal 53.8slevel spread 3.8 dBmax seam 3.8 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, quiet background, neutral tension, moderately variable
(fatigue exhaustion, relief, pleasure ecstasy · normal-paced, very low-energy, frequent disfluency, casual)I've been drinking a lot, a lot, a lot of water. I don't drink juice. (ahem) Uhm, and if I did go out, (wistful sigh) uh, on last Friday, I did have two drinks. I do have a two drink maximum. I don't drink shots or anything. I'll just get a cute little drink, call it a day. I had two, I had a lychee martini and I had, (exhausted groan) uhm, a raspberry lemon drop. That's all the alcohol I had last week actually.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as fatigue exhaustion, relief, pleasure ecstasy; style: casual, monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 9.1/10; 24.0s, EN.
EN_Xa7psxjs3ak_W000111 · in -18.9 dBFS · gain -1.1 dB · emolia-01452
(disgust, infatuation, jealousy and envy· normal-paced, normally alert, some disfluency, casual)I don't like the clubs, I hate seeing the girls with no shoes on, standing on the couches. You can stand on the couch, put your shoes on.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as disgust, infatuation, jealousy and envy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 7.4/10; 6.0s, EN.
EN_Xa7psxjs3ak_W000115 · in -19.3 dBFS · gain -0.7 dB · emolia-01452
(embarrassment, fear, shame· normal-paced, normally alert, some disfluency, casual)It's just not my crowd, okay? It's just not me. I don't wanna be in there. (low mumble) Uhm, and like, I can understand if we're celebrating my birthday, cool, whatever, but.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment, fear, shame; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 6.5/10; 8.3s, EN.
EN_Xa7psxjs3ak_W000116 · in -18.5 dBFS · gain -1.5 dB · emolia-01452
(embarrassment, infatuation, jealousy and envy· normal-paced, normally alert, some disfluency, casual)Me and my friends and my man, we went out and we had a whole bunch of shots. We had like 24 shots total amongst the four of us. And that was the last time I took a shot.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment, infatuation, jealousy and envy; style: casual, conversational; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.2/10; 8.8s, EN.
EN_Xa7psxjs3ak_W000119 · in -22.3 dBFS · gain +2.3 dB · emolia-01452
(disgust, embarrassment, jealousy and envy ·brisk, normally alert, some disfluency, conversational)I'ma be so honest with you, I hated how I felt. And I didn't even have a hangover the next day, I just felt disgusted with myself. I was so disgusted. And I will never do it again.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as disgust, embarrassment, jealousy and envy; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.2/10; 6.1s, EN.
EN_Xa7psxjs3ak_W000120 · in -19.7 dBFS · gain -0.3 dB · emolia-01452
VALS — valence stability ↑k-VN1-k5 · #16
This is a VoiceNet dimension, not an emotion: valence stability (VALS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with valence stability (VALS) low — 0.22, lower than 78 % of clips in this corpus — and ends with it high at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.59.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.17, then +0.21, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 38 s · en · emolia
k 5d_a 0.592d_b 0.592step_a 0.217step_b 0.217min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_nA6jL-Lj9EEtrack EN_nA6jL-Lj9EEtotal 38.4slevel spread 1.3 dBmax seam 1.2 dB
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, average recording, quiet background, normally alert, fairly steady, somewhat unclear, moderate pitch range, light breath
(normal-paced, slightly relaxed, some disfluency, casual)45% of them that (ahem) we're not getting ready yet.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.6/10; 3.7s, EN.
EN_nA6jL-Lj9EE_W000051 · in -16.1 dBFS · gain -3.9 dB · emolia-01070
(measured, slightly relaxed, frequent disfluency, casual)So it's probably coming out from the fact that the managers are more informed, uh, (low mumble) than non-managers, but (ahem) it's something to, to highlight anyhow.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.0/10; 8.5s, EN.
EN_nA6jL-Lj9EE_W000052 · in -17.3 dBFS · gain -2.7 dB · emolia-01070
(normal-paced, slightly relaxed, frequent disfluency, casual)And when we ask HR though, uh, (low mumble) are each organization getting ready, (low mumble) uh, to deal with this technological development?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.2/10; 8.5s, EN.
EN_nA6jL-Lj9EE_W000053 · in -17.3 dBFS · gain -2.7 dB · emolia-01070
(measured, neutral tension, frequent disfluency, casual)A very large majority, 87% of them, I think, yes, definitely we're, we're, we're ready for that and we started to reflect upon it. (low mumble)
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.0/10; 9.5s, EN.
EN_nA6jL-Lj9EE_W000054 · in -17.4 dBFS · gain -2.6 dB · emolia-01070
(elation, triumph, hope enthusiasm optimism·normal-paced, slightly relaxed, some disfluency, casual)Almost 100% in Spain are staying, actually, for HR and L&D. Yes, we're ready and our organization is getting ready for it, so.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as elation, triumph, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.4/10; 7.7s, EN.
EN_nA6jL-Lj9EE_W000055 · in -17.4 dBFS · gain -2.6 dB · emolia-01070
S_MONO — style: monologue ↑k-VN1-k5 · #17
This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: monologue (S_MONO) low — 0.22, lower than 78 % of clips in this corpus — and ends with it high at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.58.
It takes 5 clips to get there. Clip to clip the moves are -0.00, then +0.20, then +0.25, then +0.13 — not a clean run: step 1 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 37 s · zh · emolia
k 5d_a 0.576d_b 0.576step_a 0.248step_b 0.248min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00067_S01345track ZH_B00067_S01345total 37.3slevel spread 3.5 dBmax seam 3.3 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · slightly dark, slightly rough, quiet background, measured
(jealousy and envy, doubt · energised, neutral tension, moderately variable, storytelling)He's got a mind that is always asking why why why why and he's very good at coming up.
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is warm, slightly dark, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as jealousy and envy, doubt; style: storytelling, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.7/10; 5.9s, ZH.
ZH_B00067_S01345_W000008 · in -27.2 dBFS · gain +7.2 dB · emolia-03947
(longing, fatigue exhaustion, distress·normally alert, neutral tension, fairly steady, storytelling)Thirty some years, i've been watching you. I would say what it takes to make you not sleep at night. It is illness in the family.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, very full; average clarity, almost no disfluency, wide pitch range, light breath; affect is neutral, dominant, fairly guarded; reads as longing, fatigue exhaustion, distress; style: storytelling, authoritative; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.4/10; 6.2s, ZH.
ZH_B00067_S01345_W000009 · in -26.8 dBFS · gain +6.8 dB · emolia-03947
(intoxication altered states of consciousness· normally alert, relaxed, fairly steady, casual)Where m likes the game, i like the game, and even in the periods that were tough to other people.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 7.3s, ZH.
ZH_B00067_S01345_W000010 · in -30.1 dBFS · gain +10.1 dB · emolia-03947
(normally alert, slightly relaxed, steady, narration)In those charts and immense amount of information is put in very usable form. And (low mumble) if i were running a business school, we would be teaching from violand jarks.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is warm, slightly dark, slightly rough, balanced body; average clarity, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, fairly guarded; no dominant emotion; style: narration, monologue; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.0/10; 10.6s, ZH.
ZH_B00067_S01345_W000011 · in -30.3 dBFS · gain +10.3 dB · emolia-03947
(normally alert, slightly relaxed, steady, formal)In any company, the stock could give to a price so high. It would be people for the corporation to repurchase its shares.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: formal, monologue; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.1/10; 6.7s, ZH.
ZH_B00067_S01345_W000012 · in -27.9 dBFS · gain +7.9 dB · emolia-03947
STRU — structuredness of delivery ↑k-VN1-k5 · #18
This is a VoiceNet dimension, not an emotion: structuredness of delivery (STRU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with structuredness of delivery (STRU) low — 0.10, lower than 90 % of clips in this corpus — and ends with it above average at 0.60, higher than 60 % of clips in this corpus. That is a total rise of 0.50.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.08, then -0.01, then +0.25 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.73 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.73 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 52 s · en · emolia
k 5d_a 0.503d_b 0.503step_a 0.246step_b 0.246min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_PN-84Vjj8MQtrack EN_PN-84Vjj8MQtotal 51.5slevel spread 1.7 dBmax seam 1.7 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an elderly strongly masculine voice · slightly cool, rough, below-average recording, quiet background, wide pitch range
(bitterness · slow, highly aroused, relaxed, storytelling)But someway somehow he gave me these ideas but these were my ideas. My ideas. You think
full caption & clip details
An elderly strongly masculine voice; delivery is highly aroused, slow, relaxed, moderately variable; timbre is slightly cool, slightly dark, rough, thin; very clear, frequent disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as bitterness; style: storytelling, authoritative; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.0/10; 9.2s, EN.
EN_PN-84Vjj8MQ_W000049 · in -18.2 dBFS · gain -1.8 dB · emolia-01314
(doubt, confusion· slow, energised, slightly relaxed, storytelling)Where do we, (ahem) other than just this?
full caption & clip details
A child masculine voice; delivery is energised, slow, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, rough, thin; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, dominant, fairly guarded; reads as doubt, confusion; style: storytelling, dramatic; below-average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.2/10; 3.1s, EN.
EN_PN-84Vjj8MQ_W000050 · in -16.6 dBFS · gain -3.5 dB · emolia-01314
(anger, disappointment, bitterness·measured, energised, neutral tension, storytelling)Other than the idea that I came up with and said back in the debates, severance tax dollars would rise. And you know what? They laughed at me. They laughed at me. And it did. And the other thing is just this. I said,
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, very full; very clear, frequent disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, fairly guarded; reads as anger, disappointment, bitterness; style: storytelling, monologue; below-average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.7/10; 19.7s, EN.
EN_PN-84Vjj8MQ_W000051 · in -18.2 dBFS · gain -1.8 dB · emolia-01314
(anger, bitterness, pain·slow, highly aroused, tense, authoritative)My roads plan is our ticket. Other than those two things, what would this budget have done?
full caption & clip details
A middle-aged strongly masculine voice; delivery is highly aroused, slow, tense, fairly steady; timbre is slightly cool, slightly dark, rough, thin; very clear, frequent disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as anger, bitterness, pain; style: authoritative, storytelling; below-average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 10.7s, EN.
EN_PN-84Vjj8MQ_W000052 · in -17.8 dBFS · gain -2.2 dB · emolia-01314
(anger, impatience and irritability·measured, energised, neutral tension, authoritative)If you pulled those two things out of our budget that we have today, the budget that our lawmakers passed
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, fairly steady; timbre is slightly cool, slightly dark, rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, impatience and irritability; style: authoritative, storytelling; below-average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.0/10; 8.2s, EN.
EN_PN-84Vjj8MQ_W000053 · in -17.2 dBFS · gain -2.8 dB · emolia-01314
EMPH — emphasis / stress strength ↓k-VN1-k5 · #19
This is a VoiceNet dimension, not an emotion: emphasis / stress strength (EMPH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with emphasis / stress strength (EMPH) high — 0.87, higher than 87 % of clips in this corpus — and works its way down to below average at 0.31, lower than 69 % of clips in this corpus. That is a total fall of 0.56.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then -0.13, then -0.24, then -0.19 — not a clean run: step 1 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 43 s · en · emolia
k 5d_a -0.562d_b -0.562step_a 0.239step_b 0.239min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_HsGSsrkBn9utrack EN_HsGSsrkBn9utotal 42.6slevel spread 3.1 dBmax seam 2.4 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, average clarity
(normal-paced, energised, neutral tension, conversational)And it's not safe to drive. Well, let's look at the data. Yeah, car crashes are pretty darn common. However,
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; no dominant emotion; style: conversational, dramatic; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.5/10; 8.9s, EN.
EN_HsGSsrkBn9u_W000068 · in -16.2 dBFS · gain -3.8 dB · emolia-02509
(fear·brisk, normally alert, neutral tension, casual)Ice storm. That's probably not the best time to be out there driving. So yeaah, there are gonna be a lot of crashes then. But if it is sunny weather and not rush hour and, you know, all things considered, is driving all that unsafe.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as fear; style: casual, conversational; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.9/10; 16.4s, EN.
EN_HsGSsrkBn9u_W000071 · in -15.6 dBFS · gain -4.5 dB · emolia-02509
(normal-paced, normally alert, slightly relaxed, casual)So we want to help them look at these polarizations. Every time
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 4.6/10; 3.1s, EN.
EN_HsGSsrkBn9u_W000072 · in -13.1 dBFS · gain -6.9 dB · emolia-02509
(contentment·slow, very low-energy, relaxed, casual)They say an extreme word and you'll hear in your clients there are certain ones that they like to use, (ahem) uhm, always, never, have to, (low mumble) uhm,
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 10.1s, EN.
EN_HsGSsrkBn9u_W000073 · in -15.5 dBFS · gain -4.5 dB · emolia-02509
(normal-paced, normally alert, slightly relaxed, storytelling)Challenge them to take that word out of their vocabulary.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: storytelling, casual; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.1/10; 3.6s, EN.
EN_HsGSsrkBn9u_W000074 · in -14.5 dBFS · gain -5.5 dB · emolia-02509
R_MIXD — resonance: mixed ↓k-VN1-k5 · #20
This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: mixed (R_MIXD) above average — 0.66, higher than 66 % of clips in this corpus — and works its way down to low at 0.14, lower than 86 % of clips in this corpus. That is a total fall of 0.52.
It takes 5 clips to get there. Clip to clip the moves are -0.25, then +0.02, then -0.06, then -0.24 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · fr · emolia
k 5d_a -0.522d_b -0.522step_a 0.247step_b 0.247min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_r0IEKEBOs4otrack FR_r0IEKEBOs4ototal 50.7slevel spread 3.4 dBmax seam 3.4 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth, average recording, quiet background, some disfluency, light breath
(astonishment surprise, pleasure ecstasy · brisk, normally alert, slightly relaxed, dramatic)Je sais pas si vous avez vu, si vous êtes allé faire un tour sur les précommandes de la Samsung Galaxy Watch 5, mais il y a beaucoup de personnalisations possibles. En gros, on vous demande de choisir le boîtier et après vous êtes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise, pleasure ecstasy; style: dramatic, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.9/10; 10.3s, FR.
FR_r0IEKEBOs4o_W000004 · in -20.7 dBFS · gain +0.7 dB · emolia-02824
(pride·fast, normally alert, neutral tension, casual)En cuir, et cetera, et cetera. Il y a beaucoup de couleurs. Moi, j'ai opté pour cette couleur assez sympa. Bien entendu, je vous rappelle qu'il y aura deux tailles aussi sur la Galaxy Watch 5. Un boîtier de 40 mm et un boîtier de 44. Moi, j'ai pris celui de 44.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pride; style: casual, didactic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 8.4/10; 14.4s, FR.
FR_r0IEKEBOs4o_W000005 · in -20.1 dBFS · gain +0.1 dB · emolia-02824
(contentment· fast, normally alert, slightly relaxed, casual)assez sport, assez souple. Elle est en train de s'allumer. On va parler maintenant de son design, ses caractéristiques techniques. Et bien entendu, bien entendu, on la prendra en main ensemble. C'est parti.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment; style: casual, didactic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.5/10; 10.2s, FR.
FR_r0IEKEBOs4o_W000006 · in -23.0 dBFS · gain +3.0 dB · emolia-02824
(fast, energised, slightly relaxed, dramatic)Parlons maintenant du design de cette galaxie Watch 5. Je le disais lorsque je l'ai prise en main, elle est pratiquement identique, hein.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, storytelling; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 3.8/10; 6.6s, FR.
FR_r0IEKEBOs4o_W000007 · in -19.6 dBFS · gain -0.4 dB · emolia-02824
(disgust· fast, energised, neutral tension, dramatic)à la Galaxy Watch 4. Le look nous fait penser à une simple petite mise à jour par rapport au modèle précédent. Cette Galaxy Watch 5 est disponible en deux formats.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust; style: dramatic, casual; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 6.5/10; 8.7s, FR.
FR_r0IEKEBOs4o_W000008 · in -20.2 dBFS · gain +0.2 dB · emolia-02824