Manifest tier. emotion, rule B1, T=0.25, step cap 0.2. Population 3,895,947 chains (44,903 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 3,036,096.
Rule.B1 — one-sided: emotion B rises by >=T; the other axis is unconstrained Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.emotion__B1__T0.25__C0.20__INTERNAL — population 3,895,947 chains (44,903 h). SHAREABLE variant: 3,036,096. Filter.rule=='B1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.25 and step_b<=0.2 Sampled from 497,796 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.32.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.85 to 0.30 (-0.55), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.00, then +0.12 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · zh · emolia
hear it un-normalised (raw levels, max seam 4.0 dB)
k 4d_a -0.550d_b 0.317step_a 0.663step_b 0.194min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00062_S09061track ZH_B00062_S09061total 32.7slevel spread 4.0 dBmax seam 4.0 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, no background noise, normally alert, slightly relaxed, fairly steady, clear
(normal-paced, no disfluency, monologue, formal)萧南向行业协会举报了周到周到被市场的监察给问责,他答应先进整改。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.0/10; 7.1s, ZH.
ZH_B00062_S09061_W000016 · in -18.3 dBFS · gain -1.7 dB · emolia-03893
(concentration· normal-paced, no disfluency, monologue, whispered)齐天佐监督大家工作陈朗,给他端茶倒水。齐天佐很不习惯陈朗,劝他不要在后厨溜达,这会让大家不自在七天走,只好先离开了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, whispered; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.2/10; 12.6s, ZH.
ZH_B00062_S09061_W000017 · in -22.4 dBFS · gain +2.4 dB · emolia-03893
(thankfulness gratitude·fast, some disfluency, monologue, dramatic)东依一约了几天见面,感谢一直以来对他的照顾和关心。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as thankfulness gratitude; style: monologue, dramatic; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 2.6/10; 5.0s, ZH.
ZH_B00062_S09061_W000018 · in -21.1 dBFS · gain +1.1 dB · emolia-03893
(normal-paced, no disfluency, monologue, whispered)董一承认,在他最脆弱的时候,把七天当成了一号,他当面把七天的微信给删除了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, whispered; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 7.6s, ZH.
ZH_B00062_S09061_W000019 · in -21.4 dBFS · gain +1.4 dB · emolia-03893
Pride ↑ (unconstrained axis: Thankfulness Gratitude)emotion__B1__T0.25__C0.20__INTERNAL · #2
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pride clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.28.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.99 to 0.46 (-0.54), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.07, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 56 s · french · mls
hear it un-normalised (raw levels, max seam 2.7 dB)
k 4d_a -0.538d_b 0.277step_a 0.417step_b 0.199min_cos_consec —min_cos_anchor —dataset mlslang frenchspeaker 6249track 6249|paradisperdu_11_miltototal 55.5slevel spread 2.7 dBmax seam 2.7 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an elderly masculine voice · neutral-toned, slightly dark, quiet background, relaxed, steady, frequent disfluency, fairly narrow pitch, normal breath
(thankfulness gratitude, awe, contentment · measured, subdued, slurred, whispered)un être qui magnanime pût cor respondre d'ici avec le ciel mais reconnaître dans sa grati tude d'où son bien
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, awe, contentment; style: whispered, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.0/10; 11.3s, FRENCH.
6249_11031_001101 · in -27.6 dBFS · gain +7.6 dB · mls-00046
(infatuation, malevolence malice, emotional numbness· measured, subdued, slurred, whispered)et le coeur la voix les yeux dévotement dirigés là adorer révérer le dieu suprême qui le fit chef de tous ses ouvrages c'est pourquoi le père tout-puis sant éternel
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, malevolence malice, emotional numbness; style: whispered, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.2/10; 18.9s, FRENCH.
6249_11031_001930 · in -28.7 dBFS · gain +8.7 dB · mls-00046
(awe, contemplation, concentration· measured, very low-energy, somewhat unclear, monologue)distinctement à son fils parla de la sorte faisons à présent l'homme à notre image et à notre ressemblance et qu'il commande aux poissons de la mer
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, steady; timbre is neutral-toned, slightly dark, rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as awe, contemplation, concentration; style: monologue, storytelling; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.5/10; 14.8s, FRENCH.
6249_11031_001521 · in -26.0 dBFS · gain +6.0 dB · mls-00046
(pride, intoxication altered states of consciousness·slow, very low-energy, slurred, monologue)aux ce oiseaux du ciel aux bêtes des champs à toute la terre et à tous les reptiles qui se remuent sur la terre
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, neutral openness; reads as pride, intoxication altered states of consciousness; style: monologue, didactic; below-average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.0/10; 10.1s, FRENCH.
6249_11031_001735 · in -27.3 dBFS · gain +7.3 dB · mls-00046
This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Contempt clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.37.
Nothing was asked of the other axis, and in fact Malevolence Malice barely moves at all, sitting near 0.90 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.09, then +0.05, then +0.06 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 43 s · zh · emolia
hear it un-normalised (raw levels, max seam 2.1 dB)
k 5d_a 0.032d_b 0.369step_a 0.091step_b 0.177min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00010_S09367track ZH_B00010_S09367total 43.4slevel spread 3.3 dBmax seam 2.1 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, average recording, normally alert, slightly relaxed, light breath
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, malevolence malice, contempt; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 7.6/10; 14.6s, ZH.
ZH_B00010_S09367_W000374 · in -19.3 dBFS · gain -0.7 dB · emolia-03379
(contempt, impatience and irritability, bitterness·fast, moderately variable, some disfluency, dramatic)对不对?但是如果说你把那个音乐那个调改一下,对吧?万松放那个海底。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, impatience and irritability, bitterness; style: dramatic, storytelling; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.8/10; 5.9s, ZH.
ZH_B00010_S09367_W000375 · in -17.2 dBFS · gain -2.8 dB · emolia-03379
Emotional Numbness ↑ (unconstrained axis: Intoxication Altered States of Consciousness)emotion__B1__T0.25__C0.20__INTERNAL · #4
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Emotional Numbness clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.28.
Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.97 to 0.43 (-0.54), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.08, then +0.07, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.70 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.70 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 35 s · de · emolia
hear it un-normalised (raw levels, max seam 4.9 dB)
k 4d_a -0.537d_b 0.280step_a 0.662step_b 0.133min_cos_consec —min_cos_anchor —dataset emolialang despeaker DE_ISpvVLsG_e4track DE_ISpvVLsG_e4total 35.2slevel spread 5.2 dBmax seam 4.9 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording
(intoxication altered states of consciousness, concentration, fatigue exhaustion · slow, very low-energy, relaxed, monologue)(low mumble) Die Ergebnisse hier nochmal, (ahem) ähm, verschiedenen, (low mumble) ähm, Parameter, (ahem) und, (low mumble) ähm, ich beschränke mich mal hier jetzt einfach auf den Ertrag, äh, (low mumble) pro Hektar, Tonnen pro Hektar, also hier haben wir 4,2
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, audible breath; affect is neutral, submissive, slightly guarded; reads as intoxication altered states of consciousness, concentration, fatigue exhaustion; style: monologue, whispered; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 19.9s, DE.
DE_ISpvVLsG_e4_W000103 · in -24.6 dBFS · gain +4.6 dB · emolia-00043
(sourness, bitterness, disappointment·normal-paced, normally alert, slightly relaxed, casual)ja, das zeigt sich halt auch in den anderen Messergebnissen. Sehr, sehr viel gemessen worden. Das ist alles sehr teuer und aufwendig.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, bitterness, disappointment; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.7/10; 5.8s, DE.
DE_ISpvVLsG_e4_W000104 · in -19.7 dBFS · gain -0.3 dB · emolia-00043
(confusion, impatience and irritability, sourness · normal-paced, normally alert, slightly relaxed, conversational)Aber das muss man natürlich bei der Praxisanwendung dann nicht mehr jedes Mal machen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, impatience and irritability, sourness; style: conversational, casual; average recording, no background noise; genuineness 4.1/6; vocal-burst blend 0.0/10; 3.5s, DE.
DE_ISpvVLsG_e4_W000105 · in -20.3 dBFS · gain +0.3 dB · emolia-00043
(emotional numbness·measured, normally alert, slightly relaxed, formal)Gibt es ein Qualitätskriterium? (low mumble) Die meisten Kleinbauern können sich nicht erlauben.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.3/10; 5.6s, DE.
DE_ISpvVLsG_e4_W000106 · in -19.4 dBFS · gain -0.6 dB · emolia-00043
Sexual Lust ↑ (unconstrained axis: Embarrassment)emotion__B1__T0.25__C0.20__INTERNAL · #5
This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Sexual Lust clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.
Nothing was asked of the other axis, and in fact Embarrassment drifts down from 0.99 to 0.93 (-0.06), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.41 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.55 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.41, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 38 s · en · podcast
hear it un-normalised (raw levels, max seam 3.0 dB)
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, average recording, moderately variable, some disfluency, wide pitch range
(embarrassment, longing, contemplation · normal-paced, normally alert, neutral tension, casual)I'm sitting here like, how do I celebrate? How do I need to figure out how to make things exciting because before it was so easy to make things exciting?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, longing, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 7.3/10; 8.3s, EN.
598864_00033528 · in -29.3 dBFS · gain +9.3 dB · podcast-00831
(astonishment surprise, elation, jealousy and envy·brisk, energised, neutral tension, casual)I mean, you turn 18, you're an adult, you vote, you turn 20, you're no longer a teen. Turn twenty (wistful sigh) one.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as astonishment surprise, elation, jealousy and envy; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 4.9/6; vocal-burst blend 8.0/10; 22.9s, EN.
598864_00034376 · in -31.4 dBFS · gain +11.3 dB · podcast-06123
(sexual lust, pleasure ecstasy, amusement·normal-paced, normally alert, relaxed, casual)(breathy giggle) But we also love ice cream and (breathy giggle) I
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as sexual lust, pleasure ecstasy, amusement; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.3/10; 6.4s, EN.
598864_00268368 · in -28.4 dBFS · gain +8.4 dB · podcast-00805
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.38.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.98 to 0.39 (-0.59), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.17, then -0.16, then +0.17 — not a clean run: step 3 moves back the other way by 0.16 before the chain recovers.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, normally alert, some disfluency, light breath
(infatuation, affection, relief · normal-paced, neutral tension, fairly steady, casual)(ahem) Um, so anybody that is on the verge, doesn't know really anything too much about God, is just pray to him. You know, and I was gonna ask if we could end with prayer.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as infatuation, affection, relief; style: casual, conversational; average recording, some background noise; genuineness 4.5/6; vocal-burst blend 6.3/10; 9.4s, EN.
318601_00759688 · in -12.8 dBFS · gain -7.2 dB · podcast-00721
(infatuation, sexual lust, awe·brisk, neutral tension, moderately variable, casual)don't think you have to be perfect before you come to God. Come to him just as you are. You know, Jesus died on the cross, like I said earlier, he He
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as infatuation, sexual lust, awe; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 4.6/6; vocal-burst blend 9.1/10; 6.6s, EN.
318601_00760912 · in -12.1 dBFS · gain -8.0 dB · podcast-00715
(contentment, hope enthusiasm optimism, affection· brisk, slightly relaxed, fairly steady, casual)already knows everything you were gonna do, and and even if you are gonna mess up tomorrow, I mean we don't know. I I mess up a lot. You know, but the difference is that the sin starts falling off and you can't be okay with it anymore when the Holy Spirit comes and Jesus' spirit starts living in you, and the things that used to make you happy don't make you happy anymore. And that's the God's coming after you, and he may be calling you today. And I just wanted to pray with you guys if you guys are listening. And just know heaven's not for perfect people. My pastor always says heaven's not for perfect people, it's for forgiven people.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contentment, hope enthusiasm optimism, affection; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 10.0/10; 28.4s, EN.
318601_00761568 · in -11.9 dBFS · gain -8.1 dB · podcast-00535
(relief, pleasure ecstasy, amusement·normal-paced, slightly relaxed, moderately variable, conversational)God for forgiveness, and (low mumble) um that's okay with you guys. Yeah, that's it. Lord (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as relief, pleasure ecstasy, amusement; style: conversational, casual; average recording, some background noise; genuineness 5.0/6; vocal-burst blend 0.0/10; 4.9s, EN.
318601_00764696 · in -13.2 dBFS · gain -6.8 dB · podcast-00743
(hope enthusiasm optimism, awe, affection·measured, slightly relaxed, fairly steady, casual)That we're here and we're putting the spotlight on you, Jesus, and and how you can change the human heart, God.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, awe, affection; style: casual, monologue; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.8/10; 6.4s, EN.
318601_00767248 · in -11.2 dBFS · gain -8.8 dB · podcast-00330
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Confusion clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Emotional Numbness climbs from 0.71 to 0.93 (+0.22), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are -0.02, then +0.15, then +0.16 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 28 s · zh · emolia
k 4d_a 0.220d_b 0.292step_a 0.106step_b 0.161min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00055_S03668track ZH_B00055_S03668total 28.0slevel spread 1.5 dBmax seam 1.1 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, measured, normally alert, slightly relaxed
(fairly steady, some disfluency, somewhat unclear, authoritative)让我们先把三维的空间压缩成两维的空间。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 4.1/10; 4.3s, ZH.
ZH_B00055_S03668_W000007 · in -19.0 dBFS · gain -1.0 dB · emolia-03830
(pain·steady, no disfluency, clear, monologue)你可能满心期待能看到一张惊世骇俗的做梦,也没想到过的神图。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, narration; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.1/10; 7.0s, ZH.
ZH_B00055_S03668_W000008 · in -19.1 dBFS · gain -0.9 dB · emolia-03830
(fairly steady, no disfluency, clear, monologue)可是眼前就是一张随便打开一本中学数学课本就能看到的图,各位耐心点。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 3.9/10; 8.6s, ZH.
ZH_B00055_S03668_W000009 · in -19.4 dBFS · gain -0.6 dB · emolia-03830
(confusion, emotional numbness· fairly steady, some disfluency, clear, monologue)我觉得形容正是你接下去将要看到的内容,请先耐着性子听我解释一下。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, emotional numbness; style: monologue, didactic; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 4.0/10; 7.6s, ZH.
ZH_B00055_S03668_W000010 · in -20.5 dBFS · gain +0.5 dB · emolia-03830
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Emotional Numbness clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.98 to 0.33 (-0.65), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 25 s · en · emolia
k 3d_a -0.648d_b 0.274step_a 0.561step_b 0.154min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00078_S06890track EN_B00078_S06890total 25.1slevel spread 10.0 dBmax seam 10.0 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, frequent disfluency
(infatuation, pride, concentration · normal-paced, steady, somewhat unclear, monologue)Reading the table from first row to second row. First row is females, second row is males. So the first row mean is 3.7 minus the second row mean is 3.2.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, pride, concentration; style: monologue; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.4/10; 12.9s, EN.
EN_B00078_S06890_W000007 · in -22.6 dBFS · gain +2.6 dB · emolia-01731
(measured, steady, average clarity, monologue)Since that number is well below 1%, our conclusion
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, whispered; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.0/10; 4.7s, EN.
EN_B00078_S06890_W000008 · in -12.6 dBFS · gain -7.4 dB · emolia-01731
(emotional numbness, concentration· measured, fairly steady, somewhat unclear, monologue)So s1 squared plus n1 over s2 squared plus n2, where s1
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: monologue, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.4/10; 7.1s, EN.
EN_B00078_S06890_W000009 · in -20.1 dBFS · gain +0.1 dB · emolia-01731
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Emotional Numbness clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Contempt drifts down from 0.90 to 0.14 (-0.77), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.02, then +0.04, then +0.11 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 35 s · fr · emolia
k 5d_a -0.769d_b 0.254step_a 0.688step_b 0.113min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_xRFyIWrqsHUtrack FR_xRFyIWrqsHUtotal 35.1slevel spread 2.3 dBmax seam 2.0 dB
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, quiet background, normally alert, slightly relaxed, fairly steady, moderate pitch range
(contempt · measured, some disfluency, average clarity, monologue)ou les pénibilités sont nombreuses que l'on parle des pénibilités (ahem) physiques. Dominique l'a rappelé en introduction sur le bâtiment. (ahem)
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt; style: monologue, didactic; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.9/10; 7.0s, FR.
FR_xRFyIWrqsHU_W000060 · in -17.9 dBFS · gain -2.1 dB · emolia-02858
(awe, confusion, longing· measured, some disfluency, average clarity, casual)taux d'accident (low mumble) parmi les plus élevés, (ahem) euh, y a une comprise d'accidents mortels.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as awe, confusion, longing; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.6/10; 4.9s, FR.
FR_xRFyIWrqsHU_W000061 · in -15.9 dBFS · gain -4.1 dB · emolia-02858
(confusion · measured, frequent disfluency, somewhat unclear, monologue)(low mumble) euh, donc, les pénibilités physiques, les, les coiffeuses, hein, sont pas, (ahem) euh, sont, sont pas mieux loties, (ahem) euh, les TMS, il y en a énormément dans ces secteurs.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.5/10; 8.7s, FR.
FR_xRFyIWrqsHU_W000062 · in -15.6 dBFS · gain -4.4 dB · emolia-02858
(normal-paced, some disfluency, average clarity, monologue)(low mumble) mais aussi des pénibilités mentales, parce que (ahem) dans les métiers de service, en particulier dans la coiffure et dans la restauration,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.0/10; 6.5s, FR.
FR_xRFyIWrqsHU_W000063 · in -16.2 dBFS · gain -3.8 dB · emolia-02858
(emotional numbness· normal-paced, some disfluency, somewhat unclear, monologue)les exigences (ahem) émotionnelles, (low mumble) euh, en lien avec le travail auprès d'une clientèle.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: monologue, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.5/10; 7.4s, FR.
FR_xRFyIWrqsHU_W000064 · in -17.6 dBFS · gain -2.4 dB · emolia-02858
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.93 to 0.17 (-0.77), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.04, then +0.10, then +0.12, then +0.00 — a plateau around step 4, where it barely moves.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.18 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.35 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.18, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, moderate pitch range, light breath
(infatuation, astonishment surprise · normal-paced, normally alert, slightly relaxed, monologue)Like with social media now, I'll see DJs post on Instagram, like the DJ booth looks like the most fun place in the world.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, astonishment surprise; style: monologue, casual; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.2/10; 5.9s, EN.
133044_00238112 · in -23.0 dBFS · gain +3.0 dB · podcast-05069
(astonishment surprise, fatigue exhaustion, amusement· normal-paced, normally alert, slightly relaxed, casual)you can see the crowds just kind of standing there and I'm always curious. I'm like, I want to ask the GM or the manager be like, what were sales that night?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise, fatigue exhaustion, amusement; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 5.4/10; 6.2s, EN.
133044_00238760 · in -20.4 dBFS · gain +0.4 dB · podcast-05035
(elation, contentment·measured, subdued, relaxed, casual)That's right. That's right. So (low mumble) um, you know, we have (low mumble) uh one of the most (low mumble) um entertaining things that we see on a weekly basis is when we when we When people make their comments on some of the things that we say on the show or some of the clips that we release later.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as elation, contentment; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 5.7/10; 14.9s, EN.
133044_00239712 · in -17.1 dBFS · gain -2.9 dB · podcast-05035
(interest, hope enthusiasm optimism, sourness·normal-paced, very low-energy, neutral tension, casual)Oh my goodness. Like the I'm I'm sure there's gonna be some folks that are gonna be like DJs ain't nothing, whatever. Like but like seriously, prove me wrong. I'm right here. Like that like DJs are responsible for literally hundreds of millions, if not billions of dollars in food and beverage sales across the world. Right. There's something it's a universal language that you guys are doing in music. (low mumble) Um you gotta know it. You gotta have (ahem) um you gotta have access to a light.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as interest, hope enthusiasm optimism, sourness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 10.0/10; 25.2s, EN.
133044_00241371 · in -16.5 dBFS · gain -3.5 dB · podcast-06133
(hope enthusiasm optimism, elation, sexual lust· normal-paced, normally alert, neutral tension, casual)You gotta have access to one, get booked, you gotta have access to the (low mumble) um to the training, you gotta have access to the to the sounds, the music and and like you're working, like especially when you're doing a big event like New Year's, you're timing things down to a clock.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, sexual lust; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 10.0/10; 12.0s, EN.
133044_00243888 · in -17.4 dBFS · gain -2.6 dB · podcast-02176
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Concentration clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.99 to 0.81 (-0.18), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.12, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: an adult feminine voice · neutral-bright, balanced body
(hope enthusiasm optimism, contentment, thankfulness gratitude · normal-paced, normally alert, slightly relaxed, conversational)it's really important that we we, for me, it has been very important to keep the humility,
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, contentment, thankfulness gratitude; style: conversational, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.2/10; 6.2s, EN.
888934_00257008 · in -16.7 dBFS · gain -3.3 dB · podcast-03848
(contemplation, infatuation, contentment ·slow, normally alert, relaxed, conversational)to keep the curiosity, (ahem) uh, and also to listen meaningfully, and then really spur action.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, infatuation, contentment; style: conversational, casual; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.2/10; 8.5s, EN.
888934_00257672 · in -17.0 dBFS · gain -3.0 dB · podcast-03864
(contemplation, doubt, concentration·normal-paced, very low-energy, slightly relaxed, monologue)where you're not operating from a place of maybe inherent biases that you have about certain groups or certain people, certain attributes of people? So, how do you deepen that understanding? And then on top of that, you mentioned something there around insecurity. Uh, (low mumble) because sometimes if you frame your leadership or your position as a position that is supposed to know all the answers, it's supposed to give all the answers. Then when it comes time to say, I don't know. (exhausted groan) Yeah, uh, that's a hard thing to say.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, doubt, concentration; style: monologue, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 6.3/10; 28.1s, EN.
888934_00260640 · in -16.4 dBFS · gain -3.6 dB · podcast-00715
(concentration, contemplation, doubt · normal-paced, normally alert, neutral tension, conversational)So, how do you, what sort of pillars do you need to build your identity on as a leader to have that sense of deep self-worth and security, where if you don't know the answer, you say, I don't know the answer. (low mumble) Um, if there is somebody who has a better approach, you can create space for them to thrive as well. And that's really important. So, how how did I do it? Or how would I uh (ahem) see it being done? (ahem) Um I would say (low mumble) um
full caption & clip details
An elderly masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, contemplation, doubt; style: conversational, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 10.0/10; 27.8s, EN.
888934_00263488 · in -16.9 dBFS · gain -3.1 dB · podcast-02749
Fear ↑ (unconstrained axis: Emotional Numbness)emotion__B1__T0.25__C0.20__INTERNAL · #12
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Fear around average — 0.49, right about the corpus median — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.32.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.97 to 0.73 (-0.24), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.00, then +0.15 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 41 s · en · emolia
k 4d_a -0.237d_b 0.321step_a 0.761step_b 0.171min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_oLrUFHAOtZotrack EN_oLrUFHAOtZototal 41.1slevel spread 1.5 dBmax seam 1.5 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed, fairly steady
(emotional numbness · no disfluency, light breath, formal, casual)The state constitution was amended to reject same-sex marriage
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, casual; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.2/10; 3.9s, EN.
EN_oLrUFHAOtZo_W000218 · in -15.4 dBFS · gain -4.6 dB · emolia-00657
(concentration, triumph, pain·almost no disfluency, no audible breath, newsreading, formal)In January 2007, Ron Ramsey became the first Republican elected as Speaker of the State Senate since Reconstruction, as a result of the realignment of the Democratic and Republican parties in the South since the late 20th century, with Republicans now elected by conservative voters, who previously had supported Democrats.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph, pain; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 19.3s, EN.
EN_oLrUFHAOtZo_W000219 · in -16.6 dBFS · gain -3.4 dB · emolia-00657
(emotional numbness·no disfluency, light breath, formal, newsreading)In 2010, during the 2010 midterm elections, Bill Haslam was elected to succeed Breederson, who was term-limited, to become the 49th Governor of Tennessee in January 2011
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.2s, EN.
EN_oLrUFHAOtZo_W000220 · in -15.1 dBFS · gain -5.0 dB · emolia-00657
(no disfluency, light breath, formal, casual)In April and May 2010, flooding in Middle Tennessee devastated Nashville and other parts of Middle Tennessee
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 6.4s, EN.
EN_oLrUFHAOtZo_W000221 · in -15.4 dBFS · gain -4.6 dB · emolia-00657
This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Embarrassment barely there — 0.18, lower than 82 % of clips in this corpus — and ends with it around average at 0.58, higher than 58 % of clips in this corpus. That is a total rise of 0.40.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.97 to 0.23 (-0.74), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.15, then +0.11 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 34 s · en · emolia
k 4d_a -0.742d_b 0.404step_a 0.698step_b 0.150min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_c_X7vNsSUbMtrack EN_c_X7vNsSUbMtotal 33.8slevel spread 1.5 dBmax seam 1.2 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness, shame · steady, formal, authoritative)In the words of the bishops at Chalcedon, an excommunicated the pope himself.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, shame; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 4.6s, EN.
EN_c_X7vNsSUbM_W000100 · in -13.5 dBFS · gain -6.5 dB · emolia-00457
(pain·fairly steady, formal, newsreading)Meanwhile, Leo I had received the appeals of Theodoret and Flavianne, of whose death he was unaware, and had written to them and to the Emperor and Empress, informing them that all of the Acts of the Council were nullified
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.0s, EN.
EN_c_X7vNsSUbM_W000101 · in -14.7 dBFS · gain -5.3 dB · emolia-00457
(sadness, bitterness, disappointment· fairly steady, newsreading, formal)He eventually excommunicated all who had taken part in it and absolved all whom it had condemned – including Theodorette, with the exception of Domnus of Antioch, who seems to have had no wish to resume his see and retired into the monastic life that he had left many years earlier with regret.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, bitterness, disappointment; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.3s, EN.
EN_c_X7vNsSUbM_W000102 · in -14.6 dBFS · gain -5.4 dB · emolia-00457
(fairly steady, formal, casual)Catholic Encyclopedia. New York, Robert Appleton.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.1/10; 3.4s, EN.
EN_c_X7vNsSUbM_W000103 · in -15.0 dBFS · gain -5.0 dB · emolia-00457
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.55.
Nothing was asked of the other axis, and in fact Disgust drifts down from 0.98 to 0.65 (-0.33), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.97 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.97 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 43 s · en · emolia
k 4d_a -0.329d_b 0.549step_a 0.316step_b 0.196min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_v7XM6MwWsOutrack EN_v7XM6MwWsOutotal 42.6slevel spread 1.2 dBmax seam 1.2 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(disgust, concentration, emotional numbness · newsreading, formal)Horizontal segregation is likely to be increased by post-industrial restructuring of the economy
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, concentration, emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 11.8s, EN.
EN_v7XM6MwWsOu_W000007 · in -15.0 dBFS · gain -5.0 dB · emolia-01494
(newsreading, formal)The millions of housewives who entered the economy during post-industrial restructuring primarily entered into service sector jobs where they could work part time and having flexible hours
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 9.5s, EN.
EN_v7XM6MwWsOu_W000008 · in -14.8 dBFS · gain -5.2 dB · emolia-01494
(disgust, malevolence malice· newsreading, formal)While these options are often appealing to mothers, who are often
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, malevolence malice; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.7s, EN.
EN_v7XM6MwWsOu_W000009 · in -15.2 dBFS · gain -4.8 dB · emolia-01494
(newsreading, formal)The idea that nurses and teachers are often pictured as women whereas doctors and lawyers
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 10.3s, EN.
EN_v7XM6MwWsOu_W000010 · in -14.0 dBFS · gain -6.0 dB · emolia-01494
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Confusion clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.32.
Nothing was asked of the other axis, and in fact Teasing drifts down from 1.00 to 0.36 (-0.63), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.01, then +0.08, then +0.04 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 72 s · en · emolia
k 5d_a -0.634d_b 0.321step_a 0.622step_b 0.181min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00008_S05640track EN_B00008_S05640total 72.3slevel spread 3.5 dBmax seam 3.5 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, moderately variable, some disfluency
(teasing, amusement, pride · normally alert, neutral tension, average clarity, casual)Did (low mumble) you get your crap out? What'd he say? I'm like, don't apologize bro. I'm like, you did what you absolutely needed to do.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as teasing, amusement, pride; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 7.4/10; 11.3s, EN.
EN_B00008_S05640_W000056 · in -18.3 dBFS · gain -1.7 dB · emolia-00409
(embarrassment, pleasure ecstasy, intoxication altered states of consciousness· normally alert, neutral tension, average clarity, casual)I'll never, I've tried like several times like, this is the time, I'm like, this is how, I always thought like at some point people have to make a conscious effort to enjoy sports and I've tried doing it multiple times and I'm like, oh god I want to leave so bad. Man. It's so long dude. So long. The crowd is fucked.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, pleasure ecstasy, intoxication altered states of consciousness; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 6.8/10; 18.7s, EN.
EN_B00008_S05640_W000057 · in -20.3 dBFS · gain +0.3 dB · emolia-00409
(intoxication altered states of consciousness, triumph, elation· normally alert, neutral tension, somewhat unclear, casual)Dude, I didn't beat, I didn't beat whatsoever. That's my number one thing is when me and Sydney get a room, we share a room, so I'm always worried. Whenever we check in, I feel like people think we're a gay couple, first off of that. We get in (ahem) and I'm like, hey, (ahem) uh, yeah, we got here a show here, no funny business here. Every time I'm like, I don't say anything, but in my head I'm like, and we come out of the room together and I'm always like,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, triumph, elation; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 10.0/10; 22.2s, EN.
EN_B00008_S05640_W000058 · in -20.1 dBFS · gain +0.1 dB · emolia-00409
(amusement, intoxication altered states of consciousness, teasing·energised, relaxed, average clarity, casual)Dude, it's, it haunts me the whole entire weekend. We're just coming out, we check in and I'm like, two queens, (chuckle) two queens. That's what you are. (chuckle) That's what the reservation's under. (chuckle)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, intoxication altered states of consciousness, teasing; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 2.6/10; 11.1s, EN.
EN_B00008_S05640_W000059 · in -18.3 dBFS · gain -1.7 dB · emolia-00409
(confusion, embarrassment, awe·normally alert, neutral tension, average clarity, casual)Dude, it was just, it was in a room. It was, I couldn't find where it was. I walked throughout the hotel, just, you know, I went to the gym and I'm still here and I'm like, dude, this is everywhere.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as confusion, embarrassment, awe; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 10.0/10; 8.6s, EN.
EN_B00008_S05640_W000060 · in -21.8 dBFS · gain +1.8 dB · emolia-00409
This chain comes from the one-sided rule: only Teasing had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Teasing clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.
Nothing was asked of the other axis, and in fact Pride drifts down from 0.96 to 0.61 (-0.34), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.12, then -0.01, then +0.02 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.39 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.32 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.39, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice
(pride, shame · normal-paced, subdued, neutral tension, casual)that. So my 16 team in New York, (low mumble) uh, this team I I helped coach in New York, (low mumble) um, recently had a player (low mumble) um whose brother was I'm not gonna give his name, but his brother is the number one pick in the NHL draft and just named the Canadian Olympic team, was kicked out for using a racial slur in a game. So I'm like, well, you know, that's not the precedent that that we said. I'm talking to coach, and that these aren't my decisions, these are his decisions, and he just we're talking, and and I say, What did he possibly say?
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride, shame; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 7.7/10; 28.3s, EN.
455836_00520368 · in -17.6 dBFS · gain -2.4 dB · podcast-06398
(disgust, malevolence malice, intoxication altered states of consciousness·slow, normally alert, neutral tension, monologue)Of all things that he possibly could have called the kid, he called the kid a ginger.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, malevolence malice, intoxication altered states of consciousness; style: monologue, didactic; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.1/10; 8.5s, EN.
455836_00523192 · in -15.7 dBFS · gain -4.3 dB · podcast-01346
(teasing, amusement, impatience and irritability·brisk, energised, tense, casual)What? Well, yeah, that's the problem with that, is he used the hard R. If you're gonna say it, use ginger, you know what I mean? He got he got removed from the golf
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, tense, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as teasing, amusement, impatience and irritability; style: casual, playful; below-average recording, some background noise; mildly explicit content; genuineness 5.5/6; vocal-burst blend 9.4/10; 7.6s, EN.
455836_00524151 · in -8.9 dBFS · gain -11.1 dB · podcast-00095
(affection, elation, amusement · brisk, energised, neutral tension, casual)Is Dan Collins, a wonderful man who also happens to be a card carrying ginger. He's like, I have never looked at it like hate speech.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as affection, elation, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 8.1/10; 8.6s, EN.
455836_00525504 · in -12.3 dBFS · gain -7.7 dB · podcast-00077
(teasing, amusement, disgust· brisk, energised, neutral tension, casual)No. Or sort of like that. He says, This guy's running around all game, chirping everybody about everything, calling everybody fat ass, midget, this, that, blah, blah, blah. And all our player says to him is like, get out of here, you ginger. Boom. What is proper terminology? Redhead? Yeah.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as teasing, amusement, disgust; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 4.4/6; vocal-burst blend 4.1/10; 15.3s, EN.
455836_00526376 · in -11.0 dBFS · gain -9.0 dB · podcast-00107
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Concentration clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 98 % of clips in this corpus. That is a total rise of 0.28.
Nothing was asked of the other axis, and in fact Disgust drifts down from 0.89 to 0.33 (-0.56), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.00, then +0.16, then +0.00 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 60 s · en · emolia
k 5d_a -0.564d_b 0.283step_a 0.322step_b 0.159min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_QucAImnjFbotrack EN_QucAImnjFbototal 59.9slevel spread 3.0 dBmax seam 3.0 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, slightly dark, balanced body, average recording, measured, slightly relaxed, steady
(normally alert, some disfluency, somewhat unclear, monologue)If H and K are two subgroups, immediately you can say that HK is also a subgroup of G.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, no background noise; genuineness 3.7/6; vocal-burst blend 1.5/10; 5.6s, EN.
EN_QucAImnjFbo_W000082 · in -19.7 dBFS · gain -0.3 dB · emolia-02136
(normally alert, some disfluency, slurred, monologue)Because by definition Reg K is equal to the products of the form Reg K. It is governed that G is an abelian group, so Reg K is equal to K of H.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.6/10; 8.8s, EN.
EN_QucAImnjFbo_W000083 · in -18.7 dBFS · gain -1.3 dB · emolia-02136
(normally alert, some disfluency, slurred, monologue)Because by definition Reg K is equal to the products of the form Reg K. It is governed that G is an abelian group, so Reg K is equal to K of H.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.6/10; 8.8s, EN.
EN_QucAImnjFbo_W000083 · in -18.7 dBFS · gain -1.3 dB · emolia-02136
(concentration·subdued, frequent disfluency, somewhat unclear, monologue)And here, it is the set of all products of the form K into H, where K belongs to K H belongs to H. So it is nothing but K H. So immediately we get H K is equal to K H. Therefore by our necessary and sufficient condition, our H K is a subgroup of G.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 18.1s, EN.
EN_QucAImnjFbo_W000084 · in -21.7 dBFS · gain +1.7 dB · emolia-02136
(concentration · subdued, frequent disfluency, somewhat unclear, monologue)And here, it is the set of all products of the form K into H, where K belongs to K H belongs to H. So it is nothing but K H. So immediately we get H K is equal to K H. Therefore by our necessary and sufficient condition, our H K is a subgroup of G.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 18.1s, EN.
EN_QucAImnjFbo_W000084 · in -21.7 dBFS · gain +1.7 dB · emolia-02136
This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Shame clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude barely moves at all, sitting near 0.98 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.16, then -0.07, then +0.03, then +0.17 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.45 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, fairly steady, frequent disfluency, somewhat unclear
(thankfulness gratitude, embarrassment, contentment · measured, normally alert, slightly relaxed, monologue)Tam je tam je ten jediný bod. Já jsem měl minulý rok na Check Stage BTC práčem to vysvětloval ten princip. (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, embarrassment, contentment; style: monologue, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.9/10; 10.5s, CS.
476692_00413524 · in -19.7 dBFS · gain -0.3 dB · podcast-03458
(contemplation, thankfulness gratitude · measured, normally alert, relaxed, casual)(low mumble) (low mumble) Zdá se mi to. (ahem) I oni je hodně hodně ty integrace s nostrem ještě, že tam je to takový fakt dobrý váb tam, že se tam jako rychle vyvíjí nějaký věci
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, thankfulness gratitude; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.5/10; 14.4s, CS.
476692_00414568 · in -20.1 dBFS · gain +0.1 dB · podcast-03457
(longing, fatigue exhaustion, contemplation · measured, normally alert, relaxed, casual)a zajímavé věci. Takže ta kombinace lightning, nos se mi zdá strašně fascinující.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; reads as longing, fatigue exhaustion, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 2.5/10; 9.9s, CS.
476692_00416004 · in -20.2 dBFS · gain +0.2 dB · podcast-03455
(relief, triumph, hope enthusiasm optimism·normal-paced, normally alert, slightly relaxed, monologue)Dá se teda říct, že jako potenciální škálování Bitcoinů je takový, že máš base layer. (ahem) To je ten Settlement layer, on chain transakce opravdu jenom (ahem) jako prostor masivní přesuny, něco jako na dnešním Swiftu. (low mumble) Pak Lightning. Tam je trošičku problém. Prostě (low mumble) ta infrastruktura, musíš manažovat i kanály a tak. (ahem) A pak nad Lightning by tedy byl i caš (low mumble) pro normální lidi.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, triumph, hope enthusiasm optimism; style: monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 8.1/10; 25.5s, CS.
476692_00417343 · in -19.3 dBFS · gain -0.7 dB · podcast-04395
(shame, thankfulness gratitude, embarrassment· normal-paced, subdued, relaxed, casual)No, ten light a je to do toho nevidím to, Michal Nováka si by věděl mnohem líp, ale mám pocit, že časem pak ten Lightning bude interbank (low mumble) settlement spíš a nad tím bude třeba něco jiného. (low mumble)
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, thankfulness gratitude, embarrassment; style: casual, monologue; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 8.9/10; 16.6s, CS.
476692_00420016 · in -19.6 dBFS · gain -0.3 dB · podcast-03447
Impatience and Irritability ↑ (unconstrained axis: Affection)emotion__B1__T0.25__C0.20__INTERNAL · #19
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Impatience and Irritability clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.28.
Nothing was asked of the other axis, and in fact Affection drifts down from 0.95 to 0.79 (-0.15), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are -0.03, then +0.13, then +0.19 — not a clean run: step 1 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 40 s · fr · emolia
k 4d_a -0.152d_b 0.280step_a 0.095step_b 0.185min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_keKB7FIXgHMtrack FR_keKB7FIXgHMtotal 39.9slevel spread 4.4 dBmax seam 4.4 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, average recording, quiet background
(affection · normal-paced, normally alert, slightly relaxed, monologue)Mais, (low mumble) euh, voilà, tout ça pour vous dire que c'est, euh, (low mumble) vraiment à nous de faire quelque chose. C'est à nous de dire stop, (low mumble) hein.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as affection; style: monologue, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.1/10; 6.7s, FR.
FR_keKB7FIXgHM_W000017 · in -16.8 dBFS · gain -3.2 dB · emolia-02843
(confusion, embarrassment, pride·measured, very low-energy, neutral tension, casual)c'est à nous de dire, non, je refuse, je, c'est à nous de dire, (ahem) euh, voilà, (low mumble) euh, vous, vous bêtisez, il y en a marre, euh, (ahem) maintenant, on y va, (low mumble) hein. Et la deuxième, moi, qui a sauté, c'est celle-ci. Et celle-ci, elle dit,
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as confusion, embarrassment, pride; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.0/10; 15.3s, FR.
FR_keKB7FIXgHM_W000018 · in -21.3 dBFS · gain +1.3 dB · emolia-02843
(relief, triumph·normal-paced, normally alert, slightly relaxed, casual)Après l'effort, le réconfort, vous récoltez les fruits de votre travail, débarrassez-vous des situations qui vous accablent. Voilà. Ah, bah, écoutez, hein, je crois que c'est clair.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, triumph; style: casual, dramatic; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.9/10; 9.1s, FR.
FR_keKB7FIXgHM_W000019 · in -19.6 dBFS · gain -0.4 dB · emolia-02843
(impatience and irritability, anger, malevolence malice· normal-paced, normally alert, neutral tension, conversational)(surprised gasp) hein, donc, moi, je vous dis, moi, mon cheval de bataille, c'est les enfants, mais pour tous, (ahem) tous, il y en a qui se battent pour les personnes âgées, il y en a qui se (ahem) battent
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as impatience and irritability, anger, malevolence malice; style: conversational, casual; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 6.5/10; 8.4s, FR.
FR_keKB7FIXgHM_W000020 · in -16.9 dBFS · gain -3.1 dB · emolia-02843
This chain comes from the one-sided rule: only Amusement had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Amusement clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Awe drifts down from 0.99 to 0.47 (-0.52), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.13, then +0.00 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.20 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 90 s · en · evasnippets
k 4d_a -0.519d_b 0.292step_a 0.519step_b 0.159min_cos_consec —min_cos_anchor —dataset evasnippetslang enspeaker cond_podcastt_01_02_0001_5track cond_podcastt_01_02_0001_5total 90.3slevel spread 6.5 dBmax seam 5.9 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · average recording, quiet background, brisk, energised, neutral tension, moderately variable, some disfluency, average clarity
(awe, interest, astonishment surprise · very wide pitch range, normal breath, casual, dramatic)like kind of goes more into like a that was a really stupid thing that you did. Like there is there's the scene where like the the the some someone is attacking the the building that she's in and everything's going down and everything's freaking out. And she's just like standing there as everything is going down and like actively moves towards the danger. And it's like what how is she a smart how is she supposed to be the smart
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, neutral openness; reads as awe, interest, astonishment surprise; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 5.4/10; 22.9s.
cond_podcastt_01_02_0001_599717_00046504 · in -21.6 dBFS · gain +1.6 dB · evasnippets-00008
(interest, disappointment, impatience and irritability·wide pitch range, light breath, casual, conversational)You don't understand how much I love Bioware. I I I really do love Bioware's games. (ahem) Um, but Bioware did some shit that has pissed people off a lot in a good way. Um, (ahem) in the fact that in the latest Dragon Age uh (exhausted groan) installment that they did, Dragon Age Inquisition, they added in a spunky
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as interest, disappointment, impatience and irritability; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 9.3/10; 19.2s.
cond_podcastt_01_02_0001_599717_00156208 · in -21.0 dBFS · gain +1.0 dB · evasnippets-00280
(amusement, triumph, elation· wide pitch range, normal breath, casual, playful)It's pretty much the exact opposite of this situation happened to me today. I went out with some of the coworkers. I was the only male in the group. And they like one of them brought up gaming and I was like, Oh, they know so much more about this shit than me. I (chuckle) was like I'm like, oh, I play the soccer car game. I I started Minecraft recently. I don't know what the fuck I'm doing. I'm helping with a path.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, triumph, elation; style: casual, playful; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.2/10; 23.9s.
cond_podcastt_01_02_0001_599717_00217904 · in -15.1 dBFS · gain -4.9 dB · evasnippets-00299
(amusement, triumph, elation · wide pitch range, normal breath, casual, playful)It's pretty much the exact opposite of this situation happened to me today. I went out with some of the coworkers. I was the only male in the group. And they like one of them brought up gaming and I was like, Oh, they know so much more about this shit than me. I (chuckle) was like I'm like, oh, I play the soccer car game. I I started Minecraft recently. I don't know what the fuck I'm doing. I'm helping with a path.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, triumph, elation; style: casual, playful; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.2/10; 23.9s.
cond_podcastt_01_02_0001_599717_00217904 · in -15.1 dBFS · gain -4.9 dB · evasnippets-00299