Manifest tier. voicenet, rule VN1, T=0.6, step cap 0.25. Population 3,019,723 chains (40,917 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 2,069,094.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.voicenet__VN1__T0.60__C0.25__INTERNAL — population 3,019,723 chains (40,917 h). SHAREABLE variant: 2,069,094. Filter.rule=='VN1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.6 and step_b<=0.25 Sampled from 20,076 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with metallic quality (METL) at the very top of the range — 0.95, higher than 95 % of clips in this corpus — and works its way down to below average at 0.30, lower than 70 % of clips in this corpus. That is a total fall of 0.65.
It takes 5 clips to get there. Clip to clip the moves are -0.12, then -0.08, then -0.24, then -0.20 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · ko · emolia
hear it un-normalised (raw levels, max seam 2.3 dB)
k 5d_a -0.647d_b -0.647step_a 0.243step_b 0.243min_cos_consec —min_cos_anchor —dataset emolialang kospeaker KO_fFojNEkMStItrack KO_fFojNEkMStItotal 50.7slevel spread 3.0 dBmax seam 2.3 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-bright, average recording, quiet background
(helplessness, longing, sadness · measured, normally alert, slightly relaxed, authoritative)오늘은 투소감사 주의입니다. 예수를 믿는 저와 여러분들은
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, longing, sadness; style: authoritative, didactic; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.4/10; 5.5s, KO.
KO_fFojNEkMStI_W000004 · in -17.5 dBFS · gain -2.5 dB · emolia-03146
(bitterness, disappointment, longing · measured, normally alert, neutral tension, didactic)이렇게 말할 수 없는 복이 감사신앙이라서 마기는 그 사실을 잘 압니다. 마기는 우리 신자들이 감사신앙으로 무장이 있을 때 세상을 그땐 이기더라는 거예요.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as bitterness, disappointment, longing; style: didactic, storytelling; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.1/10; 14.2s, KO.
KO_fFojNEkMStI_W000005 · in -19.8 dBFS · gain -0.2 dB · emolia-03146
(longing, anger, disgust·fast, energised, neutral tension, dramatic)그래서 어떻게 하면 그 감사신앙을 뺏어갈까? 감사신앙을 무너뜨릴까?
full caption & clip details
An elderly masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, some disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, fairly guarded; reads as longing, anger, disgust; style: dramatic, storytelling; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.5/10; 6.3s, KO.
KO_fFojNEkMStI_W000006 · in -20.3 dBFS · gain +0.3 dB · emolia-03146
(contemplation, disappointment·measured, normally alert, neutral tension, conversational)그래서 이 감사, 주의를 같은 경우에 굉장히 부담스럽게 만드는 것, 이거 사단에 하는 겁니다. 여러분.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as contemplation, disappointment; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.3/10; 7.8s, KO.
KO_fFojNEkMStI_W000007 · in -20.5 dBFS · gain +0.5 dB · emolia-03146
(contemplation, longing, bitterness· measured, normally alert, slightly relaxed, didactic)어, (low mumble) 혹시 우리 가운데 오늘 주수 감사한금 내는데 얼마 낼까? 이걸로 혹시 싸운 부부 혹시 없습니까? 제가 우수게 얘기인지 모르지만, (ahem) 언젠가 이런 얘기를 들었습니다. 우리 교우 집사님 부부 가운데.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, longing, bitterness; style: didactic, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.5/10; 16.2s, KO.
KO_fFojNEkMStI_W000008 · in -20.4 dBFS · gain +0.3 dB · emolia-03146
RANG — pitch range used ↑voicenet__VN1__T0.60__C0.25__INTERNAL · #2
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with pitch range used (RANG) below average — 0.31, lower than 69 % of clips in this corpus — and ends with it at the very top of the range at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.64.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.25, then +0.12, then +0.18 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 33 s · zh · emolia
hear it un-normalised (raw levels, max seam 2.7 dB)
k 5d_a 0.641d_b 0.641step_a 0.248step_b 0.248min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00077_S09411track ZH_B00077_S09411total 32.7slevel spread 4.7 dBmax seam 2.7 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, moderate pitch range, light breath
(measured, fairly steady, some disfluency, monologue)还是说老人也觉得这件事情变得不像以前那么重视和在意了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 4.3/10; 5.4s, ZH.
ZH_B00077_S09411_W000224 · in -24.0 dBFS · gain +4.0 dB · emolia-04043
(fatigue exhaustion, disappointment, sadness·normal-paced, fairly steady, some disfluency, monologue)(ahem) 前两天我爸过完生日之后,我又给他打一个电话,就有问他。我说哎说起来我也挺惭愧,挺不好意思,也没给你办个什么六十大寿什么的。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, disappointment, sadness; style: monologue, authoritative; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.8/10; 9.0s, ZH.
ZH_B00077_S09411_W000225 · in -24.1 dBFS · gain +4.1 dB · emolia-04043
(measured, fairly steady, little disfluency, monologue)他会问我这么一句,说怎么叫半个六十大寿,你要叫谁来啊?我后来一想,我能叫谁来呢?特别是我们这一代是独生子女。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 5.0/10; 7.6s, ZH.
ZH_B00077_S09411_W000226 · in -22.0 dBFS · gain +2.0 dB · emolia-04043
(sourness, impatience and irritability, contempt·brisk, fairly steady, little disfluency, dramatic)就是你好像要凑齐要由你来操办这件事情,你本身就是排斥的。
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, impatience and irritability, contempt; style: dramatic, authoritative; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.7/10; 5.5s, ZH.
ZH_B00077_S09411_W000227 · in -19.3 dBFS · gain -0.7 dB · emolia-04043
(sourness ·normal-paced, moderately variable, some disfluency, conversational)(chuckle) 不不不,就是我现在就跟你说你说你爸要过六十岁生日。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sourness; style: conversational, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.2/10; 4.6s, ZH.
ZH_B00077_S09411_W000228 · in -20.3 dBFS · gain +0.3 dB · emolia-04043
This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with background noise level (BKGN) at the very top of the range — 0.93, higher than 93 % of clips in this corpus — and works its way down to below average at 0.29, lower than 71 % of clips in this corpus. That is a total fall of 0.64.
It takes 5 clips to get there. Clip to clip the moves are -0.10, then -0.17, then -0.17, then -0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.07 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 87 s · en · podcast
hear it un-normalised (raw levels, max seam 3.3 dB)
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normal-paced, slightly relaxed, fairly steady, some disfluency, average clarity
(contemplation, concentration · normally alert, casual, conversational)The other way I think about the future of theological education is to move in incrementally more and more toward interactive, i.e. student-centered kind of learning,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, concentration; style: casual, conversational; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.4/10; 12.0s, EN.
104834_00194560 · in -13.7 dBFS · gain -6.3 dB · podcast-00777
(interest, contemplation, concentration · normally alert, casual, monologue)integrated learning that's you know, not one course that's the life of the mind and a different course that's the flourishing of the soul and a different course that's development of your professional skills, but rather integrating all of that and also integrated with the rest of the world so that it's part of art and science and environment, and then contextualized, which we spoke about earlier, deeper integration with religious and other communities.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, contemplation, concentration; style: casual, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 3.9/10; 22.5s, EN.
104834_00195752 · in -14.8 dBFS · gain -5.2 dB · podcast-00784
(doubt· normally alert, didactic, casual)Contextual education as a thread throughout, not an episode in students' theological education. And I think that will, among other things, make us more accountable to issues of justice.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt; style: didactic, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 11.4s, EN.
104834_00198004 · in -16.1 dBFS · gain -3.9 dB · podcast-00781
(contemplation, interest, contentment·very low-energy, monologue, whispered)I love the way that you were thinking, Rachel, in terms of thinking what you were thinking as small, but in some ways that's what we're moving back to, right? How to create those essential communities and community ties and partnerships to do the common good.
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contemplation, interest, contentment; style: monologue, whispered; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 3.2/10; 13.6s, EN.
104834_00212808 · in -19.3 dBFS · gain -0.7 dB · podcast-00277
(thankfulness gratitude, contemplation, affection·subdued, casual, monologue)As those communities really struggle with what their vocation is in the world, about their ministry and they're being challenged to think about not just the people who are members of their community, but who's in their neighborhood, who are their next door neighbors, who are the people in their city. I think you know, seminaries have the same kind of responsibility. And I think those who are thinking strategically within theological institutions really do have to think about who do we really want to teach?
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, contemplation, affection; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.8/10; 27.4s, EN.
104834_00215016 · in -16.0 dBFS · gain -4.0 dB · podcast-00776
This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with resonance: throat (R_THRT) low — 0.12, lower than 88 % of clips in this corpus — and ends with it above average at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.62.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.17, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · en · podcast
hear it un-normalised (raw levels, max seam 2.2 dB)
Unchanged across all 4 clips: a young adult feminine voice · moderately variable, wide pitch range
(disgust, affection, embarrassment · brisk, normally alert, neutral tension, casual)the word. Everyone was using Brody. It's like that Grody, oh that's your mastery. Oh my gosh, do you remember Alex? He is just like so grody, so no, I will not date him.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as disgust, affection, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.9/6; vocal-burst blend 4.4/10; 10.8s, EN.
841972_00082176 · in -22.9 dBFS · gain +2.9 dB · podcast-04313
(longing, contemplation, relief·normal-paced, normally alert, relaxed, casual)Well groovy though too, like when we were in the nineties, you know, we were recycling the sixties as well. So like (surprised gasp) and even like groovy,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, contemplation, relief; style: casual, conversational; average recording, no background noise; genuineness 4.4/6; vocal-burst blend 5.2/10; 12.3s, EN.
841972_00090176 · in -24.1 dBFS · gain +4.1 dB · podcast-04316
(affection, infatuation, pleasure ecstasy· normal-paced, very low-energy, neutral tension, casual)Yeah, that's true. True. We had a lost power.
full caption & clip details
A child feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is slightly cool, bright, fairly smooth, thin; slurred, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as affection, infatuation, pleasure ecstasy; style: casual, storytelling; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.4/10; 7.9s, EN.
841972_00095236 · in -22.4 dBFS · gain +2.5 dB · podcast-04326
(longing, embarrassment, amusement· normal-paced, normally alert, slightly relaxed, casual)See, we didn't really have slam book uh once I moved to Idaho 'cause we were teeny tiny little town.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, embarrassment, amusement; style: casual, conversational; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 4.0/10; 5.8s, EN.
841972_00106128 · in -24.6 dBFS · gain +4.6 dB · podcast-04307
This is a VoiceNet dimension, not an emotion: style: dramatic (S_DRAM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with style: dramatic (S_DRAM) above average — 0.75, higher than 75 % of clips in this corpus — and works its way down to low at 0.14, lower than 86 % of clips in this corpus. That is a total fall of 0.60.
It takes 5 clips to get there. Clip to clip the moves are +0.03, then -0.25, then -0.18, then -0.21 — not a clean run: step 1 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.23 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.23, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 71 s · es · podcast
hear it un-normalised (raw levels, max seam 6.6 dB)
A child feminine voice; delivery is energised, normal-paced, relaxed, volatile; timbre is slightly cool, dark, slightly rough, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, teasing, affection; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.4/10; 5.4s, ES.
795713_00218448 · in -19.2 dBFS · gain -0.8 dB · podcast-00425
(affection, disgust, infatuation·fast, energised, slightly relaxed, casual)claro que sí, fíjate que quiero campingists who are with us, a
full caption & clip details
A child feminine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is cool, bright, fairly smooth, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly submissive, neutral openness; reads as affection, disgust, infatuation; style: casual, playful; below-average recording, some background noise; genuineness 4.0/6; vocal-burst blend 5.1/10; 4.4s, ES.
795713_00223064 · in -23.5 dBFS · gain +3.5 dB · podcast-00417
(thankfulness gratitude, bitterness, jealousy and envy·measured, normally alert, neutral tension, monologue)La capacidad tienen muchos in the pueblo, dice, que lo quieran hacer is otra cosa, eso es muy importante. Los siguientes problemas que tiene nuestro pueblo nuevoñas, como el tráfico vehicular, los tales tuttuqueros.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, bitterness, jealousy and envy; style: monologue, casual; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 6.9/10; 27.4s, ES.
795713_00388880 · in -16.9 dBFS · gain -3.1 dB · podcast-00418
(affection, hope enthusiasm optimism, thankfulness gratitude ·normal-paced, normally alert, neutral tension, monologue)también queremos aprovechar también para saludar a don William Salvador que está conectado por ahí y también te envío un cordial saludo hacia tu persona que dice buenas noches quiero felicitar al licenciado con sus proyectos lástima que no es tan conocido pero siga para adelante siempre hay una segunda oportunidad o sea es muy bonito pues pero a la larga vos haces de que tenés que trabajar
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection, hope enthusiasm optimism, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 6.9/10; 24.6s, ES.
795713_00418923 · in -16.8 dBFS · gain -3.2 dB · podcast-03057
(elation, pride, pleasure ecstasy· normal-paced, normally alert, neutral tension, casual)Felicitaciones al ideal, BNB. Dios lo bendiga, lo estoy viendo desde los Estados Unidos. Estoy en Polonovias, qué bonito, qué bonito, la verdad.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation, pride, pleasure ecstasy; style: casual; below-average recording, quiet background; genuineness 5.8/6; vocal-burst blend 8.1/10; 8.2s, ES.
795713_00445096 · in -18.0 dBFS · gain -2.0 dB · podcast-00410
This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with background noise level (BKGN) low — 0.15, lower than 85 % of clips in this corpus — and ends with it high at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.71.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.13, then +0.19, then +0.25 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 47 s · en · emolia
k 5d_a 0.712d_b 0.712step_a 0.249step_b 0.249min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_bO__gVg7t5Itrack EN_bO__gVg7t5Itotal 47.4slevel spread 1.6 dBmax seam 0.7 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, minimal breath, formal, newsreading)For the first issue, Gold obtained stories by several well-known authors, including Isaac Asimov, Fritz Lieber, and Theodore Sturgeon, as well as Part 1 of Time Quarry by Clifford D. Simak
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.1s, EN.
EN_bO__gVg7t5I_W000032 · in -15.4 dBFS · gain -4.6 dB · emolia-01772
(triumph· steady, light breath, formal, narration)Along with an essay by Gold, Galaxy's premier issue introduced a book review column by anthologist Graf Conklin, which ran until 1955, and a Willie Lay science column
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.7s, EN.
EN_bO__gVg7t5I_W000033 · in -15.4 dBFS · gain -4.6 dB · emolia-01772
(thankfulness gratitude·fairly steady, light breath, formal, newsreading)Gold sought to implement high-quality printing techniques, though the quality of the available paper was insufficient for the full benefits to be seen
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 8.2s, EN.
EN_bO__gVg7t5I_W000034 · in -14.7 dBFS · gain -5.3 dB · emolia-01772
(steady, light breath, formal, narration)Within months, the outbreak of the Korean War led to paper shortages that forced gold to find a new printer, Robert M. Ginn
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.7s, EN.
EN_bO__gVg7t5I_W000035 · in -14.5 dBFS · gain -5.5 dB · emolia-01772
(steady, light breath, formal, monologue)The new paper was of even lower quality, a disappointment to gold.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.2/10; 4.2s, EN.
EN_bO__gVg7t5I_W000036 · in -13.8 dBFS · gain -6.2 dB · emolia-01772
VOLT — loudness / volume ↑voicenet__VN1__T0.60__C0.25__INTERNAL · #7
This is a VoiceNet dimension, not an emotion: loudness / volume (VOLT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with loudness / volume (VOLT) low — 0.14, lower than 86 % of clips in this corpus — and ends with it above average at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.60.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.10, then +0.08, then +0.21 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 21 s · snippets
k 5d_a 0.600d_b 0.600step_a 0.215step_b 0.215min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch275_part0_batch275_patrack batch275_part0_batch275_patotal 20.6slevel spread 0.9 dBmax seam 0.5 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, light breath
(measured, fairly steady, no disfluency, formal)Fish and other marine animals benefit too.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; style: formal, monologue; good recording, no background noise; no dominant emotion; genuineness 0.0/6; vocal-burst blend 3.2/10; 3.0s.
batch275_part0_batch275_part0_chunk_943_1_722099 · in -25.8 dBFS · gain +5.8 dB · snippets-00913
(emotional numbness·normal-paced, fairly steady, no disfluency, narration)The Patagonia park shows what challenges have to be overcome to turn back the clock.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: narration, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 5.2s.
batch275_part0_batch275_part0_chunk_943_1_722110 · in -25.9 dBFS · gain +5.9 dB · snippets-00913
(teasing, doubt·measured, fairly steady, no disfluency, monologue)Or is it possible to make them do a little more?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as teasing, doubt; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 6.2/10; 3.2s.
batch275_part0_batch275_part0_chunk_943_1_722170 · in -26.1 dBFS · gain +6.1 dB · snippets-00913
(fear·normal-paced, fairly steady, no disfluency, narration)with desert dust, for example, the ice nuclei would grow faster
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear; style: narration, storytelling; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.7/10; 4.8s.
batch275_part0_batch275_part0_chunk_943_1_722206 · in -26.6 dBFS · gain +6.6 dB · snippets-00913
(distress· normal-paced, moderately variable, some disfluency, casual)And that's why people, young people and old people are in the streets.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as distress; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.0/10; 3.8s.
batch275_part0_batch275_part0_chunk_943_1_722283 · in -26.7 dBFS · gain +6.7 dB · snippets-00913
This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with vocal focus (FOCS) low — 0.24, lower than 76 % of clips in this corpus — and ends with it high at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.62.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.17, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 49 s · en · emolia
k 4d_a 0.621d_b 0.621step_a 0.239step_b 0.239min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00010_S04722track EN_B00010_S04722total 49.3slevel spread 2.7 dBmax seam 1.6 dB
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, fairly steady, some disfluency
(pain · normal-paced, neutral tension, casual, conversational)(low mumble) Um, I think it's the downgrade though in the lot in the market doesn't it's it's trending up. You can get three and a half. So I have best bet 10 plus three and a half.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.0/10; 7.0s, EN.
EN_B00010_S04722_W000025 · in -23.5 dBFS · gain +3.5 dB · emolia-00449
(relief·brisk, slightly relaxed, casual, monologue)Norland's defense has good stats. They have faced some limited offenses this year, (low mumble) uhm, and Tampa Bay's finally flopped against an elite team, but they've competed otherwise, and this is not an elite team. Their offense looks much better with Dave Canales running the show. I think they're gonna...
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief; style: casual, monologue; average recording, no background noise; genuineness 3.4/6; vocal-burst blend 4.8/10; 10.5s, EN.
EN_B00010_S04722_W000026 · in -25.1 dBFS · gain +5.1 dB · emolia-00449
(hope enthusiasm optimism, pride, elation· brisk, neutral tension, casual, conversational)show out here and have a think Tampa Bay deserves, especially if you're getting plus three and a half is going to be a close game and they've played well enough that they should get more respect (ahem) coming out of that game against Philly, which is that's just gonna happen against Philly.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, pride, elation; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 9.8/10; 11.5s, EN.
EN_B00010_S04722_W000027 · in -26.1 dBFS · gain +6.1 dB · emolia-00449
(disappointment, triumph, interest· brisk, slightly relaxed, casual, monologue)And Saints are perceived to have this great home-field advantage, but that's not really hasn't been the case. (low mumble) Um, the last couple years I tried, you know, I'd give any unique home-field advantage numbers to every team when I'm building my Power Ratings. And, (ahem) um, they just, they're home and away stats. They're not that different. So I don't think they have a great home-field advantage at this point. They're not, they're, they're basically the same, not, pretty close to the same team home and away at this point.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, triumph, interest; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 6.8/10; 19.8s, EN.
EN_B00010_S04722_W000028 · in -24.6 dBFS · gain +4.6 dB · emolia-00449
This is a VoiceNet dimension, not an emotion: style: technical (S_TECH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with style: technical (S_TECH) high — 0.80, higher than 80 % of clips in this corpus — and works its way down to low at 0.13, lower than 87 % of clips in this corpus. That is a total fall of 0.68.
It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.21, then -0.15, then -0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.43 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.43 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.43, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, light breath
(confusion, amusement, teasing · normally alert, slightly relaxed, fairly steady, conversational)Und dann auch immer noch die Strophe ist dann, also fängt mit dem F immer jeder Vers mit einem Fünffiertel an, geht dann mit einem Vierviertel weiter und der Refrain ist ein Dreiviertel. Und man merkt's nicht.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, amusement, teasing; style: conversational, dramatic; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 0.0/10; 9.0s, DE.
829141_00181543 · in -16.3 dBFS · gain -3.7 dB · podcast-05923
(normally alert, slightly relaxed, fairly steady, conversational)Das ist sehr krass, sonst sonst bringen einen so fünf Vierteltakte oder sieben Viertel bringt einen sonst immer so ein bisschen raus.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 8.6s, DE.
829141_00184472 · in -16.1 dBFS · gain -3.9 dB · podcast-00707
(contentment, affection, intoxication altered states of consciousness· normally alert, neutral tension, moderately variable, casual)(ahem) krass, man merkt es (low mumble) wirklich nicht.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contentment, affection, intoxication altered states of consciousness; style: casual, conversational; good recording, quiet background; genuineness 4.8/6; vocal-burst blend 1.9/10; 24.3s, DE.
829141_00185392 · in -17.3 dBFS · gain -2.7 dB · podcast-00027
(embarrassment, sadness, intoxication altered states of consciousness · normally alert, slightly relaxed, fairly steady, casual)Ja, Hammer. (low mumble) Es geht so ein bisschen um Nachkriegstrauma und Angst vor neuen Kriegen, also um so ein kollektives Gefühl der Kriegsgeneration. Also eben davor ist ja sein Papa im Krieg gestorben und (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, sadness, intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 0.3/10; 13.4s, DE.
829141_00199024 · in -16.5 dBFS · gain -3.5 dB · podcast-04052
(interest, disappointment, contemplation·very low-energy, slightly relaxed, fairly steady, casual)dieses kollektive Gefühl, dieses Nachkriegstrauma färbt halt unmittelbar auf Pink ab, als der Krieg endet. (low mumble) Und weil er eben auch viel mitbekommt über seine Mutter und über der Generation, über die Generation über ihm so ein bisschen. Ja, und es ist, wie gesagt, jetzt kein Song, der so krass doll diese Story weiterbringt, aber es ist ein weiterer Aspekt, der so ein bisschen die Zeit beschreibt, in der sich Pink auch bewegt.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, disappointment, contemplation; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.7/10; 25.4s, DE.
829141_00200364 · in -17.1 dBFS · gain -2.9 dB · podcast-04349
This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with disfluency (DFLU) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to below average at 0.30, lower than 70 % of clips in this corpus. That is a total fall of 0.63.
It takes 5 clips to get there. Clip to clip the moves are -0.18, then -0.11, then -0.19, then -0.15 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 68 s · lt · eurospeech
k 5d_a -0.635d_b -0.635step_a 0.187step_b 0.187min_cos_consec —min_cos_anchor —dataset eurospeechlang ltspeaker lithuania_lithuania_0_1305track lithuania_lithuania_0_1305total 68.2slevel spread 2.0 dBmax seam 1.7 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-bright, average recording, normal breath
(malevolence malice · normal-paced, normally alert, slightly relaxed)Ir (low mumble) iš tikrųjų (ahem) bus dar kartą tikrinami tie patys ruožai, apie kuriuos aš jau kalbėjau, kartu mes (low mumble) esame numatę (low mumble) atrankiniu būdu
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.3/10; 11.7s, LT.
lithuania_lithuania_0_13052010_2836944_2848688 · in -15.2 dBFS · gain -4.8 dB · eurospeech-01820
(anger, distress, disgust· normal-paced, normally alert, slightly relaxed, authoritative)yra ir bus, jeigu mes iš esmės nepakeisime sistemos. Todėl mes (low mumble) parengėme visą paketą pasiūlymų, kaip reikėtų sustiprinti (low mumble) atiduodamų objektų
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, distress, disgust; style: authoritative; average recording, some background noise; genuineness 3.0/6; vocal-burst blend 1.6/10; 11.9s, LT.
lithuania_lithuania_0_13052010_2864592_2876528 · in -15.8 dBFS · gain -4.2 dB · eurospeech-01820
(concentration, triumph, bitterness· normal-paced, energised, slightly relaxed, authoritative)Pirmiausia yra visiškai neteisinga, kad (low mumble) Lietuvos automobilių kelių direkcija faktiškai atsako ir dalyvauja visuose proceso etapuose - nuo pat projektavimo iki pat objekto atidavimo.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as concentration, triumph, bitterness; style: authoritative, cartoonish; average recording, some background noise; genuineness 2.0/6; vocal-burst blend 1.0/10; 16.7s, LT.
lithuania_lithuania_0_13052010_2876528_2893184 · in -14.4 dBFS · gain -5.6 dB · eurospeech-01820
(fear, malevolence malice, bitterness ·brisk, energised, slightly relaxed, authoritative)kad yra samdomos įmonės, kurios atlieka techninę priežiūrą ir kurios turėtų prisiimti atsakomybę už techninę priežiūrą, jeigu po to atiduodant objektą nustatyta, kad blogai padarytas darbas,
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as fear, malevolence malice, bitterness; style: authoritative, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.4/10; 12.4s, LT.
lithuania_lithuania_0_13052010_2905632_2918064 · in -14.7 dBFS · gain -5.3 dB · eurospeech-01820
(bitterness, triumph, anger· brisk, highly aroused, tense, authoritative)bet mes pasidomėjome, kad tokių faktų, kad techninės priežiūros įmonės prisiimtų atsakomybę, yra pavieniai atvejai, jų beveik nėra ir jos net niekuo nerizikuoja. Pasirodo, iš jų negalima atimti licencijos,
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, some disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as bitterness, triumph, anger; style: authoritative, dramatic; average recording, some background noise; genuineness 2.3/6; vocal-burst blend 3.7/10; 14.9s, LT.
lithuania_lithuania_0_13052010_2918064_2932944 · in -16.4 dBFS · gain -3.6 dB · eurospeech-01820
This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with warmth (WARM) high — 0.88, higher than 88 % of clips in this corpus — and works its way down to below average at 0.27, lower than 73 % of clips in this corpus. That is a total fall of 0.61.
It takes 4 clips to get there. Clip to clip the moves are -0.15, then -0.25, then -0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a child masculine voice · neutral-toned, average recording
(affection, sadness, longing · slow, very low-energy, relaxed, casual)shout out to LaShawn, another young one. I think he's only a year older than you, maybe.
full caption & clip details
A child masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as affection, sadness, longing; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.5/10; 6.4s, EN.
851673_00077792 · in -23.2 dBFS · gain +3.2 dB · podcast-01350
(contemplation, fatigue exhaustion, sadness · slow, lethargic, relaxed, monologue)a few years ago, I was going through (low mumble) um
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, vulnerable; reads as contemplation, fatigue exhaustion, sadness; style: monologue, casual; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.6/10; 6.8s, EN.
851673_00078840 · in -26.4 dBFS · gain +6.4 dB · podcast-04246
(shame, fear, thankfulness gratitude· slow, very low-energy, relaxed, whispered)a growth spurt that some people. (surprised gasp) As spiritual seekers call dark night of the soul. And I wasn't in a great place.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as shame, fear, thankfulness gratitude; style: whispered, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.2/10; 11.7s, EN.
851673_00079520 · in -25.8 dBFS · gain +5.8 dB · podcast-04241
(fatigue exhaustion, pride, triumph·measured, very low-energy, neutral tension, casual)And I was spending a lot of late nights. And I was out of the game. I took a break. I was burnt out from taking clients and all that. So I was doing a nonprofit organization and I was not seeing clients or anything like that.
full caption & clip details
An adolescent masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, pride, triumph; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 7.1/10; 18.0s, EN.
851673_00080688 · in -24.9 dBFS · gain +4.9 dB · podcast-02388
This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with resonance: mixed (R_MIXD) at the very bottom of the range — 0.06, lower than 94 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.64.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.24, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 50 s · en · emolia
k 4d_a 0.637d_b 0.637step_a 0.241step_b 0.241min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00003_S02153track EN_B00003_S02153total 50.0slevel spread 9.5 dBmax seam 9.5 dB
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an elderly somewhat feminine voice
(longing, helplessness, sadness · slow, very low-energy, relaxed, narration)A girl. I thought so. Speak, friend. She's someone who can never be mine. I've been close to her, but I can never be as close again.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, variable; timbre is slightly warm, slightly dark, slightly rough, slightly thin; slurred, almost no disfluency, fairly narrow pitch, audible breath; affect is negative, submissive, vulnerable; reads as longing, helplessness, sadness; style: narration, storytelling; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.6/10; 13.3s, EN.
EN_B00003_S02153_W000039 · in -20.2 dBFS · gain +0.2 dB · emolia-00321
(sadness, distress, longing · slow, very low-energy, relaxed, whispered)And I really miss (low mumble) her. Spike looks thoughtful. Yes. That's a problem.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, rough, very full; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, fairly guarded; reads as sadness, distress, longing; style: whispered, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.7/10; 7.3s, EN.
EN_B00003_S02153_W000040 · in -28.3 dBFS · gain +8.3 dB · emolia-00321
(relief, shame, longing ·measured, very low-energy, neutral tension, narration)I missed a girl at Swansea Station once. I got the time wrong. I left five minutes before the train arrived. He laughed. Thanks. William sighed. Very helpful.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, variable; timbre is slightly cool, slightly dark, slightly rough, full; clear, little disfluency, wide pitch range, light breath; affect is negative, slightly dominant, vulnerable; reads as relief, shame, longing; style: narration, storytelling; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 1.5/10; 13.2s, EN.
EN_B00003_S02153_W000041 · in -18.8 dBFS · gain -1.2 dB · emolia-00321
(emotional numbness, longing, affection· measured, normally alert, slightly relaxed, narration)That night, William ate with his friends at Tony's restaurant. They were the only people there, as usual. They often ate there, because Tony was an old friend. This was his first restaurant. He was trying hard to make it a success.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as emotional numbness, longing, affection; style: narration, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.9/10; 15.7s, EN.
EN_B00003_S02153_W000042 · in -20.6 dBFS · gain +0.6 dB · emolia-00321
FULL — fullness of tone ↓voicenet__VN1__T0.60__C0.25__INTERNAL · #13
This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with fullness of tone (FULL) high — 0.87, higher than 87 % of clips in this corpus — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.75.
It takes 5 clips to get there. Clip to clip the moves are -0.11, then -0.24, then -0.20, then -0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 34 s · fr · emolia
k 5d_a -0.752d_b -0.752step_a 0.240step_b 0.240min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_Aod4zhXqAZgtrack FR_Aod4zhXqAZgtotal 34.1slevel spread 2.7 dBmax seam 1.5 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, average recording, normally alert, slightly relaxed, light breath
(relief, pride · measured, steady, some disfluency, monologue)puis, (ahem) euh, sur conseil, j'ai déterminé qu'il aurait été mieux de commencer aux origines en modifiant le projet, pour qu'il devienne une série de livres blancs sur l'histoire des révolutions, des luttes sociales et acquis sociaux depuis 1789.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, pride; style: monologue, whispered; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 4.5/10; 12.8s, FR.
FR_Aod4zhXqAZg_W000034 · in -19.6 dBFS · gain -0.3 dB · emolia-02638
(measured, fairly steady, some disfluency, monologue)(low mumble) face aux recherches et aux événements sans fin, le projet a muté en une frise chronologique approfondie qui contient à ce jour plus de 500 événements.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.4/10; 8.1s, FR.
FR_Aod4zhXqAZg_W000035 · in -18.3 dBFS · gain -1.7 dB · emolia-02638
(fast, fairly steady, some disfluency, casual)révolution, lutte, acquis sociaux en France et dans le monde depuis 1589.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.7/10; 4.2s, FR.
FR_Aod4zhXqAZg_W000036 · in -17.8 dBFS · gain -2.2 dB · emolia-02638
(fatigue exhaustion·measured, fairly steady, frequent disfluency, conversational)Pour le moment, le projet est en mode ouvert sur (low mumble) le Discord de l'After Thinking.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion; style: conversational, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.1/10; 5.4s, FR.
FR_Aod4zhXqAZg_W000037 · in -17.0 dBFS · gain -3.0 dB · emolia-02638
(normal-paced, fairly steady, some disfluency, casual)Mais visuellement, elle ressemble à un gigantesque Google Doc.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.0/10; 3.0s, FR.
FR_Aod4zhXqAZg_W000038 · in -18.5 dBFS · gain -1.5 dB · emolia-02638
This is a VoiceNet dimension, not an emotion: resonance: mask (R_MASK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with resonance: mask (R_MASK) low — 0.20, lower than 80 % of clips in this corpus — and ends with it high at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.61.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.18, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 19 s · de · emolia
k 4d_a 0.612d_b 0.612step_a 0.227step_b 0.227min_cos_consec —min_cos_anchor —dataset emolialang despeaker DE_NoGsX4Zml8Atrack DE_NoGsX4Zml8Atotal 19.2slevel spread 4.4 dBmax seam 4.4 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, normal-paced, normally alert, slightly relaxed, some disfluency
(doubt, impatience and irritability, confusion · moderately variable, average clarity, wide pitch range, casual)Ich weiß gar nicht, soll die Lampe an sein oder nicht? Was ist die Frage?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, impatience and irritability, confusion; style: casual, conversational; average recording, some background noise; genuineness 5.0/6; vocal-burst blend 0.0/10; 5.0s, DE.
DE_NoGsX4Zml8A_W000047 · in -18.3 dBFS · gain -1.7 dB · emolia-00053
(fairly steady, somewhat unclear, moderate pitch range, casual)Aber wie können wir das, wie kriegen wir das hin, dass der, (ahem) dass die Batterie jetzt lädt?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 0.8/10; 4.0s, DE.
DE_NoGsX4Zml8A_W000048 · in -14.6 dBFS · gain -5.5 dB · emolia-00053
(astonishment surprise, doubt, confusion·moderately variable, average clarity, moderate pitch range, conversational)Das ist quasi der Batterie-Lade-Onkel.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, doubt, confusion; style: conversational, casual; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 0.0/10; 4.0s, DE.
DE_NoGsX4Zml8A_W000049 · in -19.0 dBFS · gain -1.0 dB · emolia-00053
(anger, impatience and irritability, disappointment·fairly steady, average clarity, moderate pitch range, conversational)Und die Jiton, was ihr hier seht, es ist natürlich, weil (ahem) der Motor läuft, denke ich mal, das ist so ein bisschen simuliert.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, impatience and irritability, disappointment; style: conversational, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.0/10; 5.7s, DE.
DE_NoGsX4Zml8A_W000050 · in -17.2 dBFS · gain -2.8 dB · emolia-00053
This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with vocal flexibility / inflection (VFLX) low — 0.08, lower than 92 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.61.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.24, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 54 s · da · eurospeech
k 4d_a 0.606d_b 0.606step_a 0.236step_b 0.236min_cos_consec —min_cos_anchor —dataset eurospeechlang daspeaker denmark_20151M007_2015-10-track denmark_20151M007_2015-10-total 53.6slevel spread 5.5 dBmax seam 3.7 dB
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult somewhat feminine voice · balanced body, average recording
(emotional numbness · measured, subdued, slightly relaxed, monologue)Der er ikke nogen, der har bedt om ordet, så jeg siger tak til ordføreren. Så er det fru Mette
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: monologue; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.1/10; 14.6s, DA.
denmark_20151M007_2015-10-23_1000_7395840_7410473 · in -32.1 dBFS · gain +12.1 dB · eurospeech-00248
(awe, fear, intoxication altered states of consciousness· measured, very low-energy, relaxed, casual)Da SF's ordfører ikke kunne være her i dag, (low mumble) tager jeg talen, men hilser selvfølgelig (ahem) spørgsmål velkommen.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as awe, fear, intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.3/10; 11.5s, DA.
denmark_20151M007_2015-10-23_1000_7410473_7421936 · in -28.4 dBFS · gain +8.4 dB · eurospeech-00248
(confusion, disgust, distress·normal-paced, normally alert, slightly relaxed, authoritative)Det her lovforslag er en opfølgning på energiforliget fra 2012, som havde og har SF's fulde opbakning. Dengang var en af begrundelserne for en tilskudsordning den NOx-afgift, som Venstre og regeringen nu vil afskaffe.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, disgust, distress; style: authoritative, formal; average recording, some background noise; genuineness 2.0/6; vocal-burst blend 1.6/10; 13.1s, DA.
denmark_20151M007_2015-10-23_1000_7421936_7434992 · in -26.6 dBFS · gain +6.6 dB · eurospeech-00248
(pride, disgust, bitterness·brisk, normally alert, slightly relaxed, authoritative)Vi støtter lovforslaget, fordi vi mener, at en samtidig produktion af varme og el generelt er den samfundsøkonomiske og miljømæssigt bedste måde at udnytte brændsel på. SF mener også, at det er vigtigt at sikre, at tilskud ligesom afgifter har den rigtige udformning.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pride, disgust, bitterness; style: authoritative, formal; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.0/10; 14.0s, DA.
denmark_20151M007_2015-10-23_1000_7434992_7449023 · in -27.5 dBFS · gain +7.5 dB · eurospeech-00248
TEMP — tempo ↑voicenet__VN1__T0.60__C0.25__INTERNAL · #16
This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with tempo (TEMP) low — 0.16, lower than 84 % of clips in this corpus — and ends with it high at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.63.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.20, then +0.16, then +0.07 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 53 s · en · emolia
k 5d_a 0.630d_b 0.630step_a 0.200step_b 0.200min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00035_S04804track EN_B00035_S04804total 53.3slevel spread 3.5 dBmax seam 3.5 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child masculine voice · neutral-toned, fairly smooth, good recording, slightly relaxed, light breath
(awe, confusion, astonishment surprise · normal-paced, energised, moderately variable, didactic)He might say, wow, the ground's wet. The leaves of the trees are wet. The leaves and the grass are wet. It didn't rain last night. Why is it wet?
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, confusion, astonishment surprise; style: didactic, storytelling; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 2.8/10; 9.3s, EN.
EN_B00035_S04804_W000019 · in -19.8 dBFS · gain -0.2 dB · emolia-00924
(intoxication altered states of consciousness· normal-paced, energised, fairly steady, narration)Because the water in the air condensed, came together. Condense means to change from water vapor, water vapor in the air, to water drops. And these water drops fall onto the ground, on the grass, on the leaves.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness; style: narration, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.4/10; 15.2s, EN.
EN_B00035_S04804_W000020 · in -20.3 dBFS · gain +0.3 dB · emolia-00924
(concentration, interest· normal-paced, energised, fairly steady, didactic)And so even in the morning, if it didn't rain, you will find drops of water like we can see here in the picture. So condense and evaporate are opposites. Evaporate means to spread out, spread away from each other. Condense means to come together.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, interest; style: didactic, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.3/10; 16.5s, EN.
EN_B00035_S04804_W000021 · in -20.9 dBFS · gain +0.9 dB · emolia-00924
(normal-paced, normally alert, fairly steady, conversational)And especially when we're talking about water and the stages of the water cycle.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: conversational, casual; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.5/10; 5.9s, EN.
EN_B00035_S04804_W000022 · in -22.0 dBFS · gain +2.0 dB · emolia-00924
(fatigue exhaustion·brisk, normally alert, fairly steady, casual)We've talked about water vapor a little bit already. If you take a look at this picture we can see a lake.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fatigue exhaustion; style: casual, conversational; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.1/10; 5.9s, EN.
EN_B00035_S04804_W000023 · in -18.5 dBFS · gain -1.5 dB · emolia-00924
This is a VoiceNet dimension, not an emotion: style: conversational (S_CONV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with style: conversational (S_CONV) low — 0.16, lower than 84 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.61.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.24, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, quiet background, measured, subdued, fairly steady, somewhat unclear
(hope enthusiasm optimism, triumph, shame · slightly relaxed, some disfluency, moderate pitch range, monologue)2018 the FDA came out with a new five-year action plan for supporting antimicrobial stewardship in the veterinary community. And it was driven by the concept that medically important antimicrobial drugs should only be used in animals when necessary for the treatment, control, or prevention of specific diseases. And one action item in that plan was to ensure that any medically important antimicrobial drugs
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism, triumph, shame; style: monologue, casual; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 6.1/10; 28.0s, EN.
577558_00059964 · in -24.7 dBFS · gain +4.7 dB · podcast-04541
(pride, thankfulness gratitude· slightly relaxed, frequent disfluency, moderate pitch range, monologue)that continue to remain available as over-the-counter products are brought under the oversight of licensed veterinarians. So June 2021 guidance for industry number 263 was first published by the USDA, and that was basically recommendations for the manufacturers of these medically important antibiotics
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.0/10; 22.1s, EN.
577558_00062759 · in -24.9 dBFS · gain +4.9 dB · podcast-04532
(fear· slightly relaxed, frequent disfluency, fairly narrow pitch, monologue)approved for use in animals to become prescription and not over-the-counter anymore. And that's what went into effect this June, where
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as fear; style: monologue, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.0/10; 9.4s, EN.
577558_00064965 · in -25.9 dBFS · gain +5.9 dB · podcast-04526
(awe, fatigue exhaustion·relaxed, frequent disfluency, fairly narrow pitch, casual)starting in June of this year, you know, the over-the-counter antibiotics, which you know, not a whole lot of them left at this point. It was your penicillins, your tetracyclines, that kind of thing. (low mumble) Umward,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as awe, fatigue exhaustion; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 7.0/10; 16.4s, EN.
577558_00065904 · in -25.7 dBFS · gain +5.7 dB · podcast-04529
This is a VoiceNet dimension, not an emotion: valence stability (VALS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with valence stability (VALS) at the very top of the range — 0.99, higher than 99 % of clips in this corpus — and works its way down to below average at 0.29, lower than 71 % of clips in this corpus. That is a total fall of 0.70.
It takes 5 clips to get there. Clip to clip the moves are -0.23, then -0.13, then -0.12, then -0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.03 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.03, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, some disfluency
(astonishment surprise, teasing, elation · fast, energised, neutral tension, conversational)Fait qu'ils ont intérêt. Oh, à être organisé. Puis ils veulent être organisés. Fait que ça, la motivation. Ça vient de vous. Je vous (surprised gasp) aide beaucoup eux.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as astonishment surprise, teasing, elation; style: conversational, casual; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 3.8/10; 8.8s, EN.
11108_00257184 · in -22.4 dBFS · gain +2.4 dB · podcast-02372
(sourness, affection, teasing ·brisk, normally alert, neutral tension, conversational)(low mumble) C'est des outils que vous leur donnez, mais d'entrée de jeu, en secondaire, on vous leur montrait comment (ahem) tourner.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as sourness, affection, teasing; style: conversational, casual; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.4/10; 26.2s, EN.
11108_00258067 · in -23.4 dBFS · gain +3.4 dB · podcast-06421
(infatuation·normal-paced, normally alert, slightly relaxed, casual)Secondo l'organisation, entre autres dans le sport, dans la pratique.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.6/10; 8.0s, EN.
11108_00269784 · in -25.7 dBFS · gain +5.7 dB · podcast-02377
(relief, affection, thankfulness gratitude· normal-paced, normally alert, slightly relaxed, casual)It's really. I'm not the most medicine in New York is very equilibrium.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as relief, affection, thankfulness gratitude; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.4/10; 3.2s, EN.
11108_00272472 · in -28.3 dBFS · gain +8.3 dB · podcast-02380
(affection, bitterness, jealousy and envy· normal-paced, normally alert, neutral tension, conversational)passed this défis-up adversité-là. Nous, ça fait 11 ans qu'on est in operation. Nous, on donne 2500 courses par année. On est one des plus grosses écoles dans le Bas (ahem) Saint-Laurent. Mon gars is a walky. My fille is very artistic.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as affection, bitterness, jealousy and envy; style: conversational, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.7/10; 27.2s, EN.
11108_00274712 · in -21.9 dBFS · gain +1.9 dB · podcast-02376
This is a VoiceNet dimension, not an emotion: vulnerability (VULN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with vulnerability (VULN) below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the range at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.62.
It takes 5 clips to get there. Clip to clip the moves are +0.07, then +0.24, then +0.15, then +0.16 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.21 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.27 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.21, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, light breath
(jealousy and envy, impatience and irritability, triumph · normal-paced, subdued, neutral tension, casual)they're not good. Uh Swift is dominating the snaps the last two games after you know the concussion. So expect him to be involved a lot early and a lot. In the passing game, the running game. He's their best running back by far, so you know, feed him. It's the last game. We'll see what they got. We'll see what Bevel draws up for him. Expect a while, though.
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, impatience and irritability, triumph; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.4/10; 18.9s, EN.
619530_00126608 · in -21.4 dBFS · gain +1.4 dB · podcast-03053
(triumph, jealousy and envy, hope enthusiasm optimism·brisk, energised, neutral tension, casual)Last guy got running back, Josh Jacobs versus the Broncos, only sixty-two hundred bucks. It's probably the cheapest he's been all year. He split carries with Booker last week, but (ahem) uh he wasn't feeling well. I think that is why he's so cheap. Uh (ahem) he killed him earlier in the year, so uh (ahem) (ahem) right Josh Jacobs for sixty hundred
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, jealousy and envy, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 9.7/10; 15.6s, EN.
619530_00128624 · in -20.6 dBFS · gain +0.6 dB · podcast-03047
(affection, triumph, hope enthusiasm optimism ·normal-paced, normally alert, slightly relaxed, casual)week. So he could be finally healthy and motivated because that playoff shots. Yep, so yeah, all right, like that. All right,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as affection, triumph, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 4.2/10; 5.0s, EN.
619530_00131480 · in -25.7 dBFS · gain +5.7 dB · podcast-02906
(triumph, hope enthusiasm optimism, pride· normal-paced, subdued, slightly relaxed, casual)so wide receiver. Give me Calvin really, eighty five hundred bucks against Tampa Bay's (low mumble) uh Swiss cheese secondary right now. Last time he played Tampa Bay, thirty five points, ten catches, hundred sixty yards, and a touchdown. Uh (low mumble) I don't expect Julio to play this week. I don't know if we put that in the news or not. I don't think
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as triumph, hope enthusiasm optimism, pride; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 7.1/10; 18.1s, EN.
619530_00131984 · in -24.8 dBFS · gain +4.8 dB · podcast-03049
(infatuation, sexual lust, anger· normal-paced, normally alert, slightly relaxed, casual)it doesn't matter. Calvin Ridley's hot, riding the hot end. Give me Calvin Really.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as infatuation, sexual lust, anger; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.7/10; 3.9s, EN.
619530_00134328 · in -23.3 dBFS · gain +3.3 dB · podcast-02914
This is a VoiceNet dimension, not an emotion: valence stability (VALS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.60.
The chain starts with valence stability (VALS) below average — 0.31, lower than 69 % of clips in this corpus — and ends with it at the very top of the range at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.65.
It takes 5 clips to get there. Clip to clip the moves are +0.07, then +0.12, then +0.22, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.26 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.26 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · en · emolia
k 5d_a 0.648d_b 0.648step_a 0.237step_b 0.237min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_PWwSTdkIOJ4track EN_PWwSTdkIOJ4total 31.8slevel spread 3.2 dBmax seam 3.2 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, slightly relaxed, light breath
(doubt · normal-paced, normally alert, fairly steady, casual)So yes, we're, (surprised gasp) uh, what question, Risa, for the no date?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt; style: casual, formal; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.2/10; 3.8s, EN.
EN_PWwSTdkIOJ4_W000285 · in -16.3 dBFS · gain -3.7 dB · emolia-03958
(normal-paced, normally alert, fairly steady, monologue)Only in the reference do you use more than a year if it's a periodical that has a month and day. But in the in-text citation, you only use the year.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, ASMR; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.2/10; 8.5s, EN.
EN_PWwSTdkIOJ4_W000289 · in -19.4 dBFS · gain -0.6 dB · emolia-03958
(slow, normally alert, fairly steady, monologue)So for the poll, (low mumble) uh, Risa, the answer, (ahem) uh, our answer is
full caption & clip details
A young adult feminine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; no dominant emotion; style: monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.0/10; 5.5s, EN.
EN_PWwSTdkIOJ4_W000290 · in -16.2 dBFS · gain -3.8 dB · emolia-03958
(normal-paced, normally alert, fairly steady, monologue)(ahem) For the poll, the answer is B, and for the slide on the in-text citations, it's A. Yes.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.4/10; 8.0s, EN.
EN_PWwSTdkIOJ4_W000291 · in -18.3 dBFS · gain -1.7 dB · emolia-03958
(doubt· normal-paced, energised, moderately variable, dramatic)All right. Good questions. Good questions. All right. Moving on to the references.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as doubt; style: dramatic, playful; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.4/10; 5.4s, EN.
EN_PWwSTdkIOJ4_W000292 · in -17.6 dBFS · gain -2.4 dB · emolia-03958