Manifest tier. voicenet, rule VN1, T=0.7, step cap 0.25. Population 299,136 chains (4,276 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 198,019.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.voicenet__VN1__T0.70__C0.25__INTERNAL — population 299,136 chains (4,276 h). SHAREABLE variant: 198,019. Filter.rule=='VN1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.7 and step_b<=0.25 Sampled from 2,094 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with vocal flexibility / inflection (VFLX) at the very bottom of the range — 0.04, lower than 96 % of clips in this corpus — and ends with it high at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.74.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.21, then +0.21, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 25 s · en · emolia
hear it un-normalised (raw levels, max seam 2.3 dB)
k 5d_a 0.739d_b 0.739step_a 0.213step_b 0.213min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_ZUudcoN1zKotrack EN_ZUudcoN1zKototal 25.3slevel spread 3.6 dBmax seam 2.3 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · normally alert
(intoxication altered states of consciousness · normal-paced, relaxed, fairly steady, casual)No, it's, (low mumble) uh, Samuel Hayden becomes the new iPhone. I really said should be if Twitch wasn't rotted.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 2.7/10; 5.7s, EN.
EN_ZUudcoN1zKo_W000014 · in -20.9 dBFS · gain +0.9 dB · emolia-00484
(jealousy and envy, triumph, teasing· normal-paced, neutral tension, fairly steady, casual)I know the chain gun is the best weapon to deal with that (low mumble) guy, but hey, I like a good fight. So that's why I'm a little disappointed in him because that was easy as hell.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as jealousy and envy, triumph, teasing; style: casual, conversational; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.5/10; 7.5s, EN.
EN_ZUudcoN1zKo_W000015 · in -21.3 dBFS · gain +1.3 dB · emolia-00484
(intoxication altered states of consciousness, sexual lust, confusion·measured, fully relaxed, moderately variable, casual)Whiplash, whiplash, come on now, whip my ass. Oh, me.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, fully relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, thin; slurred, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, sexual lust, confusion; style: casual, conversational; below-average recording, some background noise; genuineness 3.8/6; vocal-burst blend 0.7/10; 4.9s, EN.
EN_ZUudcoN1zKo_W000016 · in -19.3 dBFS · gain -0.7 dB · emolia-00484
(astonishment surprise, impatience and irritability, confusion · measured, slightly relaxed, fairly steady, casual)What? Are you sure you just said a bad?
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise, impatience and irritability, confusion; style: casual, conversational; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.8/10; 3.1s, EN.
EN_ZUudcoN1zKo_W000017 · in -19.9 dBFS · gain -0.1 dB · emolia-00484
(astonishment surprise, embarrassment, intoxication altered states of consciousness·normal-paced, fully relaxed, moderately variable, casual)I didn't even realize. Oh, yeah. I collected the extra one. Okay.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, embarrassment, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.8/10; 3.6s, EN.
EN_ZUudcoN1zKo_W000018 · in -17.6 dBFS · gain -2.4 dB · emolia-00484
STRU — structuredness of delivery ↑voicenet__VN1__T0.70__C0.25__INTERNAL · #2
This is a VoiceNet dimension, not an emotion: structuredness of delivery (STRU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with structuredness of delivery (STRU) at the very bottom of the range — 0.06, lower than 94 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.72.
It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.15, then +0.08, then +0.24 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.45 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.53 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.45, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 44 s · en · podcast
hear it un-normalised (raw levels, max seam 1.1 dB)
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, moderately variable
(relief, astonishment surprise, embarrassment · normal-paced, normally alert, relaxed, conversational)you see, it's just like immediately you're like, oh, (ahem) I okay,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as relief, astonishment surprise, embarrassment; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.8/10; 3.6s, EN.
896074_00169488 · in -28.8 dBFS · gain +8.8 dB · podcast-02547
(sexual lust, impatience and irritability, malevolence malice·brisk, energised, slightly tense, casual)And then once you get there, it hits you like a fucking brick wall to the face.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, no disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, guarded; reads as sexual lust, impatience and irritability, malevolence malice; style: casual, ranting; good recording, no background noise; mildly explicit content; genuineness 2.6/6; vocal-burst blend 4.8/10; 4.2s, EN.
896074_00170400 · in -29.4 dBFS · gain +9.3 dB · podcast-02561
(pleasure ecstasy, relief, elation·normal-paced, normally alert, neutral tension, casual)when we got off the plane and we actually got to like go drive around, we landed super late. So when we got in, it was like two o'clock, one o'clock in the morning when we landed. (low mumble) Um, so when we were driving through the neighborhoods, it was like so cool. Because I was so excited to go to sleep so I could wake up and it'd be morning already. So
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as pleasure ecstasy, relief, elation; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 8.5/10; 18.0s, EN.
896074_00170952 · in -30.4 dBFS · gain +10.4 dB · podcast-02544
(contemplation, interest, awe·brisk, normally alert, neutral tension, casual)Yeah, and it's just (ahem) the way everything looks, the way that things are like set up, and just like the atmosphere in general, the people, everything is so different, and it's just like it's like one of those things that like you have to experience it
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as contemplation, interest, awe; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 8.9/10; 11.9s, EN.
896074_00174032 · in -29.4 dBFS · gain +9.4 dB · podcast-02542
(longing, affection, sadness·normal-paced, normally alert, neutral tension, casual)the same way that like no matter what we did explaining Italy to our family and our friends, like if you weren't there, you know,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as longing, affection, sadness; style: casual, conversational; good recording, no background noise; genuineness 4.2/6; vocal-burst blend 6.2/10; 6.2s, EN.
896074_00175368 · in -29.7 dBFS · gain +9.7 dB · podcast-02544
RANG — pitch range used ↑voicenet__VN1__T0.70__C0.25__INTERNAL · #3
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with pitch range used (RANG) low — 0.20, lower than 80 % of clips in this corpus — and ends with it at the very top of the range at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.73.
It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.16, then +0.25, then +0.24 — an uneven climb, but always in the same direction.
The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.57 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.63 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.57, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 66 s · en · podcast
hear it un-normalised (raw levels, max seam 7.4 dB)
Unchanged across all 5 clips: an adult masculine voice · neutral-toned
(anger, disgust, intoxication altered states of consciousness · measured, subdued, neutral tension, casual)(ahem) to cook. I mean, to wash chicken before you cook it. So look it up. Look it up. Like I was saying earlier, Black Friday is (low mumble) uh is a fucking scam, sham
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as anger, disgust, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 2.2/10; 15.8s, EN.
833119_00074520 · in -29.6 dBFS · gain +9.6 dB · podcast-03661
(fear, sadness, sourness·slow, normally alert, slightly relaxed, monologue)nobody's really getting a deal. I saw this TikTok where the girl was going through
full caption & clip details
A child masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as fear, sadness, sourness; style: monologue, storytelling; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 8.2s, EN.
833119_00076644 · in -28.4 dBFS · gain +8.4 dB · podcast-03661
(measured, normally alert, neutral tension, casual)and it said Black Friday deal on this on the on the sign. But if you pulled the sign back, behind it was the same price, but it was the regular price.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 0.9/10; 13.6s, EN.
833119_00077464 · in -32.6 dBFS · gain +12.6 dB · podcast-03666
(anger, contempt, impatience and irritability· measured, energised, neutral tension, casual)That's what they're doing. That's what they gotta do. I get it. That's that's their tactics. But now people are hip to the tactics. And guess what? A lot of people still don't care. This TV is $50 off. Dude, no, it's not. It's been the same all year. They just put a sticker that says $50 off.
full caption & clip details
An adult masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as anger, contempt, impatience and irritability; style: casual, dramatic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.7/10; 23.8s, EN.
833119_00080476 · in -28.4 dBFS · gain +8.4 dB · podcast-03663
(teasing, amusement, impatience and irritability ·normal-paced, highly aroused, very tense, casual)You ain't getting a discount.
full caption & clip details
An adult masculine voice; delivery is highly aroused, normal-paced, very tense, volatile; timbre is neutral-toned, dark, gravelly, thin; very slurred, some disfluency, wide pitch range, breathless; affect is deeply negative, slightly dominant, guarded; reads as teasing, amusement, impatience and irritability; style: casual, playful; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.4/10; 3.7s, EN.
833119_00083224 · in -35.8 dBFS · gain +15.8 dB · podcast-02893
RANG — pitch range used ↑voicenet__VN1__T0.70__C0.25__INTERNAL · #4
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with pitch range used (RANG) low — 0.11, lower than 89 % of clips in this corpus — and ends with it high at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.72.
It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.20, then +0.23, then +0.21 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 80 s · lt · eurospeech
hear it un-normalised (raw levels, max seam 6.2 dB)
k 5d_a 0.717d_b 0.717step_a 0.224step_b 0.224min_cos_consec —min_cos_anchor —dataset eurospeechlang ltspeaker lithuania_lithuania_2_1711track lithuania_lithuania_2_1711total 80.5slevel spread 8.4 dBmax seam 6.2 dB
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · quiet background, normally alert, somewhat unclear
(thankfulness gratitude, triumph, concentration · measured, slightly relaxed, fairly steady, monologue)už tai, kad buvo pastatytas paminklas Sausio 13-osios atminimui prie Televizijos bokšto, kad prieš tai buvo sutvarkyta teritorija. Tas paminklas buvo statomas iš (low mumble) Lietuvos radijo ir televizijos centro lėšų.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, triumph, concentration; style: monologue, casual; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.5/10; 15.8s, LT.
lithuania_lithuania_2_17112005_12553568_12569408 · in -10.4 dBFS · gain -9.6 dB · eurospeech-01928
(concentration, triumph, interest· measured, neutral tension, fairly steady, cartoonish)ką jūs planuojate ten daryti, kur tas projektas yra, kas kreipėsi. Vien tik per Vyriausybės valandą pasakyti, kad skirtume 450 tūkst. Lt, nepakanka. Yra tam tikra tvarka ir sistema. Manau, jūs turėtumėte inicijuoti šiuos dalykus.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph, interest; style: cartoonish, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.6/10; 17.1s, LT.
lithuania_lithuania_2_17112005_12584080_12601216 · in -9.6 dBFS · gain -10.4 dB · eurospeech-01928
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, triumph, pride; style: monologue, authoritative; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.3/10; 17.6s, LT.
lithuania_lithuania_2_17112005_12628192_12645840 · in -15.8 dBFS · gain -4.2 dB · eurospeech-01928
(fatigue exhaustion, pain, intoxication altered states of consciousness· normal-paced, neutral tension, moderately variable, casual)(ahem) Dėkoju, (ahem) pirmininke. (low mumble) (ahem) Gerbiamasis premjere, mes žinome, kad Lietuvoje skursta vaikai, ir šiandien mes atėmėme galimybę paremti konservatorijos
full caption & clip details
An elderly feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, pain, intoxication altered states of consciousness; style: casual, monologue; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.2/10; 15.1s, LT.
lithuania_lithuania_2_17112005_12660688_12675776 · in -18.0 dBFS · gain -2.0 dB · eurospeech-01928
(thankfulness gratitude, pride, pain · normal-paced, neutral tension, moderately variable, casual)(low mumble) (ahem) Priėmėme protokolinį nutarimą, kad iki sausio 1 dienos Vyriausybė turi parengti gabių vaikų ir jaunimo ugdymo programą ir skirti 2 mln.
full caption & clip details
An elderly feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is positive, slightly dominant, neutral openness; reads as thankfulness gratitude, pride, pain; style: casual, playful; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 7.2/10; 14.2s, LT.
lithuania_lithuania_2_17112005_12675776_12689952 · in -14.8 dBFS · gain -5.2 dB · eurospeech-01928
This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with warmth (WARM) low — 0.09, lower than 91 % of clips in this corpus — and ends with it high at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.74.
It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.21, then +0.17, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.67 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.67 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.67, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 75 s · en · podcast
hear it un-normalised (raw levels, max seam 1.0 dB)
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, moderately variable, some disfluency, average clarity
(awe, interest, astonishment surprise · brisk, energised, neutral tension, casual)also a whole new cast of characters who have no idea who they are, but all these stormtroopers look familiar. And this guy with a gun, he's
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as awe, interest, astonishment surprise; style: casual, playful; below-average recording, some background noise; mildly explicit content; genuineness 5.7/6; vocal-burst blend 6.8/10; 6.7s, EN.
469229_00166640 · in -21.2 dBFS · gain +1.2 dB · podcast-00074
(astonishment surprise, awe, elation·normal-paced, normally alert, neutral tension, casual)and it looks a lot more gritty than anything we've ever seen before. This guy with a
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, awe, elation; style: casual, conversational; poor recording, some background noise; genuineness 5.0/6; vocal-burst blend 4.6/10; 5.2s, EN.
469229_00167423 · in -21.8 dBFS · gain +1.8 dB · podcast-00098
(doubt, amusement, teasing· normal-paced, normally alert, neutral tension, casual)There's this new robot who kind of sounds like C3PO, but is 10 times taller than C3PO, and I don't think there's a guy in that suit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as doubt, amusement, teasing; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 4.4/10; 6.8s, EN.
469229_00168384 · in -21.4 dBFS · gain +1.4 dB · podcast-00078
(hope enthusiasm optimism, elation, interest·brisk, normally alert, slightly relaxed, casual)Darth Vader. So there you have it, folks. If you haven't seen it before, well, let's just put it this way. If you haven't seen any Star Wars before, before the next (low mumble) uh Key Listener show airs, I need you to watch all seven of them because the Geek End is one of the staples of the show. Kevin Lickness and I, who is one of the guest co-hosts, he and I introduced this segment. We're resident nerds, and we in turn love to talk about nerdy stuff. So the geek end ultimately is the end of the show and where we geek out.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, interest; style: casual, conversational; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.4/10; 28.9s, EN.
469229_00173908 · in -22.4 dBFS · gain +2.4 dB · podcast-06347
(hope enthusiasm optimism, elation, thankfulness gratitude·normal-paced, energised, slightly tense, casual)And (low mumble) uh yeah, it's very important that you are up to date on all things nerd. The cool thing is in 2016, it's okay to be a nerd. People embrace it actually. And um, (low mumble) you're almost cooler if you're a nerd than not. So go out there, become nerdy, and do the deed. Watch Star Wars. And and if you think something's nerdy, it probably is. And you probably should watch it, because we're probably gonna be talking about it. (ahem)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly tense, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, full; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, thankfulness gratitude; style: casual, conversational; good recording, quiet background; mildly explicit content; genuineness 2.7/6; vocal-burst blend 6.0/10; 26.6s, EN.
469229_00176792 · in -22.5 dBFS · gain +2.5 dB · podcast-06354
This is a VoiceNet dimension, not an emotion: clarity / intelligibility (CLRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with clarity / intelligibility (CLRT) low — 0.08, lower than 92 % of clips in this corpus — and ends with it high at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.74.
It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.23, then +0.12 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 58 s · zh · emolia
k 5d_a 0.741d_b 0.741step_a 0.232step_b 0.232min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00040_S01743track ZH_B00040_S01743total 58.1slevel spread 3.6 dBmax seam 2.1 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · average recording, audible breath
(bitterness, contemplation, malevolence malice · measured, very low-energy, relaxed, storytelling)言下之意,就是我把这点价值毁掉了,辜负了你我现在呀已经不仅仅是尴尬了,而是有了犯罪感,也不用等以后的地狱。我现在已经在地狱了。萧亚文说,你先好好听着,我还没有说到地狱。
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is warm, slightly dark, rough, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as bitterness, contemplation, malevolence malice; style: storytelling, narration; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 4.4/10; 28.8s, ZH.
ZH_B00040_S01743_W000051 · in -20.0 dBFS · gain -0.0 dB · emolia-03675
This is a VoiceNet dimension, not an emotion: style: cartoonish (S_CART) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with style: cartoonish (S_CART) low — 0.10, lower than 90 % of clips in this corpus — and ends with it high at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.73.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.16, then +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 27 s · en · emolia
k 5d_a 0.727d_b 0.727step_a 0.225step_b 0.225min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_KUCNHX5WKYgtrack EN_KUCNHX5WKYgtotal 26.8slevel spread 3.2 dBmax seam 1.1 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult somewhat masculine voice · neutral tension, moderately variable, some disfluency, average clarity
(jealousy and envy, impatience and irritability, fatigue exhaustion · normal-paced, energised, wide pitch range, casual)I'm always plotting on quitting gaming, but at least I don't just do it right away now.
full caption & clip details
A young adult somewhat masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as jealousy and envy, impatience and irritability, fatigue exhaustion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.6/6; vocal-burst blend 3.3/10; 5.1s, EN.
EN_KUCNHX5WKYg_W000108 · in -17.3 dBFS · gain -2.7 dB · emolia-00786
(sourness, jealousy and envy, contempt· normal-paced, normally alert, moderate pitch range, casual)Are these false phaser cannons? I mean the torpedo turrets seem like they're better, so... Why bother with the crappy ones, right?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, jealousy and envy, contempt; style: casual, storytelling; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.1/10; 7.1s, EN.
EN_KUCNHX5WKYg_W000110 · in -17.5 dBFS · gain -2.5 dB · emolia-00786
(fatigue exhaustion, impatience and irritability, helplessness· normal-paced, normally alert, wide pitch range, casual)We probably need a second shipyard. I'm gonna run out of dilithium way too fast.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fatigue exhaustion, impatience and irritability, helplessness; style: casual, storytelling; average recording, some background noise; mildly explicit content; genuineness 3.1/6; vocal-burst blend 1.1/10; 4.8s, EN.
EN_KUCNHX5WKYg_W000112 · in -16.5 dBFS · gain -3.5 dB · emolia-00786
(awe, pride, elation· normal-paced, energised, wide pitch range, casual)I got one little starship out here. I have one little starship.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as awe, pride, elation; style: casual, storytelling; average recording, some background noise; genuineness 3.1/6; vocal-burst blend 0.2/10; 4.9s, EN.
EN_KUCNHX5WKYg_W000113 · in -15.4 dBFS · gain -4.6 dB · emolia-00786
(impatience and irritability, anger, sourness·brisk, energised, wide pitch range, casual)It's difficult to find decent channels to raid. They have, it is, and (ahem) I'm-
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as impatience and irritability, anger, sourness; style: casual, playful; below-average recording, some background noise; genuineness 3.5/6; vocal-burst blend 1.9/10; 4.3s, EN.
EN_KUCNHX5WKYg_W000115 · in -14.3 dBFS · gain -5.7 dB · emolia-00786
This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with resonance: throat (R_THRT) high — 0.81, higher than 81 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.04, lower than 96 % of clips in this corpus. That is a total fall of 0.77.
It takes 5 clips to get there. Clip to clip the moves are -0.19, then -0.18, then -0.22, then -0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.79 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.79 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 66 s · en · emolia
k 5d_a -0.768d_b -0.768step_a 0.217step_b 0.217min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00081_S09184track EN_B00081_S09184total 65.9slevel spread 3.1 dBmax seam 3.1 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, some disfluency, light breath
(astonishment surprise, amusement, teasing · normal-paced, normally alert, slightly relaxed, casual)Isn't actually that weird. John, I'll see you on Tuesday.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, amusement, teasing; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.5/10; 3.3s, EN.
EN_B00081_S09184_W000006 · in -16.8 dBFS · gain -3.2 dB · emolia-01795
(sourness, awe, astonishment surprise · normal-paced, energised, slightly relaxed, casual)This shouldn't have happened. The telescope that spotted Oumuamua does a whole sky survey and is not designed to catch things like this. In fact, it didn't spot the object until it was well past its peak brightness.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, awe, astonishment surprise; style: casual, dramatic; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.0/10; 11.3s, EN.
EN_B00081_S09184_W000007 · in -19.2 dBFS · gain -0.8 dB · emolia-01795
(doubt, interest, triumph· normal-paced, normally alert, slightly relaxed, casual)And that will be a big deal. But maybe the best argument against the alien hypothesis is simply that the universe is weird. Many times we've seen things in the night sky and had no way to explain it and thought, oh, this is it.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as doubt, interest, triumph; style: casual, conversational; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.7/10; 13.5s, EN.
EN_B00081_S09184_W000008 · in -16.7 dBFS · gain -3.3 dB · emolia-01795
(hope enthusiasm optimism, pleasure ecstasy, elation·brisk, energised, neutral tension, authoritative)To make this happen, (ahem) uh, to, to start observing once we realize that we had something weird on the radar. Educational videos are exempt from the four-minute rule, obviously. Just saying it. And the projects were awesome next week. Next week! Next weekend, you guys! So if you're thinking about making a video promoting your charity of choice, please
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, pleasure ecstasy, elation; style: authoritative, conversational; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.5/10; 18.1s, EN.
EN_B00081_S09184_W000009 · in -19.8 dBFS · gain -0.2 dB · emolia-01795
(interest, awe, astonishment surprise· brisk, energised, neutral tension, casual)See, weirdly enough, photons do not have mass, but they do have momentum, and when they hit something, they give it a tiny push. A big sheet of something very reflective but very lightweight can take advantage of that fact and use it for propulsion, something we humans have actually experimented with. But look, that can't be right, right?
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as interest, awe, astonishment surprise; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 1.8/6; vocal-burst blend 4.3/10; 19.1s, EN.
EN_B00081_S09184_W000010 · in -18.6 dBFS · gain -1.4 dB · emolia-01795
BRGT — brightness of timbre ↑voicenet__VN1__T0.70__C0.25__INTERNAL · #9
This is a VoiceNet dimension, not an emotion: brightness of timbre (BRGT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with brightness of timbre (BRGT) at the very bottom of the range — 0.01, lower than 99 % of clips in this corpus — and ends with it above average at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.72.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.08, then +0.24, then +0.18 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.57 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert
(pleasure ecstasy, relief, contentment · normal-paced, slightly relaxed, fairly steady, casual)belle journée, une très belle fin de semaine. C'est fantastique. Et puis, (low mumble) semble-t-il que son français est très bon maintenant.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as pleasure ecstasy, relief, contentment; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.3/10; 6.2s, FR.
32418_00067256 · in -30.8 dBFS · gain +10.8 dB · podcast-00899
(distress, helplessness· normal-paced, slightly relaxed, moderately variable, casual)Laval en français. (ahem) Donc, ça justifie encore plus la qualité de son français. Valérie, écoute, (ahem) 34 ans, ça fait longtemps que tu patines, tu es de l'AB. (ahem) Tu as vécu à Chicoutimi, je ne me trompe pas, c'est ça?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as distress, helplessness; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 6.2/10; 12.2s, FR.
32418_00068800 · in -27.5 dBFS · gain +7.5 dB · podcast-00902
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as triumph; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 4.2/10; 5.5s, FR.
32418_00070383 · in -23.8 dBFS · gain +3.8 dB · podcast-00904
(fatigue exhaustion, sourness, impatience and irritability·measured, neutral tension, moderately variable, conversational)du monde, mais je pense que tu en as plus que ça, même si on compte le courte épiste et le longue piste. (low mumble) Carrière
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, sourness, impatience and irritability; style: conversational, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 0.0/10; 6.2s, FR.
32418_00070935 · in -30.1 dBFS · gain +10.1 dB · podcast-00913
(impatience and irritability, astonishment surprise, pride·normal-paced, neutral tension, moderately variable, casual)Peu importe, là. Écoute, je posais la question à Laurent, mais est-ce que. (ahem) Il n'y a pas de statistiques, j'ai pas ressorti ça, mais j'ai compté 41 médailles en Coupe du Monde courte piste. Ça, c'est déjà quelque chose d'assez exceptionnel. Pis là, tu en as ajouté en longue piste aussi. C'est
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, astonishment surprise, pride; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 6.9/10; 14.2s, FR.
32418_00072935 · in -23.4 dBFS · gain +3.4 dB · podcast-04863
This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with vocal flexibility / inflection (VFLX) high — 0.84, higher than 84 % of clips in this corpus — and works its way down to low at 0.10, lower than 90 % of clips in this corpus. That is a total fall of 0.74.
It takes 5 clips to get there. Clip to clip the moves are -0.24, then -0.18, then -0.18, then -0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 27 s · zh · emolia
k 5d_a -0.737d_b -0.737step_a 0.243step_b 0.243min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00033_S07811track ZH_B00033_S07811total 27.4slevel spread 0.8 dBmax seam 0.6 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, clear
(shame, fear · normal-paced, fairly steady, moderate pitch range, formal)江一难听出楚天南话里的意思。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, fear; style: formal, monologue; very good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.7/10; 3.1s, ZH.
ZH_B00033_S07811_W000008 · in -16.2 dBFS · gain -3.8 dB · emolia-03603
This is a VoiceNet dimension, not an emotion: aesthetic pleasantness (ESTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with aesthetic pleasantness (ESTH) above average — 0.72, higher than 72 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.01, lower than 99 % of clips in this corpus. That is a total fall of 0.72.
It takes 5 clips to get there. Clip to clip the moves are -0.19, then -0.09, then -0.21, then -0.23 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.37 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.32 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.37, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly rough
(sourness, contempt, teasing · normal-paced, energised, neutral tension, casual)It's seriously the best Canadian movie you ever watch in your life.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as sourness, contempt, teasing; style: casual, conversational; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.2/10; 5.2s, EN.
121580_00014268 · in -28.1 dBFS · gain +8.2 dB · podcast-02849
(contempt, malevolence malice, sourness · normal-paced, normally alert, neutral tension, casual)You know what they do? They go straight to that source. They go straight to the brewing factory, and all kinds of chaos ensues.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contempt, malevolence malice, sourness; style: casual, monologue; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.1/10; 7.1s, EN.
121580_00017408 · in -30.0 dBFS · gain +10.0 dB · podcast-02570
(fatigue exhaustion· normal-paced, normally alert, fully relaxed, casual)I'm looking at the movie right now, 83.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as fatigue exhaustion; style: casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.8/10; 3.0s, EN.
121580_00018224 · in -28.6 dBFS · gain +8.6 dB · podcast-02564
(impatience and irritability, jealousy and envy, anger·measured, subdued, neutral tension, casual)It's not that I don't like it. It's just that to me, the creepiest of all decades has to be the 80s. When y'all have the 80 cent orchestration for movies, I'm sorry, but Scarface (low mumble) to me, the only thing that made Scarface at all scary was the soundtrack.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, jealousy and envy, anger; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.8/10; 22.6s, EN.
121580_00021040 · in -26.6 dBFS · gain +6.6 dB · podcast-02577
(teasing, amusement, intoxication altered states of consciousness·slow, very low-energy, relaxed, casual)Man, you probably just (ahem) offended half the audience right here. (low mumble)
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, very dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, intoxication altered states of consciousness; style: casual, conversational; below-average recording, quiet background; genuineness 5.7/6; vocal-burst blend 0.0/10; 14.4s, EN.
121580_00023592 · in -29.1 dBFS · gain +9.1 dB · podcast-02847
This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with style: authoritative (S_AUTH) low — 0.22, lower than 78 % of clips in this corpus — and ends with it at the very top of the range at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.77.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.20, then +0.13, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 77 s · da · eurospeech
k 5d_a 0.767d_b 0.767step_a 0.234step_b 0.234min_cos_consec —min_cos_anchor —dataset eurospeechlang daspeaker denmark_20181M030_2018-12-track denmark_20181M030_2018-12-total 77.0slevel spread 2.5 dBmax seam 2.5 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly rough, balanced body, quiet background
(pride · measured, very low-energy, relaxed, monologue)(low mumble) Det var i de her (low mumble) Mig og Charlie-tider, hvor der blev kørt meget på knallert. (low mumble) Der var desværre også mange uheld. Det er rigtigt, at den undersøgelse, som Motorcykel Importør Foreningen lægger frem,
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; reads as pride; style: monologue, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.4/10; 12.2s, DA.
denmark_20181M030_2018-12-04_1300_6273376_6285552 · in -24.3 dBFS · gain +4.3 dB · eurospeech-00344
(concentration, anger, pride · measured, subdued, slightly relaxed, monologue)viser (ahem) 0 (exhausted groan) dødsulykker (low mumble) i et par år. Så er der altså også et år, hvor der har været et par stykker. Og når det handler om sikkerhed, skal vi altså være på den sikre side, om man så må sige. Vi har ikke lyst til at udfordre (ahem) skæbnen på det område, og derfor er vi lidt tilbageholdende. Så vi kigger på undersøgelsen.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, anger, pride; style: monologue, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.8/10; 18.5s, DA.
denmark_20181M030_2018-12-04_1300_6285552_6304032 · in -22.6 dBFS · gain +2.6 dB · eurospeech-00344
(impatience and irritability, contempt, fatigue exhaustion· measured, normally alert, slightly relaxed, monologue)Tak til hr. Rasmus Prehn. Der er ikke flere korte bemærkninger. Vi går videre til hr. Kristian Pihl Lorentzen, Venstre. Kl. 14:46 (Ordfører)
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, contempt, fatigue exhaustion; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.1/10; 13.0s, DA.
denmark_20181M030_2018-12-04_1300_6304032_6317023 · in -24.3 dBFS · gain +4.3 dB · eurospeech-00344
(concentration, anger· measured, normally alert, neutral tension, monologue)Den 19. maj 1976 var en stor dag i mit liv, en dag, jeg mindes med stor glæde, for da havde jeg slagtet min sparegris og bevæbnet med 5.000 kr. i kontanter tog jeg ud til cykelhandler Josefsen i Voldby ved Hammel,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, anger; style: monologue, cartoonish; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.5/10; 17.2s, DA.
denmark_20181M030_2018-12-04_1300_6317023_6334175 · in -21.8 dBFS · gain +1.8 dB · eurospeech-00344
(impatience and irritability, anger, interest·normal-paced, energised, neutral tension, cartoonish)og så købte jeg kontant en splinterny Puck Grand Prix knallert, og så kørte jeg stolt hjem på den. Det var i sandhed en stor dag. Og det var det jo, fordi det var et utroligt fremskridt, at man som sådan en knægt, der kom fra en landsby langt, langt ude på landet, væk fra storbyerne, hvor der ikke var ret mange busser at køre med,
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as impatience and irritability, anger, interest; style: cartoonish, authoritative; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.5/10; 15.6s, DA.
denmark_20181M030_2018-12-04_1300_6334175_6349744 · in -22.2 dBFS · gain +2.2 dB · eurospeech-00344
This is a VoiceNet dimension, not an emotion: aesthetic pleasantness (ESTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with aesthetic pleasantness (ESTH) at the very top of the range — 0.91, higher than 91 % of clips in this corpus — and works its way down to low at 0.20, lower than 80 % of clips in this corpus. That is a total fall of 0.71.
It takes 5 clips to get there. Clip to clip the moves are -0.22, then -0.24, then -0.20, then -0.06 — a plateau around step 4, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.09 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · quiet background
(astonishment surprise · normal-paced, normally alert, neutral tension, casual)how does it actually and let's be real? There are a lot of CEOs and a lot of executives that get really nervous about
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 5.3/10; 5.4s, EN.
906807_00211176 · in -26.3 dBFS · gain +6.3 dB · podcast-02623
(pleasure ecstasy·brisk, normally alert, neutral tension, casual)employees of a personal brand. And I don't want to give give away everything yet, but that is one of the key pieces we're gonna talk about. So for the ICs and the managers that want to do it, and then also for the executives, like you're gonna want to hear this because I think that is one of the
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 10.0/10; 12.5s, EN.
906807_00211728 · in -29.6 dBFS · gain +9.6 dB · podcast-01214
(contemplation, doubt, interest·normal-paced, subdued, neutral tension, casual)things that Nate and I have done a great job on, and you and I talked about this a lot at lunch a couple weeks ago. Um You (low mumble) know, how how do we deal with it when we come in the room and maybe I get more attention than he does, or vice versa, when I'm used to eat all the attention? Like how how do we deal with that? How do two big personalities and we'll call it what it is? How do two big egos coexist in a way that is beneficial to all parties involved? I
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contemplation, doubt, interest; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 10.0/10; 21.8s, EN.
906807_00212976 · in -29.9 dBFS · gain +9.9 dB · podcast-01208
(amusement, contempt, impatience and irritability·brisk, normally alert, neutral tension, casual)appreciate that you checked your own words on that. You're like, no, not just personalities, egos. It is what it is.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as amusement, contempt, impatience and irritability; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 3.0/10; 5.4s, EN.
906807_00215160 · in -21.6 dBFS · gain +1.6 dB · podcast-02616
(contemplation·measured, very low-energy, relaxed, casual)Because you you don't build a brand like that, like at all, without some level of ego. People tell me all they want, like, oh no, no, it's all about no, like
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, very dark, rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contemplation; style: casual, conversational; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.8/10; 10.2s, EN.
906807_00215840 · in -30.4 dBFS · gain +10.4 dB · podcast-01207
This is a VoiceNet dimension, not an emotion: audible breath / respiration (RESP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with audible breath / respiration (RESP) at the very top of the range — 0.96, higher than 96 % of clips in this corpus — and works its way down to low at 0.21, lower than 79 % of clips in this corpus. That is a total fall of 0.76.
It takes 5 clips to get there. Clip to clip the moves are -0.21, then -0.19, then -0.17, then -0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.14 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · quiet background, average clarity
(fatigue exhaustion, pain, helplessness · normal-paced, normally alert, neutral tension, casual)split or meant went through moments of trauma and that we have to meet and resolve, like almost like parts work. My
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, pain, helplessness; style: casual, whispered; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.7/10; 7.6s, EN.
679493_00187348 · in -18.5 dBFS · gain -1.5 dB · podcast-03183
(concentration, contemplation, interest·measured, very low-energy, slightly relaxed, casual)these points of fracture can become for lack of better description, like a little bit more shamanic, where there is an enmeshment of other energies at play which are not necessarily of the self, like maybe programs or just other energies at play. Is that something that you ever encounter in your in your work or in your personal healing process?
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration, contemplation, interest; style: casual, ASMR; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.2/10; 28.4s, EN.
679493_00189176 · in -17.2 dBFS · gain -2.8 dB · podcast-05364
(normal-paced, energised, neutral tension, casual)that divine intervention that you need so that you can see more clearly. And sometimes that is even enough pass away, whatever influence has been at work through you.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.3/10; 14.5s, EN.
679493_00209467 · in -12.7 dBFS · gain -7.3 dB · podcast-03179
(hope enthusiasm optimism, interest, teasing· normal-paced, normally alert, neutral tension, casual)Okay, you need to change like yesterday. So what I often like to tell people, it's like if you really feel like you are defaced, like you sense that there is more than what is just yours there. I'm inviting you to
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, interest, teasing; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 7.9/10; 18.0s, EN.
679493_00211967 · in -12.8 dBFS · gain -7.2 dB · podcast-03208
(teasing, malevolence malice, anger· normal-paced, energised, neutral tension, storytelling)(low mumble) um let me put it this way. Sit your ass down, connect with your highest guidance and just just utter the words. I feel like there is m way more than just my own self at play here. And I am tired and I'm no longer willing, I'm no longer giving permission
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as teasing, malevolence malice, anger; style: storytelling, playful; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.2/10; 15.2s, EN.
679493_00213764 · in -14.0 dBFS · gain -6.0 dB · podcast-03177
R_HEAD — resonance: head ↑voicenet__VN1__T0.70__C0.25__INTERNAL · #15
This is a VoiceNet dimension, not an emotion: resonance: head (R_HEAD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with resonance: head (R_HEAD) low — 0.19, lower than 81 % of clips in this corpus — and ends with it at the very top of the range at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.78.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.15, then +0.19, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 46 s · en · emolia
k 5d_a 0.780d_b 0.780step_a 0.241step_b 0.241min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_WNDEIS_PjTutrack EN_WNDEIS_PjTutotal 46.2slevel spread 3.7 dBmax seam 3.7 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, quiet background, normally alert
(relief, hope enthusiasm optimism, pleasure ecstasy · normal-paced, slightly relaxed, fairly steady, casual)Yeah, I fully agree. It's also about getting buy-in from the team.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as relief, hope enthusiasm optimism, pleasure ecstasy; style: casual, conversational; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.8/10; 3.8s, EN.
EN_WNDEIS_PjTu_W000076 · in -20.9 dBFS · gain +0.9 dB · emolia-01504
(shame, elation· normal-paced, neutral tension, fairly steady, conversational)Because you could have done all the personas by yourself and presented them to the team and then say, this is what you have to use them now on.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, elation; style: conversational, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.5/10; 7.5s, EN.
EN_WNDEIS_PjTu_W000077 · in -24.6 dBFS · gain +4.6 dB · emolia-01504
(contemplation· normal-paced, neutral tension, fairly steady, casual)They would maybe, maybe they would embrace it. Maybe they would look at it and think, (ahem) oh, this is cool, but not really use it. But once you had them go through the data, help like do the personas themselves, (low mumble) um, they.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.7/10; 14.0s, EN.
EN_WNDEIS_PjTu_W000078 · in -24.6 dBFS · gain +4.6 dB · emolia-01504
(sourness, affection, embarrassment· normal-paced, slightly relaxed, fairly steady, casual)Can be a bit difficult for the UX person, especially if you're alone and you're always facilitating and getting people to get to the conclusions by themselves, because you do a lot of hard work.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sourness, affection, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.2/10; 10.6s, EN.
EN_WNDEIS_PjTu_W000080 · in -24.3 dBFS · gain +4.3 dB · emolia-01504
(embarrassment, pleasure ecstasy, relief·brisk, neutral tension, moderately variable, conversational)When you're facilitating, say a workshop to create personas and then the designers walk off and say, oh, I made these cool things and now I use them in my work. And you can feel a bit overridden.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as embarrassment, pleasure ecstasy, relief; style: conversational, playful; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 5.6/10; 9.6s, EN.
EN_WNDEIS_PjTu_W000081 · in -24.6 dBFS · gain +4.5 dB · emolia-01504
This is a VoiceNet dimension, not an emotion: vulnerability (VULN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with vulnerability (VULN) high — 0.86, higher than 86 % of clips in this corpus — and works its way down to low at 0.13, lower than 87 % of clips in this corpus. That is a total fall of 0.73.
It takes 5 clips to get there. Clip to clip the moves are -0.13, then -0.16, then -0.20, then -0.25 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.29 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.29 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.29, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, average clarity, light breath
(pleasure ecstasy, elation, affection · normal-paced, neutral tension, moderately variable, conversational)All the weekends that I was trying to do, I didn't know, because I'm going to find my emotions. So it's a course of the mind.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pleasure ecstasy, elation, affection; style: conversational, casual; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 2.9/10; 7.3s, EN.
54499_00233328 · in -27.5 dBFS · gain +7.5 dB · podcast-01088
(astonishment surprise, relief, teasing· normal-paced, neutral tension, moderately variable, casual)I'm not paying. I'm shocked you're talking about the morning I'm going to do it a bit long time. After I went to work and we have the (ahem) F the (ahem) chamber and the matter with me, I've already had the course, I just had marked, I had a nice one and I'm shocked. I've got the photo.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as astonishment surprise, relief, teasing; style: casual, conversational; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 5.8/10; 20.6s, EN.
54499_00235952 · in -27.3 dBFS · gain +7.3 dB · podcast-01095
(disappointment, teasing, shame·fast, neutral tension, moderately variable, casual)alimentally that's a danger. Putain, I'm not in that, the guy is trying to get that.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, teasing, shame; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.1/10; 9.1s, EN.
54499_00238576 · in -30.8 dBFS · gain +10.8 dB · podcast-01077
(astonishment surprise, contentment, pleasure ecstasy·normal-paced, relaxed, moderately variable, conversational)It's a chance. And we (low mumble) are at the (ahem) time. But you're trying to do a moment to equilibrate to have all these passions, and just you see the passion and you have that.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, contentment, pleasure ecstasy; style: conversational, casual; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.8/10; 17.5s, EN.
54499_00240536 · in -28.6 dBFS · gain +8.7 dB · podcast-01077
(pride, bitterness, disappointment· normal-paced, neutral tension, fairly steady, casual)I'm not a great way to my passion. No, yeah, but then more. I said to get this line of I said to read my students a year, I'll pass my permission, it's a promise to all the time.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, bitterness, disappointment; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 5.5/10; 13.0s, EN.
54499_00242608 · in -31.1 dBFS · gain +11.2 dB · podcast-01117
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with arousal / activation (AROU) at the very top of the range — 0.95, higher than 95 % of clips in this corpus — and works its way down to low at 0.20, lower than 80 % of clips in this corpus. That is a total fall of 0.76.
It takes 5 clips to get there. Clip to clip the moves are -0.25, then -0.16, then -0.11, then -0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.38 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.47 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.38, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, quiet background
(sexual lust, teasing, impatience and irritability · normal-paced, normally alert, neutral tension, casual)On the bike? He's like, Do you know what you're doing? I'm like, Yeah, I'm just trying to find this location. He's like, Well, you almost hit me. So he gave me a ticket.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, teasing, impatience and irritability; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 7.0/10; 5.6s, EN.
666507_00183600 · in -18.9 dBFS · gain -1.1 dB · podcast-04513
(embarrassment, infatuation, shame· normal-paced, normally alert, neutral tension, casual)Oh, that's it's understandable because I almost fucking hit her. Damn. No, so with mine, I was (low mumble) um The process of moving here actually, and I remember I load up all my shit, and I had I was towing my boat at the time, and (low mumble) um this boat was (low mumble) uh I called it root beer float because it was like
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as embarrassment, infatuation, shame; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 9.5/10; 18.8s, EN.
666507_00184184 · in -23.5 dBFS · gain +3.5 dB · podcast-04514
(embarrassment, astonishment surprise, intoxication altered states of consciousness· normal-paced, normally alert, fully relaxed, casual)oh god, this boat. I remember I bought this boat for I think Oh, (contented sigh) it
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment, astonishment surprise, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 0.2/10; 4.3s, EN.
666507_00186640 · in -22.6 dBFS · gain +2.6 dB · podcast-04509
(intoxication altered states of consciousness, embarrassment · normal-paced, normally alert, relaxed, casual)I bought this boat, it was like three grand, and (low mumble) uh came with all this shit. But anyway, I was towing it and it's like an old nineteen eighties, I think it
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 9.2/10; 9.4s, EN.
666507_00187616 · in -27.2 dBFS · gain +7.2 dB · podcast-04517
(astonishment surprise, confusion, contemplation·measured, very low-energy, relaxed, casual)boat, and I was towing it, and out of nowhere, like you know where like Rome is between like Jordan, like kind of after Jordan Valley area.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, very dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, confusion, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 8.5/10; 9.4s, EN.
666507_00188664 · in -27.1 dBFS · gain +7.1 dB · podcast-04509
This is a VoiceNet dimension, not an emotion: emphasis / stress strength (EMPH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with emphasis / stress strength (EMPH) below average — 0.26, lower than 74 % of clips in this corpus — and ends with it at the very top of the range at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.70.
It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.15, then +0.19, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.21 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.21 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.21, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a child somewhat masculine voice · slightly rough, quiet background, normally alert
(infatuation · normal-paced, neutral tension, moderately variable, casual)E essas dimensões, por sua vez são subdivididas em elementos. E ao todo temos 17 elementos que constituem o nosso sistema de gestão.
full caption & clip details
A child somewhat masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is cool, dark, slightly rough, thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: casual, storytelling; below-average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.2/10; 10.0s, PT.
863300_00007352 · in -24.0 dBFS · gain +4.0 dB · podcast-00475
(thankfulness gratitude· normal-paced, neutral tension, fairly steady, casual)Hoje vamos dar início a uma nova dimensão, a dimensão técnico. Hoje vamos falar do quarto elemento do VPS. Você sabe qual é esse quarto elemento, Lud?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: casual, conversational; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 4.9/10; 10.3s, PT.
863300_00008392 · in -20.1 dBFS · gain +0.1 dB · podcast-00865
(contentment, elation, infatuation·measured, neutral tension, moderately variable, casual)Sim, claro! O quarto elemento do VPS é o primeiro da dimensão técnico, que é percepção e gerenciamento de riscos. E para falar sobre ele, nós temos um convidado muito especial, Marcelo Augusto Fasa, gerente geral da Eletrovia DFC.
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, elation, infatuation; style: casual, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.1/10; 19.0s, PT.
863300_00009504 · in -23.9 dBFS · gain +4.0 dB · podcast-00480
(pleasure ecstasy, affection, contentment · measured, neutral tension, fairly steady, casual)Seja muito bem-vindo, Marcelo, obrigado por participar aqui conosco e gostaria que você começasse se apresentando para os nossos ouvintes.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pleasure ecstasy, affection, contentment; style: casual, monologue; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.6/10; 22.8s, PT.
863300_00011432 · in -24.7 dBFS · gain +4.7 dB · podcast-00865
(pride, relief, malevolence malice·brisk, slightly relaxed, fairly steady, authoritative)(low mumble) Muito bacana, Marcelo. Obrigado por sua resposta. E assim como os outros elementos do VPS, esse quarto elemento é subdividido em subelementos. Você poderia nos dar um overview de quantos são e o que eles abordam?
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, relief, malevolence malice; style: authoritative, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 7.1/10; 18.4s, PT.
863300_00023582 · in -17.6 dBFS · gain -2.4 dB · podcast-00863
This is a VoiceNet dimension, not an emotion: style: dramatic (S_DRAM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with style: dramatic (S_DRAM) low — 0.15, lower than 85 % of clips in this corpus — and ends with it at the very top of the range at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.78.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.20, then +0.21, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 78 s · hr · eurospeech
k 5d_a 0.780d_b 0.780step_a 0.233step_b 0.233min_cos_consec —min_cos_anchor —dataset eurospeechlang hrspeaker croatia_20170712152203-260track croatia_20170712152203-260total 78.1slevel spread 3.5 dBmax seam 1.4 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · slightly dark, quiet background, measured, fairly steady
(pride, triumph, thankfulness gratitude · subdued, slightly relaxed, some disfluency, monologue)Ljude će to zanimat. Ko tamo do vjerovnika ima pravo prosuđivati koji će dobit novac, i da li su oni koji su u Vjerovničkom vijeću ispred malih dioničara privilegirani, dok neki drugi čekaju mjesecima a čekat će vjerojatno još.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, triumph, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 2.3/10; 14.9s, HR.
croatia_20170712152203-26013_16930239_16945184 · in -32.9 dBFS · gain +12.9 dB · eurospeech-01477
(doubt·normally alert, slightly relaxed, some disfluency, monologue)Tako da za vas bi bilo u interesu da se ovo što prije rasvijetli za nas. Što dalje odmiče vrijeme, vjerujte bit će vam teže gledati to šta ne date ljudima informaciju. **Reiner, Željko (HDZ)** Hvala. Slijedeća replika uvaženog zastupnika Nikole Grmoje.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: monologue, storytelling; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.2/10; 16.7s, HR.
croatia_20170712152203-26013_16945184_16961904 · in -32.5 dBFS · gain +12.6 dB · eurospeech-01477
(concentration, thankfulness gratitude, contentment·subdued, slightly relaxed, frequent disfluency, monologue)Poštovani potpredsjedniče. Kolegice nije bitno sada koliko je zastupnika tu, bitno je da mi ovdje iznosimo argumente koji su neoborivi. (low mumble) Cijelo vrijeme je priča da je nekakva intencija SDP-a rušenje Vlade.
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, thankfulness gratitude, contentment; style: monologue, cartoonish; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.3/10; 19.5s, HR.
croatia_20170712152203-26013_16961904_16981424 · in -32.1 dBFS · gain +12.1 dB · eurospeech-01477
(awe, thankfulness gratitude, contentment ·normally alert, slightly relaxed, frequent disfluency, authoritative)Ova Vlada neće pasti. Ova Vlada neće pasti, ova Vlada je kupljena. To su ljudi koji su kupljeni, koji dolaze tu po potrebi glasovati.
croatia_20170712152203-26013_16981424_16994192 · in -30.6 dBFS · gain +10.7 dB · eurospeech-01477
(shame, pride, thankfulness gratitude · normally alert, neutral tension, almost no disfluency, storytelling)Ovo je iako tanka većina, ovo je najčvršća moguća većina i ovdje se ne radi o pokušaju rušenja Vlade. Vlada neće pasti. Ali ono što mi želimo
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is warm, slightly dark, rough, booming; average clarity, almost no disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as shame, pride, thankfulness gratitude; style: storytelling, cartoonish; below-average recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.7/10; 13.6s, HR.
croatia_20170712152203-26013_16994192_17007744 · in -29.3 dBFS · gain +9.3 dB · eurospeech-01477
This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.70.
The chain starts with resonance: nasal (R_NASL) low — 0.19, lower than 81 % of clips in this corpus — and ends with it at the very top of the range at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.71.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.24, then +0.18, then +0.06 — a plateau around step 4, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.00 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.17 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.00, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth, quiet background, moderately variable, some disfluency, average clarity
(contemplation, interest, awe · normal-paced, normally alert, slightly relaxed, conversational)me there. But speaking of lower levels, you know, I didn't even necessarily think we would go in this direction, but it's just inspiring me to ask you this question. We are seeing so much level two now. All over the place. I mean, it's just metastasizing and especially in social media with a click of a button where that stuff can spread everywhere. And I guess I would love a little perspective on that, on the idea of
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, interest, awe; style: conversational, casual; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 8.9/10; 26.1s, EN.
152573_00228121 · in -20.9 dBFS · gain +0.9 dB · podcast-06357
(concentration, interest, fear· normal-paced, energised, neutral tension, conversational)level two. We're all we're all in it. We're seeing it. We're sort of being poked and prodded. What can we do maybe to tr try to shift that energy a little bit in the social media space or in the political social space?
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, interest, fear; style: conversational, casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 5.1/10; 26.6s, EN.
152573_00230729 · in -21.2 dBFS · gain +1.2 dB · podcast-06345
(fear, jealousy and envy, doubt·brisk, energised, neutral tension, casual)That is great. I'm gonna
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as fear, jealousy and envy, doubt; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 7.3/10; 28.8s, EN.
152573_00240503 · in -20.0 dBFS · gain +0.0 dB · podcast-05629
(fatigue exhaustion, amusement, jealousy and envy · brisk, energised, neutral tension, casual)like I had to go with the visual of putting on the level five glasses because I was about to fight or flight, you know, that's kind of where we were, and then shifted to the level five.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as fatigue exhaustion, amusement, jealousy and envy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 6.2/10; 8.2s, EN.
152573_00243552 · in -17.4 dBFS · gain -2.6 dB · podcast-01189
(relief, embarrassment, hope enthusiasm optimism· brisk, energised, neutral tension, casual)this is a practice. So, you know, with all my students and clients, they're like, oh my gosh, I got totally sucked in. I actually said this to my partner yesterday.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as relief, embarrassment, hope enthusiasm optimism; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 6.9/10; 8.3s, EN.
152573_00244679 · in -20.6 dBFS · gain +0.6 dB · podcast-04111