VN1 at chain length k=4, all corpora, at the mining floor.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. rule x chain length Sampled from 1,333,380 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
S_AUTH — style: authoritative ↑k-VN1-k4 · #1
This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: authoritative (S_AUTH) low — 0.11, lower than 89 % of clips in this corpus — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.40.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.04, then +0.17 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.26 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.26 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 41 s · fr · emolia
hear it un-normalised (raw levels, max seam 3.7 dB)
k 4d_a 0.396d_b 0.396step_a 0.185step_b 0.185min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_CsCnN6HcmCutrack FR_CsCnN6HcmCutotal 40.7slevel spread 3.7 dBmax seam 3.7 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · fairly smooth, balanced body, average recording, normally alert, light breath
(shame · slow, relaxed, fairly steady, conversational)comme cliente, j'ai mes tatas, j'ai mes cousines, j'ai mes amis, euh, (low mumble) donc, euh, (low mumble)
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as shame; style: conversational, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.9/10; 5.5s, FR.
FR_CsCnN6HcmCu_W000031 · in -17.0 dBFS · gain -3.0 dB · emolia-02889
(disgust, impatience and irritability, contempt·normal-paced, neutral tension, moderately variable, conversational)Mais tu sais, ce que tu dis là, c'est parce que j'ai déjà entendu dire, en fait, tu veux créer un business, mais oui, tu dois répondre aux besoins des gens autour de toi. Donc, c'est vrai que si tu regardes par rapport à ton entourage, tu voudras le vendre à ton entourage. C'est logique après.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disgust, impatience and irritability, contempt; style: conversational, playful; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.9/10; 16.4s, FR.
FR_CsCnN6HcmCu_W000032 · in -20.7 dBFS · gain +0.7 dB · emolia-02889
(measured, slightly relaxed, fairly steady, conversational)non, mais au départ, fin, tu te dis, euh, (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; style: conversational, casual; average recording, no background noise; no dominant emotion; genuineness 3.8/6; vocal-burst blend 2.1/10; 3.0s, FR.
FR_CsCnN6HcmCu_W000033 · in -19.0 dBFS · gain -1.0 dB · emolia-02889
(malevolence malice·normal-paced, slightly relaxed, fairly steady, playful)Oui, mais est-ce que c'est la cible qui est pas bonne? Parce que, (ahem) regarde, si tu fais, (ahem) euh, par exemple, si tu fais une étude de marché, et quand tu étudies bien ton marché, tout à l'heure, tu vas parler de ces tatas, ben, elles se maquillent. Tu veux vendre du make-up, les personnes qui sont autour de toi se maquillent.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as malevolence malice; style: playful; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.0/10; 15.4s, FR.
FR_CsCnN6HcmCu_W000034 · in -20.7 dBFS · gain +0.7 dB · emolia-02889
METL — metallic quality ↑k-VN1-k4 · #2
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with metallic quality (METL) above average — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the range at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.36.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.02, then +0.15 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 70 s · hr · eurospeech
hear it un-normalised (raw levels, max seam 1.1 dB)
k 4d_a 0.355d_b 0.355step_a 0.184step_b 0.184min_cos_consec —min_cos_anchor —dataset eurospeechlang hrspeaker croatia_20140708120641-549track croatia_20140708120641-549total 69.8slevel spread 1.1 dBmax seam 1.1 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-bright, quiet background, energised, neutral tension, moderately variable, some disfluency, wide pitch range
(triumph, pride, malevolence malice · normal-paced, somewhat unclear, light breath, cartoonish)Palili smo te šume da bi deposedirali korisnike koji su na primjer koncesije dobili da bi tamo držali koze, da bi čistili šume i onda se to pretvaralo u vinograde i tome slično. Zadnji puta kada je bio na dnevnom redu ovaj zakon govorio sam,
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as triumph, pride, malevolence malice; style: cartoonish, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.5/10; 19.1s, HR.
croatia_20140708120641-5497_1445775_1464831 · in -9.9 dBFS · gain -10.1 dB · eurospeech-01393
(triumph, elation, pride ·brisk, average clarity, light breath, dramatic)citiram sam sebe da razlika između ovog zakona i onog Zakona o golfu je sljedeća. Ovdje se govori o pravu građenja. To pravo može biti na 30, 40, 99 godina. Tamo je bila riječ o eksproprijaciji. Normalno, može se ekspropriirati isključivo samo
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as triumph, elation, pride; style: dramatic, authoritative; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 8.9/10; 16.9s, HR.
croatia_20140708120641-5497_1464831_1481760 · in -9.4 dBFS · gain -10.6 dB · eurospeech-01393
(contempt, pride, elation · brisk, average clarity, light breath, cartoonish)samo i jedino privatno zemljište. Državno zemljište se ne da ekspropriirati. Kod eksproprijacije kakva god bila pod navodnicima pravična naknada vi tu naknadu morate platiti vlasniku parcele ili predmetne nekretnine, šume ili poljoprivrednog
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, pride, elation; style: cartoonish, authoritative; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 5.0/10; 16.9s, HR.
croatia_20140708120641-5497_1481760_1498672 · in -9.3 dBFS · gain -10.7 dB · eurospeech-01393
(triumph, contempt, concentration·normal-paced, average clarity, normal breath, cartoonish)zemljišta svejedno, dok kod prava građenja tu vrijede gospodo neki drugi principi, tu vrijeme neki posve drugi principi da dalje više ne citiram. Znači, netko će dobiti 100, netko će dobiti 200,
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as triumph, contempt, concentration; style: cartoonish, authoritative; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.9/10; 16.5s, HR.
croatia_20140708120641-5497_1498672_1515168 · in -10.4 dBFS · gain -9.6 dB · eurospeech-01393
GEND — perceived gender ↑k-VN1-k4 · #3
This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with perceived gender (GEND) below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the range at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.54.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.21, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 36 s · zh · emolia
hear it un-normalised (raw levels, max seam 0.6 dB)
k 4d_a 0.543d_b 0.543step_a 0.206step_b 0.206min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00073_S03390track ZH_B00073_S03390total 35.7slevel spread 0.6 dBmax seam 0.6 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · balanced body, average recording, measured, normally alert, audible breath
(intoxication altered states of consciousness, confusion · relaxed, moderately variable, frequent disfluency, casual)除夕之夜,到底是先杀猪呢,还是先杀驴呢?通天教主开口说完之后。
full caption & clip details
A child masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness, confusion; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 3.5/10; 8.9s, ZH.
ZH_B00073_S03390_W000015 · in -17.1 dBFS · gain -2.9 dB · emolia-04002
(sexual lust, intoxication altered states of consciousness, longing·slightly relaxed, fairly steady, some disfluency, casual)(ahem) 然后旁边那帮人,对对对,先杀猪啊,对,先杀猪啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, intoxication altered states of consciousness, longing; style: casual, storytelling; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 4.6/10; 5.6s, ZH.
ZH_B00073_S03390_W000016 · in -17.6 dBFS · gain -2.4 dB · emolia-04002
(intoxication altered states of consciousness, sexual lust, infatuation·relaxed, moderately variable, frequent disfluency, storytelling)啊,对对对啊,这个猪肉鲜美多汁,可以让凡人饱腹吗?西方二圣也开口表达通天教主,马上呵呵一笑嗯。 (low mumble)
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, very wide pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness, sexual lust, infatuation; style: storytelling, playful; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.3/10; 12.3s, ZH.
ZH_B00073_S03390_W000017 · in -17.0 dBFS · gain -3.0 dB · emolia-04002
(sexual lust, intoxication altered states of consciousness, malevolence malice· relaxed, moderately variable, frequent disfluency, storytelling)诸位也不愧是圣人哦,所言即是哦,那驴估计也是想先杀猪哦。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, neutral openness; reads as sexual lust, intoxication altered states of consciousness, malevolence malice; style: storytelling, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.2/10; 8.5s, ZH.
ZH_B00073_S03390_W000018 · in -17.6 dBFS · gain -2.4 dB · emolia-04002
METL — metallic quality ↑k-VN1-k4 · #4
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with metallic quality (METL) below average — 0.27, lower than 73 % of clips in this corpus — and ends with it high at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.51.
It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.23, then +0.21 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 64 s · en · podcast
hear it un-normalised (raw levels, max seam 0.5 dB)
k 4d_a 0.506d_b 0.506step_a 0.227step_b 0.227min_cos_consec 0.8700min_cos_anchor 0.8700dataset podcastlang enspeaker 669327track 669327total 64.0slevel spread 0.5 dBmax seam 0.5 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · slightly cool, neutral-bright, average recording, quiet background, some disfluency
(affection, bitterness, contemplation · measured, normally alert, slightly relaxed, cartoonish)Ayuda to evolution, but no to abreat your sentiments to no percibes. In this moment is when you get a message of other planos, which no need for a mensage angelic, a person who has been the same, or would be of your life, but in this moment you are able to recipe a conflict and
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, bitterness, contemplation; style: cartoonish, narration; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 4.5/10; 27.6s, EN.
669327_00232244 · in -19.8 dBFS · gain -0.2 dB · podcast-05417
(emotional numbness, sourness, teasing· measured, very low-energy, neutral tension, storytelling)this conflict has to be more grave, it would be a tonter.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, moderate pitch range, audible breath; affect is mildly negative, slightly dominant, neutral openness; reads as emotional numbness, sourness, teasing; style: storytelling, whispered; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.6/10; 6.1s, EN.
669327_00235000 · in -20.1 dBFS · gain +0.1 dB · podcast-04888
(disappointment, sadness, sourness ·normal-paced, normally alert, neutral tension, cartoonish)It's a conflict, because you're accustomed to start with males, and it's more feliz, and no cuad.
full caption & clip details
An elderly feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as disappointment, sadness, sourness; style: cartoonish, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.9/10; 16.3s, EN.
669327_00235612 · in -19.5 dBFS · gain -0.5 dB · podcast-05430
(fatigue exhaustion, distress, sadness ·measured, normally alert, neutral tension, casual)If you want to say you can give a novel of me,
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, audible breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fatigue exhaustion, distress, sadness; style: casual, storytelling; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.5/10; 13.6s, EN.
669327_00237240 · in -19.6 dBFS · gain -0.4 dB · podcast-05422
VULN — vulnerability ↓k-VN1-k4 · #5
This is a VoiceNet dimension, not an emotion: vulnerability (VULN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with vulnerability (VULN) above average — 0.66, higher than 66 % of clips in this corpus — and works its way down to low at 0.14, lower than 86 % of clips in this corpus. That is a total fall of 0.51.
It takes 4 clips to get there. Clip to clip the moves are -0.21, then -0.08, then -0.22 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · en · emolia
hear it un-normalised (raw levels, max seam 0.9 dB)
k 4d_a -0.514d_b -0.514step_a 0.218step_b 0.218min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_inkX_lbXUWItrack EN_inkX_lbXUWItotal 24.7slevel spread 0.9 dBmax seam 0.9 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, normally alert, slightly relaxed, clear, light breath
(normal-paced, fairly steady, almost no disfluency, narration)He said it in October 2001, when the Afghanistan war was not older than three days.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.4/10; 5.4s, EN.
EN_inkX_lbXUWI_W000225 · in -19.0 dBFS · gain -1.0 dB · emolia-01329
(normal-paced, fairly steady, little disfluency, narration)Now, here in Germany you have intelligent people, courageous people, who understand that international law may not be breached just like that. Who understand that one can just say
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, whispered; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.5/10; 10.2s, EN.
EN_inkX_lbXUWI_W000226 · in -18.1 dBFS · gain -1.9 dB · emolia-01329
(relief, pride, triumph· normal-paced, fairly steady, no disfluency, formal)Yes, that country is behind it. I can't prove it, but that's how it is.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, pride, triumph; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.5/10; 3.6s, EN.
EN_inkX_lbXUWI_W000227 · in -18.8 dBFS · gain -1.2 dB · emolia-01329
(disappointment·measured, steady, almost no disfluency, formal)You can't do that. You must submit evidence. This evidence was not delivered.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, very full; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disappointment; style: formal, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.4/10; 5.0s, EN.
EN_inkX_lbXUWI_W000228 · in -19.0 dBFS · gain -1.0 dB · emolia-01329
R_NASL — resonance: nasal ↑k-VN1-k4 · #6
This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: nasal (R_NASL) around average — 0.47, lower than 53 % of clips in this corpus — and ends with it high at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.37.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.03, then +0.12 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.24 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.24, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, quiet background
(affection, teasing, sourness · fast, normally alert, slightly relaxed, monologue)sei carina e simpatica adesso, sarei stata carina, simpatica, allora, è affascinante. Poi è chiaro che tutte le mamme faranno la gara a fare la festa per invitare la signora Ferragni.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, teasing, sourness; style: monologue, authoritative; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.9/10; 11.0s, IT.
183603_00026503 · in -25.4 dBFS · gain +5.4 dB · podcast-01543
(amusement, sourness, teasing ·normal-paced, normally alert, neutral tension, casual)parlato di caccole. Pensavo anche al moccolo. Caldissima, perché tu sai che io non riesco a stare ferma. (childlike giggle) Allora, ieri con Alessandro Bertoletti che fa con me i video games chiacchieravamo in maniera molto simpatica. Un po' come con te ci prendiamo in giro. Sul fatto che dobbiamo recuperare soldi per comprare una X qualcosa, non c'è più
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, sourness, teasing; style: casual, conversational; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.8/10; 26.8s, IT.
183603_00032664 · in -24.2 dBFS · gain +4.2 dB · podcast-00654
(affection, contentment, pleasure ecstasy· normal-paced, normally alert, slightly relaxed, casual)Cioè lo si sente sempre. Dici ok, ragazze, belle bellissime. Ok,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection, contentment, pleasure ecstasy; style: casual, playful; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.5/10; 4.2s, IT.
183603_00039560 · in -22.7 dBFS · gain +2.7 dB · podcast-01563
(contempt, sourness, affection · normal-paced, subdued, neutral tension, monologue)perché tutte le ragazzine che senti adesso. Tutte le ragazzine che senti, anch'io ho cominciato sulla tua scia a uscire dal mio micro profilo di Facebook dove pubblico solo argomenti di economia pallosissimi, aggiungeresti tu. Ma spesso le ragazze dicono interessantissimi, chiarissimi e leggibilissimi. Pilloledì.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, sourness, affection; style: monologue, cartoonish; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.2/10; 24.6s, IT.
183603_00082976 · in -24.2 dBFS · gain +4.2 dB · podcast-01515
S_PLAY — style: playful ↓k-VN1-k4 · #7
This is a VoiceNet dimension, not an emotion: style: playful (S_PLAY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: playful (S_PLAY) above average — 0.66, higher than 66 % of clips in this corpus — and works its way down to below average at 0.25, lower than 75 % of clips in this corpus. That is a total fall of 0.41.
It takes 4 clips to get there. Clip to clip the moves are -0.15, then -0.03, then -0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: an adult masculine voice · slightly rough, balanced body, quiet background, normally alert, fairly steady
(intoxication altered states of consciousness, fatigue exhaustion, contemplation · measured, neutral tension, frequent disfluency, casual)Yeah, but like (wistful sigh) you know, come when I went to Italy one time, like you could tell it's like some things are just kind of like you know, in the because I don't really speak a whole lot of Spanish, but I can see like the similarities between like Spanish and like Italian and stuff like that.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, contemplation; style: casual, monologue; average recording, quiet background; genuineness 5.8/6; vocal-burst blend 8.2/10; 15.3s, EN.
673834_00326080 · in -44.1 dBFS · gain +24.1 dB · podcast-03746
(contemplation, relief· measured, relaxed, frequent disfluency, casual)Yeah, but like one thing that I definitely you know was lost in translation for me, as far as like you know that goes, was (low mumble) um you know
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, relief; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.4/10; 8.4s, EN.
673834_00328800 · in -45.2 dBFS · gain +25.2 dB · podcast-03102
(relief, jealousy and envy, astonishment surprise·normal-paced, neutral tension, frequent disfluency, casual)at work, especially, you know, just doing my own thing, not really minding my own business, and it's like I was told to do something by one supervisor, and then another supervisor comes up to me and says, Oh, hey, did you do this? It's like no, I was told to do that. Like, oh well, they told me that you were supposed to do this. That's a lost in translation kind of example
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, jealousy and envy, astonishment surprise; style: casual, monologue; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 5.6/10; 22.3s, EN.
673834_00329632 · in -42.9 dBFS · gain +22.9 dB · podcast-03090
(emotional numbness, fatigue exhaustion, disappointment·measured, slightly relaxed, some disfluency, monologue)happened to me a lot at work, you know. It's just nobody really communicated the same thing, or they just interpreted it the
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, fairly guarded; reads as emotional numbness, fatigue exhaustion, disappointment; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.1/10; 5.8s, EN.
673834_00331960 · in -44.4 dBFS · gain +24.4 dB · podcast-03746
S_RANT — style: ranting ↓k-VN1-k4 · #8
This is a VoiceNet dimension, not an emotion: style: ranting (S_RANT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: ranting (S_RANT) above average — 0.67, higher than 67 % of clips in this corpus — and works its way down to below average at 0.35, lower than 65 % of clips in this corpus. That is a total fall of 0.32.
It takes 4 clips to get there. Clip to clip the moves are -0.24, then +0.11, then -0.19 — not a clean run: step 2 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 24 s · en · emolia
k 4d_a -0.322d_b -0.322step_a 0.240step_b 0.240min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B0noYV5dD5Ytrack EN_B0noYV5dD5Ytotal 24.3slevel spread 1.4 dBmax seam 1.4 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an elderly masculine voice · neutral-toned, slightly dark, average recording, quiet background
(slow, normally alert, relaxed, casual)Uh, (low mumble) not long before that we read about the National Archives that spent tons of money
full caption & clip details
An elderly masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.1/10; 7.5s, EN.
EN_B0noYV5dD5Y_W000035 · in -19.1 dBFS · gain -0.9 dB · emolia-01056
(slow, very low-energy, relaxed, monologue)(low mumble) Uhm, on a (low mumble) (low mumble) computer project that apparently isn't working very well. (low mumble)
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is neutral, submissive, neutral openness; no dominant emotion; style: monologue, ASMR; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.9/10; 7.9s, EN.
EN_B0noYV5dD5Y_W000036 · in -20.3 dBFS · gain +0.3 dB · emolia-01056
(measured, normally alert, slightly relaxed, casual)You may recall Metro 5, which was a system that, (low mumble) uh,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 3.8s, EN.
EN_B0noYV5dD5Y_W000037 · in -18.9 dBFS · gain -1.1 dB · emolia-01056
(measured, normally alert, slightly relaxed, casual)What's supposed to give us wireless access across the city of Portland? That is...
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.3/10; 4.7s, EN.
EN_B0noYV5dD5Y_W000038 · in -19.5 dBFS · gain -0.5 dB · emolia-01056
EXPL — expressiveness ↑k-VN1-k4 · #9
This is a VoiceNet dimension, not an emotion: expressiveness (EXPL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with expressiveness (EXPL) low — 0.13, lower than 87 % of clips in this corpus — and ends with it around average at 0.48, lower than 52 % of clips in this corpus. That is a total rise of 0.35.
It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.02, then +0.21 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.73 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.73 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · en · emolia
k 4d_a 0.350d_b 0.350step_a 0.207step_b 0.207min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_KCw_DM_AUg0track EN_KCw_DM_AUg0total 37.2slevel spread 1.3 dBmax seam 1.3 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, moderate pitch range, light breath
(normal-paced, fairly steady, some disfluency, monologue)The useful hints in the classifying the soil once we have the grain size distribution and Atterberg limits data.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.7/10; 6.0s, EN.
EN_KCw_DM_AUg0_W000143 · in -12.3 dBFS · gain -7.7 dB · emolia-01350
(measured, fairly steady, frequent disfluency, didactic)Always (low mumble) begin on the left hand side with A1A group.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.6/10; 4.9s, EN.
EN_KCw_DM_AUg0_W000144 · in -11.4 dBFS · gain -8.6 dB · emolia-01350
(concentration· measured, moderately variable, frequent disfluency, didactic)So once the, the chart which actually (low mumble) a1a group, when one, once we (low mumble) (low mumble) eliminate a1a group then go to a2 like that and check each of the criteria. If any criterion is not met.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: didactic, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.2/10; 14.0s, EN.
EN_KCw_DM_AUg0_W000145 · in -11.3 dBFS · gain -8.7 dB · emolia-01350
(concentration ·normal-paced, fairly steady, some disfluency, didactic)(low mumble) Uh, step to the right and repeat the process. So the chart which is actually not given but if you have a chart always begin on the left hand side with A1A group and check each side each of the criteria.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: didactic, authoritative; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.0/10; 11.8s, EN.
EN_KCw_DM_AUg0_W000146 · in -12.7 dBFS · gain -7.3 dB · emolia-01350
GEND — perceived gender ↑k-VN1-k4 · #10
This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with perceived gender (GEND) around average — 0.51, right about the corpus median — and ends with it above average at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.24.
It takes 4 clips to get there. Clip to clip the moves are +0.05, then +0.13, then +0.06 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · zh · emolia
k 4d_a 0.238d_b 0.238step_a 0.135step_b 0.135min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00072_S03753track ZH_B00072_S03753total 24.8slevel spread 5.2 dBmax seam 5.2 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · fairly smooth, thin, average recording, energised, moderately variable, wide pitch range
An adult masculine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, fear, malevolence malice; style: storytelling, dramatic; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 5.3/10; 8.7s, ZH.
ZH_B00072_S03753_W000033 · in -24.0 dBFS · gain +4.0 dB · emolia-03995
(pain, confusion·normal-paced, fully relaxed, no disfluency, storytelling)赖昌星,那原本就聪明过人。
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as pain, confusion; style: storytelling, casual; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 5.1/10; 3.3s, ZH.
ZH_B00072_S03753_W000034 · in -18.8 dBFS · gain -1.2 dB · emolia-03995
(impatience and irritability·fast, slightly tense, some disfluency, storytelling)又练大人情,办起一些微妙的事儿来,那简直就是得心应手,滴水不漏。
full caption & clip details
An adult masculine voice; delivery is energised, fast, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as impatience and irritability; style: storytelling, dramatic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 6.0/10; 6.6s, ZH.
ZH_B00072_S03753_W000035 · in -21.6 dBFS · gain +1.6 dB · emolia-03995
(disgust, contempt, malevolence malice· fast, slightly relaxed, no disfluency, storytelling)许甘露、赖昌星两人间的关系也由此是突飞猛进。
full caption & clip details
An adult masculine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; very clear, no disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, contempt, malevolence malice; style: storytelling, dramatic; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 4.8/10; 5.7s, ZH.
ZH_B00072_S03753_W000036 · in -22.8 dBFS · gain +2.8 dB · emolia-03995
WARM — warmth ↑k-VN1-k4 · #11
This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with warmth (WARM) low — 0.17, lower than 83 % of clips in this corpus — and ends with it above average at 0.67, higher than 67 % of clips in this corpus. That is a total rise of 0.50.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.14, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 68 s · sl · eurospeech
k 4d_a 0.501d_b 0.501step_a 0.200step_b 0.200min_cos_consec —min_cos_anchor —dataset eurospeechlang slspeaker slovenia_slovenia_23_Rednatrack slovenia_slovenia_23_Rednatotal 68.5slevel spread 0.7 dBmax seam 0.7 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, normally alert, moderate pitch range, light breath
(triumph, thankfulness gratitude, affection · brisk, slightly relaxed, fairly steady, authoritative)želijo pridobiti dovoljenje za dajanje zagotovila o trajnostnem poročanju, potrebno strokovno znanje, predvsem pa izkušnje pri zagotavljanju trajnostnega poročanja. Določbe bodo začele veljati od leta 2024, odvisne so od velikosti družbe, v celoti pa bodo začele veljati s poslovnim letom 2028.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as triumph, thankfulness gratitude, affection; style: authoritative, monologue; average recording, some background noise; genuineness 0.9/6; vocal-burst blend 3.4/10; 17.9s, SL.
slovenia_slovenia_23_Redna_1_24092024_3115504_3133376 · in -23.6 dBFS · gain +3.6 dB · eurospeech-02665
(bitterness, pain, distress·normal-paced, neutral tension, fairly steady, monologue)Navedeno pomeni nove obremenitve za gospodarstvo. V obrazložitvi predloga zakona je zapisano, da je namen direktive in posledično predloga zakona izboljšati družbeno odgovornost podjetij, ki naj bi pri svojem poslovanju upoštevala tudi družbena in okoljska vprašanja.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, pain, distress; style: monologue, authoritative; below-average recording, quiet background; genuineness 2.8/6; vocal-burst blend 6.2/10; 15.3s, SL.
slovenia_slovenia_23_Redna_1_24092024_3133376_3148703 · in -23.6 dBFS · gain +3.6 dB · eurospeech-02665
(bitterness, pride, sourness· normal-paced, neutral tension, moderately variable, monologue)navedenem se seveda poraja dvom, saj zgolj dodatne obveznosti, kar se revizije tiče, ne bodo izboljšale družbene odgovornosti podjetij. Veliko je drugih vzvodov, ki podjetja spodbujajo k trajnosti in okoljski odgovornosti, revizija pa je lahko v praksi zgolj dodatno breme.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, pride, sourness; style: monologue, casual; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.9/10; 17.1s, SL.
slovenia_slovenia_23_Redna_1_24092024_3148703_3165808 · in -22.9 dBFS · gain +2.9 dB · eurospeech-02665
(thankfulness gratitude, affection, bitterness ·brisk, neutral tension, fairly steady, monologue)Predlog zakona tudi ureja pristojnost Slovenskega inštituta za revizijo v delu sprejemanja oziroma zagotavljanja prevodov pravil notranjega revidiranja in spremembe, ki se nanašajo na Agencijo za javni nadzor nad revidiranjem, kot na primer sejnine in imenovanje direktorja ter znižanje števila članov strokovnega sveta,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, affection, bitterness; style: monologue, authoritative; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 6.5/10; 17.7s, SL.
slovenia_slovenia_23_Redna_1_24092024_3165808_3183520 · in -23.5 dBFS · gain +3.5 dB · eurospeech-02665
R_MIXD — resonance: mixed ↑k-VN1-k4 · #12
This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: mixed (R_MIXD) below average — 0.36, lower than 64 % of clips in this corpus — and ends with it high at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.52.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.21, then +0.08 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 35 s · en · emolia
k 4d_a 0.523d_b 0.523step_a 0.238step_b 0.238min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_-wTP40GY0kctrack EN_-wTP40GY0kctotal 34.6slevel spread 1.1 dBmax seam 0.6 dB
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · fairly smooth, balanced body, energised, moderately variable, some disfluency, wide pitch range, light breath
(anger, bitterness, impatience and irritability · brisk, slightly tense, clear, casual)Whatever you can do to get people to come into your stream, I've worn Batman, Captain America, Thor, Deadpool, I've worn all kinds of costumes, and that worked to get me to go off in the algorithm. The key is to figure out
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as anger, bitterness, impatience and irritability; style: casual, playful; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 3.9/10; 13.7s, EN.
EN_-wTP40GY0kc_W000420 · in -17.6 dBFS · gain -2.4 dB · emolia-00930
(sourness, contempt, impatience and irritability ·normal-paced, slightly tense, average clarity, casual)What are you going to do with your appearance that'll help you stand out, but it'll also feel uniquely you?
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as sourness, contempt, impatience and irritability; style: casual, storytelling; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.5/10; 6.5s, EN.
EN_-wTP40GY0kc_W000421 · in -18.2 dBFS · gain -1.8 dB · emolia-00930
(jealousy and envy, relief, teasing·brisk, neutral tension, average clarity, casual)And (ahem) I like just having the bare shoulders, it works consistently for me. Joseph says, can't hate the boobs, right? Israel says, thanks for the shorts.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as jealousy and envy, relief, teasing; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.5/10; 8.4s, EN.
EN_-wTP40GY0kc_W000422 · in -18.7 dBFS · gain -1.3 dB · emolia-00930
(elation· brisk, slightly relaxed, clear, authoritative)And the last thing you can do with your livestream is make the was live version of this watchable.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation; style: authoritative, dramatic; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.4/10; 5.5s, EN.
EN_-wTP40GY0kc_W000423 · in -18.2 dBFS · gain -1.8 dB · emolia-00930
ARSH — harshness of articulation ↑k-VN1-k4 · #13
This is a VoiceNet dimension, not an emotion: harshness of articulation (ARSH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with harshness of articulation (ARSH) below average — 0.30, lower than 70 % of clips in this corpus — and ends with it high at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.53.
It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.20, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 54 s · hr · eurospeech
k 4d_a 0.534d_b 0.534step_a 0.229step_b 0.229min_cos_consec —min_cos_anchor —dataset eurospeechlang hrspeaker croatia_20230419161221-153track croatia_20230419161221-153total 54.2slevel spread 6.8 dBmax seam 4.0 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-bright, average recording, quiet background, some disfluency, light breath
(shame, disgust, embarrassment · normal-paced, normally alert, neutral tension, casual)prema policijskom službeniku plaća do 100 eura, a prema carinskom službeniku 10.000 eura. Maksimalan iznos 100 eura naprema maksimalnog iznosa ovaj 10.000 eura.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, disgust, embarrassment; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 6.7/10; 12.8s, HR.
croatia_20230419161221-15313_3384560_3397328 · in -22.4 dBFS · gain +2.4 dB · eurospeech-01586
(doubt, confusion· normal-paced, normally alert, neutral tension, monologue)Jakšić. **Jakšić, Mišel (SDP)** Hvala poštovani potpredsjedniče. Poštovana državna tajnice da li imate ikakvu projekciju s obzirom da ste rekli da je cilj ovih kazni da se odvrati
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt, confusion; style: monologue, authoritative; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.8/10; 11.7s, HR.
croatia_20230419161221-15313_3397328_3409024 · in -18.4 dBFS · gain -1.6 dB · eurospeech-01586
(thankfulness gratitude, interest·brisk, normally alert, slightly relaxed, authoritative)od činjenja prekršaja? Koja je struktura ljudi koji rade te prekršaje i na temelju čega onda donosite zaključak da upravo ti ljudi više te prekršaje neće raditi, a znamo i sami da dobar dio tih ljudi ne može platiti maltene nikakvu kaznu,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, interest; style: authoritative, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 5.4/10; 14.3s, HR.
croatia_20230419161221-15313_3409024_3423312 · in -17.9 dBFS · gain -2.1 dB · eurospeech-01586
(relief, hope enthusiasm optimism, shame· brisk, energised, neutral tension, dramatic)a očekujemo da će sad plaćati drakonske kazne i vjerujemo da nećemo puniti zatvore i s druge strane vjerujemo da to neće kod svih običnih građana ove zemlje koji nemaju problema prekršajne naravni opet izazivati podozrivosti,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, hope enthusiasm optimism, shame; style: dramatic, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 6.0/10; 15.0s, HR.
croatia_20230419161221-15313_3423312_3438352 · in -15.6 dBFS · gain -4.4 dB · eurospeech-01586
STNC — stance / assertiveness ↓k-VN1-k4 · #14
This is a VoiceNet dimension, not an emotion: stance / assertiveness (STNC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with stance / assertiveness (STNC) high — 0.83, higher than 83 % of clips in this corpus — and works its way down to below average at 0.33, lower than 67 % of clips in this corpus. That is a total fall of 0.50.
It takes 4 clips to get there. Clip to clip the moves are -0.12, then -0.22, then -0.16 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 18 s · en · emolia
k 4d_a -0.504d_b -0.504step_a 0.221step_b 0.221min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_29keXzPXgMYtrack EN_29keXzPXgMYtotal 18.3slevel spread 2.4 dBmax seam 2.4 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, light breath
(normal-paced, some disfluency, average clarity, casual)Alright, so I'm getting a little off track. Anyway, (ahem) the point is it's one of these things where.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.8/10; 4.6s, EN.
EN_29keXzPXgMY_W000115 · in -21.1 dBFS · gain +1.1 dB · emolia-02098
(disappointment, bitterness, impatience and irritability· normal-paced, some disfluency, average clarity, casual)(low mumble) You have a business that's continuing to grow, but they're not making any more money. They're actually making less.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, bitterness, impatience and irritability; style: casual, conversational; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.0/10; 6.6s, EN.
EN_29keXzPXgMY_W000116 · in -18.6 dBFS · gain -1.4 dB · emolia-02098
(slow, frequent disfluency, somewhat unclear, casual)So that needs to be fixed. Right?
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; style: casual, conversational; good recording, no background noise; no dominant emotion; genuineness 3.0/6; vocal-burst blend 0.4/10; 3.3s, EN.
EN_29keXzPXgMY_W000117 · in -20.6 dBFS · gain +0.6 dB · emolia-02098
(normal-paced, some disfluency, average clarity, casual)Electric vehicles is pretty small when you really think about it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 3.6/10; 3.3s, EN.
EN_29keXzPXgMY_W000123 · in -18.9 dBFS · gain -1.1 dB · emolia-02098
FULL — fullness of tone ↓k-VN1-k4 · #15
This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with fullness of tone (FULL) above average — 0.60, higher than 60 % of clips in this corpus — and works its way down to below average at 0.27, lower than 73 % of clips in this corpus. That is a total fall of 0.33.
It takes 4 clips to get there. Clip to clip the moves are -0.12, then -0.02, then -0.19 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, moderately variable, some disfluency, average clarity
(sexual lust, fear, jealousy and envy · brisk, normally alert, slightly relaxed, casual)And I think Zen Zen is treated with like a lot of kindness by the film. She's got a lot of anxiety. Uh (ahem) there's a scene where like she rips off her mother's wig on accident and sees her bald and sort of freaks out and starts like screaming and having a panic attack. It would be very easy to like be like, oh, she's just this autistic weirdo, she can't handle change.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as sexual lust, fear, jealousy and envy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.2/6; vocal-burst blend 9.2/10; 25.6s, EN.
958059_00144112 · in -19.1 dBFS · gain -0.9 dB · podcast-06487
(jealousy and envy, affection, interest·normal-paced, normally alert, slightly relaxed, casual)But while she's freaking out, it cuts to memories of her as a child playing with her mother's hair, like when she was a little baby, like this like really intimate connection. And so what you've got is this like really kind of beautiful but also like sad moment where Zen is upset because her mother has changed and this like intimate relationship they have has changed because of the lack of hair. But because she's not very verbal because she doesn't have a whole lot of verbal skills, she can't articulate that.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, affection, interest; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 3.0/6; vocal-burst blend 6.5/10; 25.0s, EN.
958059_00146672 · in -19.9 dBFS · gain -0.1 dB · podcast-06483
(interest, disgust, impatience and irritability· normal-paced, normally alert, neutral tension, casual)So she's just upset. If you read stuff about autism or you interact with people who who are on the autis autism spectrum, you know, this is like a consistent issue is lacking the techniques to verbalize. It paints Zen as someone who's real and just has trouble communicating rather than like a space alien. You know, there's a there's another section where dude, it's so funny too.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as interest, disgust, impatience and irritability; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 8.1/10; 21.2s, EN.
958059_00149168 · in -18.3 dBFS · gain -1.7 dB · podcast-06485
(awe, jealousy and envy, amusement·brisk, energised, neutral tension, casual)There's this section where she like is just sitting there and a fly flies near her and she just grabs it out of the air and puts it in her mouth. It's like totally like, look how good her reflexes are. But she gets sick because of it. And so as an adult, she's afraid of flies. And so again, you know, it could be like, oh, what a weirdo, why is she afraid of flies? It's almost always got context, which I think makes it a little more kind of a representation.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as awe, jealousy and envy, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 8.7/10; 21.2s, EN.
958059_00151288 · in -16.1 dBFS · gain -3.9 dB · podcast-06481
S_PLAY — style: playful ↓k-VN1-k4 · #16
This is a VoiceNet dimension, not an emotion: style: playful (S_PLAY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: playful (S_PLAY) high — 0.76, higher than 76 % of clips in this corpus — and works its way down to below average at 0.42, lower than 58 % of clips in this corpus. That is a total fall of 0.35.
It takes 4 clips to get there. Clip to clip the moves are -0.13, then -0.06, then -0.15 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 34 s · en · emolia
k 4d_a -0.345d_b -0.345step_a 0.154step_b 0.154min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_7igiv_Lfmz8track EN_7igiv_Lfmz8total 34.4slevel spread 3.8 dBmax seam 3.8 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(measured, frequent disfluency, casual, playful)Uh, (low mumble) next to him is Robert Pinsky, a (low mumble) three-time U.S. poet laureate.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.3/10; 5.3s, EN.
EN_7igiv_Lfmz8_W000008 · in -18.9 dBFS · gain -1.1 dB · emolia-02401
(normal-paced, some disfluency, casual, monologue)The author of 19 books, which makes me tired just thinking about.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.5/10; 3.7s, EN.
EN_7igiv_Lfmz8_W000009 · in -22.8 dBFS · gain +2.8 dB · emolia-02401
(normal-paced, some disfluency, casual, monologue)(low mumble) Uhm, and the William Fairfield Warren Distinguished Professor of English and Creative Writing at Boston University.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.0/10; 5.8s, EN.
EN_7igiv_Lfmz8_W000010 · in -19.5 dBFS · gain -0.5 dB · emolia-02401
(pride, hope enthusiasm optimism, elation· normal-paced, some disfluency, casual)Uh, (low mumble) he has won the William Carlos Williams Prize from the Poetry Society of America, the Harold Washington Award from the City of Chicago, and a Lifetime Achievement Award from, uh, (low mumble) the Pan American Center. (low mumble) Uhm, and you are currently, uh, (low mumble) working on furthering digital education. Is that right? By teaching the art of poetry through edX.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride, hope enthusiasm optimism, elation; style: casual; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.4/10; 19.2s, EN.
EN_7igiv_Lfmz8_W000011 · in -22.2 dBFS · gain +2.2 dB · emolia-02401
EMPH — emphasis / stress strength ↑k-VN1-k4 · #17
This is a VoiceNet dimension, not an emotion: emphasis / stress strength (EMPH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with emphasis / stress strength (EMPH) low — 0.24, lower than 76 % of clips in this corpus — and ends with it high at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.58.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.24, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.28 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.23 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.28, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 43 s · en · podcast
k 4d_a 0.582d_b 0.582step_a 0.242step_b 0.242min_cos_consec 0.2303min_cos_anchor 0.2773dataset podcastlang enspeaker 882341track 882341total 42.6slevel spread 12.5 dBmax seam 9.4 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, some disfluency
(contentment, pleasure ecstasy, relief · normally alert, slightly relaxed, fairly steady, casual)So just like a nice light low simmer, (ahem) um, just letting everything combine.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, pleasure ecstasy, relief; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.0/6; vocal-burst blend 2.9/10; 4.9s, EN.
882341_00086664 · in -27.5 dBFS · gain +7.5 dB · podcast-00967
(fatigue exhaustion, hope enthusiasm optimism, contentment · normally alert, slightly relaxed, fairly steady, casual)Yeah, there are multiple ways to do it. So it's like if you're making another solve, or you can do it over the heat, which is the fastest way to incorporate any of the properties from the herbs into the tallow. So that would be about three to twelve hours of simmering. But if you have the patience, you really can put the herbs in it and store it away in a jar in your cabinet for weeks, and it'll do the same thing. It's just gonna be gentler in a less fast process.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, hope enthusiasm optimism, contentment; style: casual, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 6.7/10; 28.4s, EN.
882341_00087176 · in -36.9 dBFS · gain +16.9 dB · podcast-05290
(doubt, astonishment surprise, infatuation· normally alert, slightly relaxed, fairly steady, casual)Right, right. But can you imagine the smells and the aroma with that?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, astonishment surprise, infatuation; style: casual, conversational; good recording, no background noise; genuineness 4.6/6; vocal-burst blend 4.1/10; 3.8s, EN.
882341_00090040 · in -29.4 dBFS · gain +9.4 dB · podcast-00959
(astonishment surprise, relief, elation·energised, neutral tension, moderately variable, conversational)Oh my gosh. Okay. So we've got we combined the tallow and you said it
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, slightly vulnerable; reads as astonishment surprise, relief, elation; style: conversational, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.0/10; 5.0s, EN.
882341_00090728 · in -24.3 dBFS · gain +4.3 dB · podcast-00950
RANG — pitch range used ↓k-VN1-k4 · #18
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with pitch range used (RANG) at the very top of the range — 0.93, higher than 93 % of clips in this corpus — and works its way down to below average at 0.36, lower than 64 % of clips in this corpus. That is a total fall of 0.57.
It takes 4 clips to get there. Clip to clip the moves are -0.20, then -0.14, then -0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.20 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.49 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.20, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 73 s · it · podcast
k 4d_a -0.570d_b -0.570step_a 0.226step_b 0.226min_cos_consec 0.4873min_cos_anchor 0.1972dataset podcastlang itspeaker 463154track 463154total 73.1slevel spread 3.2 dBmax seam 3.2 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice
(interest, elation, pleasure ecstasy · fast, energised, neutral tension, dramatic)infettivo. E una sala operatoria dove è impossibile avere le (ahem) anestesie e lavorare senza anestesia, diventa più nemmeno che una camera di tortura, il che distrugge prospetticamente anche le persone che magari riescono a essere operate perché avranno un trauma così profondo che la loro vita, la loro proiezione di futuro
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as interest, elation, pleasure ecstasy; style: dramatic, storytelling; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 8.5/10; 24.5s, IT.
463154_00180272 · in -16.4 dBFS · gain -3.6 dB · podcast-06434
(fast, normally alert, neutral tension, conversational)sarà praticamente annichilita. Una sala a parto impossibilitata a compiere il suo lavoro, cancella il futuro. È questa la prospettiva.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: conversational, storytelling; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.5/10; 8.2s, IT.
463154_00182721 · in -19.6 dBFS · gain -0.4 dB · podcast-04177
(bitterness, shame, contentment·brisk, normally alert, slightly relaxed, monologue)Nicoletta, io ti ringrazio. Ti saluto ricordando che il Washington Post, ad esempio, ha pubblicato nomi, in alcuni casi, anche le foto dei 18.500 bambini uccisi nella striscia di Gaza, ed è proprio, (ahem) soprattutto le immagini, ma anche l'elenco dei nomi, dà proprio l'idea di questo obiettivo di distruzione del futuro, che forse all'interno di una immane tragedia, l'aspetto peggiore ancora. Grazie mille per questa chiacchierata.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, shame, contentment; style: monologue, authoritative; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 9.1/10; 27.7s, IT.
463154_00193040 · in -18.3 dBFS · gain -1.7 dB · podcast-06291
(disgust, sadness, disappointment·normal-paced, normally alert, slightly relaxed, casual)sono persone, e questa operazione del Washington Post, va in qualche modo a provare a diminuire, cancellare la disumanizzazione, appunto, di
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, sadness, disappointment; style: casual; average recording, some background noise; genuineness 4.2/6; vocal-burst blend 3.5/10; 12.2s, IT.
463154_00196404 · in -18.3 dBFS · gain -1.7 dB · podcast-00240
COGL — cognitive load ↓k-VN1-k4 · #19
This is a VoiceNet dimension, not an emotion: cognitive load (COGL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with cognitive load (COGL) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.08, lower than 92 % of clips in this corpus. That is a total fall of 0.57.
It takes 4 clips to get there. Clip to clip the moves are -0.17, then -0.21, then -0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 43 s · zh · emolia
k 4d_a -0.568d_b -0.568step_a 0.206step_b 0.206min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00070_S07960track ZH_B00070_S07960total 43.5slevel spread 1.2 dBmax seam 1.2 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · average recording, quiet background
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, confusion, impatience and irritability; style: formal, authoritative; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.1/10; 4.5s, ZH.
ZH_B00070_S07960_W000041 · in -19.6 dBFS · gain -0.4 dB · emolia-03972
(intoxication altered states of consciousness, confusion, impatience and irritability ·measured, normally alert, slightly relaxed, cartoonish)守门后才敢嘛传达太后的命令吗?太后这究竟在将军这一个人等一轮?真的来哦,一堆说话。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness, confusion, impatience and irritability; style: cartoonish, authoritative; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.0/10; 11.5s, ZH.
ZH_B00070_S07960_W000042 · in -20.7 dBFS · gain +0.7 dB · emolia-03972
A child masculine voice; delivery is energised, normal-paced, tense, moderately variable; timbre is slightly cool, dark, very rough, thin; slurred, frequent disfluency, very wide pitch range, audible breath; affect is positive, slightly dominant, guarded; reads as triumph, pride; style: storytelling, cartoonish; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.9/10; 8.6s, ZH.
ZH_B00070_S07960_W000044 · in -19.6 dBFS · gain -0.4 dB · emolia-03972
EXPL — expressiveness ↑k-VN1-k4 · #20
This is a VoiceNet dimension, not an emotion: expressiveness (EXPL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with expressiveness (EXPL) above average — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the range at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.21.
It takes 4 clips to get there. Clip to clip the moves are -0.05, then +0.14, then +0.12 — not a clean run: step 1 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 41 s · en · emolia
k 4d_a 0.208d_b 0.208step_a 0.144step_b 0.144min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00068_S05977track EN_B00068_S05977total 41.3slevel spread 3.5 dBmax seam 3.5 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · slightly thin, quiet background, normal breath
(concentration · normal-paced, very low-energy, relaxed, casual)And now I'm gonna go ahead and circle right in and give that little dimension to his ear.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration; style: casual, whispered; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.2/10; 7.7s, EN.
EN_B00068_S05977_W000001 · in -21.3 dBFS · gain +1.3 dB · emolia-01554
(sexual lust, affection· normal-paced, normally alert, slightly relaxed, casual)His arm, we're still going to keep the basic over the shape of his arm. Make it nice and skinny, but then it's going to be
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, affection; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 8.0s, EN.
EN_B00068_S05977_W000002 · in -21.0 dBFS · gain +1.0 dB · emolia-01554
(concentration, contentment, intoxication altered states of consciousness·measured, very low-energy, relaxed, whispered)And then I'm gonna, after this part, after I swoop this part here, you're gonna swoop right back in. Round here. It does not connect. And then we're gonna go right back in. And you make it thicker.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, slightly thin; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration, contentment, intoxication altered states of consciousness; style: whispered, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.0/10; 17.8s, EN.
EN_B00068_S05977_W000003 · in -22.9 dBFS · gain +2.9 dB · emolia-01554
(intoxication altered states of consciousness, affection, pleasure ecstasy·normal-paced, very low-energy, relaxed, casual)Bring this around, this little finger, so you see I go in, bring it out, and I go right back in.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, affection, pleasure ecstasy; style: casual, whispered; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.3/10; 7.3s, EN.
EN_B00068_S05977_W000004 · in -19.4 dBFS · gain -0.6 dB · emolia-01554