Manifest tier. emotion, rule B1, T=0.8, step cap 0.25. Population 2,257 chains (30 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 1,798.
Rule.B1 — one-sided: emotion B rises by >=T; the other axis is unconstrained Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.emotion__B1__T0.80__C0.25__INTERNAL — population 2,257 chains (30 h). SHAREABLE variant: 1,798. Filter.rule=='B1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.8 and step_b<=0.25 Sampled from 396 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Contempt ↑ (unconstrained axis: Intoxication Altered States of Consciousness)emotion__B1__T0.80__C0.25__INTERNAL · #1
This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Contempt barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.84.
Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.89 to 0.36 (-0.53), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.24, then +0.22, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 81 s · da · eurospeech
hear it un-normalised (raw levels, max seam 1.1 dB)
k 5d_a -0.528d_b 0.840step_a 0.926step_b 0.238min_cos_consec —min_cos_anchor —dataset eurospeechlang daspeaker denmark_20201M093_2021-04-track denmark_20201M093_2021-04-total 80.9slevel spread 1.1 dBmax seam 1.1 dB
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(normal-paced, some disfluency, average clarity, conversational)Tak for svaret. (low mumble) Vi vil selvfølgelig meget gerne diskutere detaljerne med (ahem) ressortministeren. Når jeg alligevel har bedt om, at det er udenrigsministeren, der står her i dag, er det jo, fordi det er Udenrigsministeriet, der står for
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.5/10; 11.2s, DA.
denmark_20201M093_2021-04-14_1300_634224_645392 · in -20.1 dBFS · gain +0.1 dB · eurospeech-00425
(intoxication altered states of consciousness, concentration·measured, frequent disfluency, somewhat unclear, formal)alt med traktaten og de overdragne beføjelser osv. Og i mit parti er (ahem) vi meget bekymrede for, at EU-retten ligesom æder sig ind på stadig flere (low mumble) områder, som vi jo unægtelig (low mumble) har set det, siden vores medlemskab (low mumble) begyndte.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, concentration; style: formal; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.1/10; 17.4s, DA.
denmark_20201M093_2021-04-14_1300_645392_662800 · in -21.2 dBFS · gain +1.2 dB · eurospeech-00425
(concentration, disappointment, contemplation·normal-paced, some disfluency, average clarity, monologue)Nu er vi så kommet hertil, hvor det handler (low mumble) om forhold for børnene, og jeg er ikke i tvivl om, at der er mange børn rundtomkring – sikkert især i de tidligere Sovjetlande – som har det skidt og i hvert fald betragtelig værre, end man har (low mumble) i Danmark. Men er det en berettigelse til, at man så pludselig kan begynde at lave garantier?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disappointment, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 18.1s, DA.
denmark_20201M093_2021-04-14_1300_662800_680944 · in -21.1 dBFS · gain +1.1 dB · eurospeech-00425
(concentration, triumph, interest· normal-paced, some disfluency, average clarity, monologue)(low mumble) Jeg mindes i hvert fald ikke, at det nogen sinde er blevet forelagt for de danske vælgere, (low mumble) at hvis man stemmer ja til en konkret (ahem) traktat, (low mumble) ja, så får (ahem) EU-Kommissionen beføjelse på børneområdet, socialområdet, eller hvad det nu måtte være af den art.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph, interest; style: monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.0/10; 18.9s, DA.
denmark_20201M093_2021-04-14_1300_680944_699808 · in -20.8 dBFS · gain +0.8 dB · eurospeech-00425
(contempt, concentration, disgust·measured, frequent disfluency, average clarity, formal)beføjelse, for det er kun et koordinerende tiltag. Og det er jo den nederste (low mumble) grad af de beføjelser, som man har overdraget i traktaten, og jeg er helt opmærksom på, at det så ikke er lovgivning, og at det ikke har en juridisk
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, concentration, disgust; style: formal; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 14.7s, DA.
denmark_20201M093_2021-04-14_1300_699808_714496 · in -20.6 dBFS · gain +0.6 dB · eurospeech-00425
This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Fatigue Exhaustion essentially absent — 0.07, lower than 93 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.81.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.95 to 0.73 (-0.21), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.23, then +0.25, then +0.11 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · en · emolia
hear it un-normalised (raw levels, max seam 2.2 dB)
k 5d_a -0.212d_b 0.807step_a 0.265step_b 0.248min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_bxIsVHtMl1ytrack EN_bxIsVHtMl1ytotal 31.8slevel spread 2.5 dBmax seam 2.2 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, slightly dark, balanced body, average recording
(emotional numbness, longing · slow, normally alert, slightly relaxed, didactic)The every, everything's this realm until the on finally callback.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, longing; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.3/10; 5.9s, EN.
EN_bxIsVHtMl1y_W000294 · in -18.2 dBFS · gain -1.8 dB · emolia-01770
(emotional numbness ·measured, normally alert, slightly relaxed, didactic)Uh, (low mumble) and the on-finally callback is that that is just providing a function.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: didactic, formal; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.6/10; 4.9s, EN.
EN_bxIsVHtMl1y_W000295 · in -16.0 dBFS · gain -4.0 dB · emolia-01770
(contemplation, doubt, concentration·slow, very low-energy, relaxed, didactic)So the idea that the on finally callback is determining the promise.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, neutral openness; reads as contemplation, doubt, concentration; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.4/10; 8.0s, EN.
EN_bxIsVHtMl1y_W000296 · in -15.6 dBFS · gain -4.4 dB · emolia-01770
(fear, emotional numbness·measured, normally alert, slightly relaxed, didactic)Which is apparently the current semantics that they're trying to change. It sounds like the current semantics is insane.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, emotional numbness; style: didactic, whispered; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.5/10; 7.5s, EN.
EN_bxIsVHtMl1y_W000297 · in -17.6 dBFS · gain -2.4 dB · emolia-01770
(measured, normally alert, slightly relaxed, casual)And the semantics they're proposing is compellingly the only possible right
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 1.7/10; 4.9s, EN.
EN_bxIsVHtMl1y_W000298 · in -17.9 dBFS · gain -2.1 dB · emolia-01770
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Pride)emotion__B1__T0.80__C0.25__INTERNAL · #3
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.91.
Nothing was asked of the other axis, and in fact Pride drifts down from 0.99 to 0.76 (-0.23), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.22, then +0.23, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 77 s · en · emolia
hear it un-normalised (raw levels, max seam 4.8 dB)
k 5d_a -0.231d_b 0.915step_a 0.463step_b 0.246min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_6GP50-GLO7ktrack EN_6GP50-GLO7ktotal 76.9slevel spread 7.8 dBmax seam 4.8 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged somewhat feminine voice
(pride, contentment, relief · measured, very low-energy, slightly relaxed, whispered)We're working out our unit quantity of a hundred millilitres. So put in a hundred ml up here, highlighted both of those and drag those down. So let's look at how I calculated the price. We click on this cell at the top.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, contentment, relief; style: whispered, didactic; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.3/10; 19.4s, EN.
EN_6GP50-GLO7k_W000011 · in -18.6 dBFS · gain -1.4 dB · emolia-02399
(concentration·slow, very low-energy, slightly relaxed, whispered)E3, we can see the formula up here. What I've done is I've used B3, which is the price currently for this particular item, I've then divided that by C3 which is the amount
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: whispered, didactic; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.8/10; 14.9s, EN.
EN_6GP50-GLO7k_W000012 · in -23.0 dBFS · gain +3.0 dB · emolia-02399
(slow, very low-energy, slightly relaxed, whispered)That we've got and what this does is this calculates the price per millilitre.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly cool, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.2/10; 7.0s, EN.
EN_6GP50-GLO7k_W000013 · in -18.2 dBFS · gain -1.8 dB · emolia-02399
(normal-paced, normally alert, neutral tension, dramatic)So having calculated the price per milliliter, we then multiply that by our new
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: dramatic, conversational; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 3.7/10; 5.5s, EN.
EN_6GP50-GLO7k_W000014 · in -15.1 dBFS · gain -4.9 dB · emolia-02399
(intoxication altered states of consciousness, fatigue exhaustion·measured, very low-energy, slightly relaxed, ASMR)quantity, which is a hundred millilitres in F3 and that will give us the price per hundred mils. Having put that formula in this top cell here, we can just drag, click on that, click on the corner. Whoops, didn't mean to do that. Just click back. Right, just click on the corner here and drag that down. And what that does is that puts the formula in
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, fatigue exhaustion; style: ASMR, whispered; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.7/10; 29.5s, EN.
EN_6GP50-GLO7k_W000015 · in -18.6 dBFS · gain -1.4 dB · emolia-02399
Disgust ↑ (unconstrained axis: Intoxication Altered States of Consciousness)emotion__B1__T0.80__C0.25__INTERNAL · #4
This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Disgust barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.82.
Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.66 to 0.58 (-0.08), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.21, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 23 s · snippets
hear it un-normalised (raw levels, max seam 1.9 dB)
k 5d_a -0.080d_b 0.822step_a 0.138step_b 0.222min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch84_part0_batch84_parttrack batch84_part0_batch84_parttotal 23.2slevel spread 2.9 dBmax seam 1.9 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, fairly steady, no disfluency, narration)the gang has also made close ties with cartels in South America.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.2/10; 3.6s.
batch84_part0_batch84_part0_chunk_1756_1_1623773 · in -25.7 dBFS · gain +5.7 dB · snippets-01323
(normal-paced, fairly steady, little disfluency, casual)The gang is the second largest mafia in Italy and is certainly amongst the richest organizations in the world.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.3/10; 6.2s.
batch84_part0_batch84_part0_chunk_1756_1_1623784 · in -23.8 dBFS · gain +3.8 dB · snippets-01323
(relief, astonishment surprise, thankfulness gratitude· normal-paced, fairly steady, almost no disfluency, monologue)It's been so beneficial that the family rakes in almost five billion dollars a year in revenue.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, astonishment surprise, thankfulness gratitude; style: monologue, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.5/10; 5.2s.
batch84_part0_batch84_part0_chunk_1756_1_1623880 · in -23.7 dBFS · gain +3.7 dB · snippets-01323
(measured, fairly steady, no disfluency, formal)the naming rights to ten NFL stadiums.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.9/10; 3.2s.
batch84_part0_batch84_part0_chunk_1756_1_1623915 · in -24.3 dBFS · gain +4.3 dB · snippets-01323
(disgust, emotional numbness, distress·normal-paced, steady, no disfluency, formal)prostitution, arms trafficking, money laundering and loan sharking.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, emotional numbness, distress; style: formal, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.0/10; 4.3s.
batch84_part0_batch84_part0_chunk_1756_1_1623960 · in -22.8 dBFS · gain +2.8 dB · snippets-01323
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Sourness)emotion__B1__T0.80__C0.25__INTERNAL · #5
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.85.
Nothing was asked of the other axis, and in fact Sourness drifts down from 1.00 to 0.60 (-0.40), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.15, then +0.24, then +0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.52 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 59 s · en · podcast
hear it un-normalised (raw levels, max seam 2.2 dB)
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, light breath
(sourness, bitterness, impatience and irritability · normal-paced, energised, neutral tension, casual)together. No, they really suck. Like (ahem) if name one thing the American military accomplished. They they didn't they didn't stop the Holocaust. The Soviets did that, and that's the biggest thing they take credit for. I mean, sure and then they dropped some fucking nukes on Japan and the people are still suffering
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as sourness, bitterness, impatience and irritability; style: casual, playful; average recording, some background noise; genuineness 6.0/6; vocal-burst blend 7.1/10; 17.2s, EN.
511300_00286720 · in -17.7 dBFS · gain -2.3 dB · podcast-02322
(disgust, anger, contempt· normal-paced, normally alert, neutral tension, conversational)from that shit. And we wanna we wanna go and say that this building that fell over is the biggest act of terror act of terrorism. Like we invented a weapon of mass destruction. We
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as disgust, anger, contempt; style: conversational, casual; average recording, quiet background; genuineness 5.9/6; vocal-burst blend 5.1/10; 9.6s, EN.
511300_00288448 · in -17.9 dBFS · gain -2.1 dB · podcast-02319
(embarrassment· normal-paced, normally alert, relaxed, casual)were the first ones to do it. Technically they were technically they already existed in Germany, but we're talking, you know, the the sarin gas that he had, it wouldn't it couldn't do what that nuke did. And after
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 3.3/10; 11.9s, EN.
511300_00289400 · in -19.6 dBFS · gain -0.4 dB · podcast-02323
(amusement· normal-paced, normally alert, slightly relaxed, casual)the after we dropped those nukes, we started limiting things like sarin gas,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as amusement; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.5/10; 3.5s, EN.
511300_00290584 · in -19.8 dBFS · gain -0.2 dB · podcast-02326
(intoxication altered states of consciousness, amusement, embarrassment·measured, normally alert, relaxed, casual)Syria, it happened. (low mumble) Um yeah, I mean (low mumble) the US military. They give some cool planes though. Oh also did you know did you know that the they they suspected that bin Laden was gonna attack on the
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, amusement, embarrassment; style: casual, playful; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 1.7/10; 16.2s, EN.
511300_00291200 · in -17.6 dBFS · gain -2.4 dB · podcast-02326
This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Fatigue Exhaustion barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.83.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.90 to 0.28 (-0.61), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.21, then +0.20, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.09 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.09 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 48 s · en · emolia
k 5d_a -0.613d_b 0.826step_a 0.564step_b 0.213min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_t3zsbnPp3Bctrack EN_t3zsbnPp3Bctotal 48.2slevel spread 3.6 dBmax seam 2.2 dB
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, fairly steady
(normal-paced, normally alert, slightly relaxed, monologue)Is technology creating new spaces for work and changing others that (low mumble) are being (ahem) left redundant from the private sector perspective?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.7/10; 9.4s, EN.
EN_t3zsbnPp3Bc_W000330 · in -22.4 dBFS · gain +2.4 dB · emolia-01571
(slow, subdued, slightly relaxed, monologue)Yeah, (low mumble) the future is globalization. (low mumble) Uh, technology is enabling us, uh, (low mumble) to leverage cross-border trade, uh, (low mumble) drive innovation. (low mumble) Uhm, so it's, it's, (low mumble) uhm,
full caption & clip details
An adult masculine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_t3zsbnPp3Bc_W000331 · in -20.4 dBFS · gain +0.4 dB · emolia-01571
(interest, concentration·normal-paced, normally alert, slightly relaxed, monologue)In some ways it's creating jobs, I know it's taking away jobs, but it's creating different types of jobs. So it reverts back to the big question on education. How do we train the workforce to participate in more of a globalization (ahem) network?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: monologue, formal; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.0/10; 12.4s, EN.
EN_t3zsbnPp3Bc_W000332 · in -20.7 dBFS · gain +0.7 dB · emolia-01571
(intoxication altered states of consciousness, contemplation·measured, subdued, relaxed, casual)My view is, is, is similar. I think, you know, more and more we are a service economy. And, (low mumble) uh, you know, we kind of, you know, say that data is the, is the new wild, so.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, submissive, neutral openness; reads as intoxication altered states of consciousness, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 4.8/10; 10.9s, EN.
EN_t3zsbnPp3Bc_W000333 · in -21.1 dBFS · gain +1.1 dB · emolia-01571
(fatigue exhaustion, intoxication altered states of consciousness · measured, subdued, slightly relaxed, casual)That needs to kind of, you know, really, (low mumble) uh, keep going.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.1/10; 3.5s, EN.
EN_t3zsbnPp3Bc_W000334 · in -18.8 dBFS · gain -1.2 dB · emolia-01571
This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Contempt barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.81.
Nothing was asked of the other axis, and in fact Affection drifts down from 0.82 to 0.60 (-0.22), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.21, then +0.22, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.10 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, slightly relaxed, fairly steady
(normal-paced, normally alert, frequent disfluency, conversational)(wistful sigh) Uh we have special handling of cats, we use (low mumble) uh local anesthetics to draw the blood from them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: conversational, ASMR; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.4/10; 6.7s, EN.
886199_00010612 · in -20.5 dBFS · gain +0.5 dB · podcast-05367
(measured, normally alert, no disfluency, monologue)So we are using the most modern ways how to treat cats.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: monologue, ASMR; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 5.4/10; 4.2s, EN.
886199_00011756 · in -20.5 dBFS · gain +0.5 dB · podcast-05359
(normal-paced, normally alert, some disfluency, casual)Okay. And do they they come through the same entrance but they just go into a different place?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.6/6; vocal-burst blend 2.1/10; 4.1s, EN.
886199_00012808 · in -20.4 dBFS · gain +0.4 dB · podcast-04643
(normal-paced, normally alert, frequent disfluency, casual)exactly. So are you saying that cats are so sensitive to the smell that if they smell a dog that then they will become much more nervous. So this is the reason why you have that separated area so they feel more comfortable.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 1.2/10; 13.4s, EN.
886199_00013312 · in -20.6 dBFS · gain +0.6 dB · podcast-05360
(contempt, emotional numbness· normal-paced, very low-energy, frequent disfluency, casual)Yeah, yeah. Uh (low mumble) it's not just the smell, it's also vision, see the dog or hear the dog. So they don't see and hear the dogs when they are waiting, because there is completely separated area. Okay,
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as contempt, emotional numbness; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 1.9/10; 14.0s, EN.
886199_00014688 · in -20.7 dBFS · gain +0.7 dB · podcast-05363
This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Astonishment Surprise barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.85.
Nothing was asked of the other axis, and in fact Interest drifts down from 0.98 to 0.66 (-0.32), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.16, then +0.22, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.23 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.25 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.23, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, average clarity, light breath
(interest, contemplation, doubt · measured, subdued, relaxed, casual)harnessing that power, trying to capture that lightning in a bottle. And I will say this if they're able to do it with a couple of other shows that are kind of ripe for the picking,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as interest, contemplation, doubt; style: casual, conversational; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 4.9/10; 13.5s, EN.
183770_00105768 · in -18.6 dBFS · gain -1.4 dB · podcast-02784
(measured, normally alert, slightly relaxed, whispered)we're gonna be in a good spot. Okay. So I I will say the possibility of us getting a continuation of Spider-Man.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.6/10; 9.7s, EN.
183770_00107352 · in -18.0 dBFS · gain -2.0 dB · podcast-02781
(normal-paced, normally alert, slightly relaxed, casual)Spider-Man 98 has been thrown around. Okay, and we have seen our friend Spider Guy.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.2/10; 5.1s, EN.
183770_00108480 · in -17.4 dBFS · gain -2.6 dB · podcast-02768
(fatigue exhaustion· normal-paced, normally alert, slightly relaxed, casual)(ahem) um, and in in the the final final season here again, spoiler alert. In in the final episode, rather, we saw Peter Parker and Mary Jane.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion; style: casual, conversational; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.2/10; 9.0s, EN.
183770_00109040 · in -17.3 dBFS · gain -2.7 dB · podcast-02813
(astonishment surprise, relief, triumph· normal-paced, normally alert, neutral tension, casual)I did, and then I we watched (ahem) uh a breakdown video, and I was like, Oh, that's why they look familiar. That's okay. Now I got it. There it is.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, relief, triumph; style: casual, conversational; good recording, quiet background; genuineness 5.0/6; vocal-burst blend 3.6/10; 7.2s, EN.
183770_00110288 · in -20.1 dBFS · gain +0.1 dB · podcast-02791
Fear ↑ (unconstrained axis: Embarrassment)emotion__B1__T0.80__C0.25__INTERNAL · #9
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Fear barely there — 0.17, lower than 83 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.83.
Nothing was asked of the other axis, and in fact Embarrassment barely moves at all, sitting near 0.97 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.19, then +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 49 s · en · emolia
k 5d_a 0.018d_b 0.833step_a 0.019step_b 0.229min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00055_S09869track EN_B00055_S09869total 49.0slevel spread 1.8 dBmax seam 0.8 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, moderately variable, some disfluency, wide pitch range
(embarrassment, shame, affection · brisk, normally alert, slightly relaxed, conversational)I'm like, are we good? 100%. I think it's important in any relationship to have arguments and conflict, but I came from a few friendships that...
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, shame, affection; style: conversational, casual; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 7.6/10; 8.2s, EN.
EN_B00055_S09869_W000006 · in -20.8 dBFS · gain +0.8 dB · emolia-01313
(embarrassment, contemplation, longing· brisk, energised, neutral tension, casual)I couldn't argue in them, I couldn't really speak my mind, so I remember when you and I became friends, you were kind of like a bad atta. How? Can we say that?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, contemplation, longing; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 7.1/10; 8.4s, EN.
EN_B00055_S09869_W000007 · in -20.3 dBFS · gain +0.3 dB · emolia-01313
(embarrassment, pleasure ecstasy, thankfulness gratitude· brisk, energised, neutral tension, casual)You're very vocal in how you're feeling. And so, and I, I told you, I came from these friendships where I felt like I couldn't really speak up. They didn't really, it was just a weird thing where we couldn't really have open communication like that. So when I came into our friendship, do you remember the first
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, slightly vulnerable; reads as embarrassment, pleasure ecstasy, thankfulness gratitude; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 9.6/10; 13.6s, EN.
EN_B00055_S09869_W000008 · in -19.9 dBFS · gain -0.1 dB · emolia-01313
(elation, relief, pleasure ecstasy · brisk, energised, neutral tension, casual)Little thing that we got into the first argument. I thought, I thought, I thought it was over for us. I thought it was way before Girls Gone Bible.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as elation, relief, pleasure ecstasy; style: casual, playful; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.7/10; 8.8s, EN.
EN_B00055_S09869_W000009 · in -19.0 dBFS · gain -1.0 dB · emolia-01313
(fear, infatuation, distress·normal-paced, energised, neutral tension, casual)I literally start bawling my eyes out because you said something like, I'm concerned for this friend. You said something like that and it triggered me so hard. I thought you were breaking up with me.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly vulnerable; reads as fear, infatuation, distress; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.9/10; 9.4s, EN.
EN_B00055_S09869_W000010 · in -19.6 dBFS · gain -0.5 dB · emolia-01313
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Concentration barely there — 0.15, lower than 85 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.83.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 1.00 to 0.82 (-0.18), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.19, then +0.22, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.12 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.12 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 76 s · en · emolia
k 5d_a -0.181d_b 0.828step_a 0.547step_b 0.234min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN__SMs5f7SbzQtrack EN__SMs5f7SbzQtotal 75.5slevel spread 2.7 dBmax seam 2.3 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, slightly relaxed, fairly steady
(hope enthusiasm optimism, contentment, thankfulness gratitude · brisk, normally alert, some disfluency, casual)Hello and welcome back to a special episode of Community Hotline. My name is James Ofsink and we're here today to talk about the Gresham Safety Measure. (low mumble) Uh, we're fortunate to be joined by Gresham City Manager Nina Vetter and Gresham Fire Chief Scott Lewis. Thank you so much both for being here. Thank you.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, minimal breath; affect is positive, slightly submissive, slightly guarded; reads as hope enthusiasm optimism, contentment, thankfulness gratitude; style: casual, conversational; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 3.4/10; 16.4s, EN.
EN__SMs5f7SbzQ_W000000 · in -15.7 dBFS · gain -4.3 dB · emolia-01709
(doubt·normal-paced, normally alert, some disfluency, casual)Nina, can you tell us a little bit about the primary goals of the upcoming levy?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.9/10; 4.4s, EN.
EN__SMs5f7SbzQ_W000001 · in -13.3 dBFS · gain -6.7 dB · emolia-01709
(hope enthusiasm optimism, contentment, thankfulness gratitude· normal-paced, very low-energy, some disfluency, casual)Yeah, absolutely. So our Gresham safety levy that's on the ballot in May of this year, (low mumble) uhm, is a five-year operating levy that would fund our critical safety services at an average cost of about $28 (ahem) per household. (ahem) Uhm, so our community safety levy is really focused on a few key areas. Police, fire, homelessness response, and mental health response.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, minimal breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, contentment, thankfulness gratitude; style: casual, monologue; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 1.7/10; 24.7s, EN.
EN__SMs5f7SbzQ_W000002 · in -15.5 dBFS · gain -4.5 dB · emolia-01709
(disappointment, helplessness, sadness· normal-paced, normally alert, some disfluency, monologue)(ahem) Since the city has been in a budget crisis for the last few decades, (ahem) we have not been able to provide the community really with the services (ahem) that they need, that they demand, and the services that are needed in these changing times (ahem) as safety has evolved. So the Gresham Safety Levy would allow us to
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, helplessness, sadness; style: monologue, casual; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.6/10; 19.0s, EN.
EN__SMs5f7SbzQ_W000003 · in -16.1 dBFS · gain -3.9 dB · emolia-01709
(concentration, pain·brisk, normally alert, almost no disfluency, formal)Stabilize our current services that we do provide in safety and add strategically positions across those critical safety areas so that we can be more responsive and we can be proactive.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, pain; style: formal, dramatic; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 10.4s, EN.
EN__SMs5f7SbzQ_W000004 · in -15.2 dBFS · gain -4.8 dB · emolia-01709
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Relief barely there — 0.15, lower than 85 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.82.
Nothing was asked of the other axis, and in fact Triumph drifts down from 0.92 to 0.20 (-0.72), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.23, then +0.18, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.75 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.75 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 61 s · zh · emolia
k 5d_a -0.725d_b 0.821step_a 0.837step_b 0.233min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00047_S01791track ZH_B00047_S01791total 61.1slevel spread 7.5 dBmax seam 7.5 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · average recording, normally alert, moderately variable
(triumph · measured, slightly relaxed, no disfluency, storytelling)我拿着放着笔的墨水瓶,欢快的哼着歌,一蹦一跳的跑过来。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as triumph; style: storytelling, cartoonish; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.0/10; 6.1s, ZH.
ZH_B00047_S01791_W000002 · in -30.1 dBFS · gain +10.1 dB · emolia-03743
(measured, slightly relaxed, some disfluency, cartoonish)爸爸双手抱住脑袋,沮丧的尖叫着,然后气冲冲的离开了。我知道爸爸去找罐子了怎么办?又要挨打了,我得想办法补救一下。有啦。
full caption & clip details
A child strongly feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, bright, fairly smooth, slightly thin; clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; no dominant emotion; style: cartoonish, storytelling; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 4.7/10; 18.1s, ZH.
ZH_B00047_S01791_W000003 · in -29.7 dBFS · gain +9.7 dB · emolia-03743
(intoxication altered states of consciousness· measured, slightly relaxed, some disfluency, storytelling)我把墨水瓶放在一边,拿着笔开始在弄脏的地方画画。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, bright, fairly smooth, balanced body; slurred, some disfluency, very wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness; style: storytelling, playful; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.4/10; 5.8s, ZH.
ZH_B00047_S01791_W000004 · in -32.3 dBFS · gain +12.3 dB · emolia-03743
(astonishment surprise, doubt, contempt· measured, neutral tension, some disfluency, cartoonish)臭小子又闯祸,看我怎么收拾你。他举起棍子,正准备打我,却看到了我画的小猴子,忍不住夸道。咦画的不错嘛,爸爸自己也手痒痒了,说看着儿子老爸给你录一首。
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, doubt, contempt; style: cartoonish, storytelling; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.9/10; 23.4s, ZH.
ZH_B00047_S01791_W000005 · in -24.8 dBFS · gain +4.8 dB · emolia-03743
(relief·normal-paced, slightly relaxed, some disfluency, storytelling)说完,他拿起墨水瓶浇洒、墨汁地毯,中间出现了一条蛇。
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is cool, bright, fairly smooth, balanced body; crisply articulate, some disfluency, very wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as relief; style: storytelling, cartoonish; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.1/10; 7.2s, ZH.
ZH_B00047_S01791_W000006 · in -31.6 dBFS · gain +11.6 dB · emolia-03743
This chain comes from the one-sided rule: only Triumph had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Triumph barely there — 0.08, lower than 92 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.82.
Nothing was asked of the other axis, and in fact Sadness drifts down from 0.98 to 0.43 (-0.55), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.24, then +0.20, then +0.15 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 29 s · snippets
k 5d_a -0.547d_b 0.823step_a 0.547step_b 0.244min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch260_part3_batch260_patrack batch260_part3_batch260_patotal 28.7slevel spread 17.8 dBmax seam 17.8 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, moderate pitch range
(sadness, distress, disappointment · brisk, fairly steady, no disfluency, storytelling)This was disappointing as they were putting a lot of effort into finding who had killed Jemima.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sadness, distress, disappointment; style: storytelling, whispered; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.3/10; 5.0s.
batch260_part3_batch260_part3_chunk_808_1_787811 · in -24.5 dBFS · gain +4.5 dB · snippets-00837
(normal-paced, fairly steady, no disfluency, narration)Last month's bonus episode focused on the deep freeze murder of Anne Noblet.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: narration, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.6/10; 4.6s.
batch260_part3_batch260_part3_chunk_808_1_787866 · in -26.2 dBFS · gain +6.2 dB · snippets-00837
(measured, steady, almost no disfluency, narration)they're being used together, for example, using music therapy alongside of chemotherapy for a child with acute lymphocytic leukemia.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.9/10; 7.0s.
batch260_part3_batch260_part3_chunk_808_1_788156 · in -34.6 dBFS · gain +14.6 dB · snippets-00837
(awe, sourness, emotional numbness· measured, fairly steady, no disfluency, ASMR)you'd say it's a relative who would claim to be a spirit guide or an ascended master.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, sourness, emotional numbness; style: ASMR, casual; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 2.3/10; 4.7s.
batch260_part3_batch260_part3_chunk_808_1_788226 · in -16.8 dBFS · gain -3.2 dB · snippets-00837
(triumph·fast, fairly steady, some disfluency, casual)you know, but in the case of that boy, I mean we grew up together in a in a in an area in Abeokuta. So, I know him.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as triumph; style: casual, storytelling; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 6.5/10; 6.8s.
batch260_part3_batch260_part3_chunk_808_1_788412 · in -27.7 dBFS · gain +7.7 dB · snippets-00837
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.86.
Nothing was asked of the other axis, and in fact Elation drifts down from 0.95 to 0.13 (-0.82), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.22, then +0.21, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 54 s · en · emolia
k 5d_a -0.824d_b 0.858step_a 0.830step_b 0.230min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_-UiWTavWGiYtrack EN_-UiWTavWGiYtotal 54.4slevel spread 3.8 dBmax seam 2.4 dB
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, some disfluency, light breath
(elation, hope enthusiasm optimism · normal-paced, neutral tension, moderately variable, casual)(low mumble) Uh, it kept getting pushed further and further out, but we are very happy to announce that this is now totally live and up and anyone can go and download it. So the idea is this is not a research paper, this is just kind of a technical how-to of
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 6.6/10; 12.6s, EN.
EN_-UiWTavWGiY_W000022 · in -22.1 dBFS · gain +2.1 dB · emolia-02480
(relief, hope enthusiasm optimism, contentment· normal-paced, slightly relaxed, fairly steady, casual)(ahem) Uhm, talk to me about this, the paper is up now, please go download it, (low mumble) uhm, get in contact with us, share your thoughts, questions, anything like that. So, this is just kind of an aside, but we're really happy to announce that we pushed to get it ready for BroCon and it is out and available now.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, hope enthusiasm optimism, contentment; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.0/10; 13.8s, EN.
EN_-UiWTavWGiY_W000024 · in -22.2 dBFS · gain +2.2 dB · emolia-02480
(confusion, embarrassment· normal-paced, slightly relaxed, fairly steady, casual)So, how do we do IR with Bro? (low mumble) Uhm, we are kind of an old school team. We don't have a sim really, except for Gmail. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as confusion, embarrassment; style: casual, playful; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.7/10; 7.5s, EN.
EN_-UiWTavWGiY_W000025 · in -19.9 dBFS · gain -0.1 dB · emolia-02480
(triumph, longing, elation· normal-paced, neutral tension, fairly steady, casual)We've, we've messed around with a lot of different things over the years. We've, we've had Splunk, we've had, uh, (low mumble) you know, the, the Elk stack, we've messed around with that, we've been looking at Spark, we've been looking at all this stuff, but ultimately our team often gets back to,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, longing, elation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 6.8/10; 13.0s, EN.
EN_-UiWTavWGiY_W000026 · in -20.9 dBFS · gain +0.9 dB · emolia-02480
(emotional numbness·brisk, slightly relaxed, fairly steady, monologue)And we pull all of our logs from most of our bros into like a central log repository and we, and then we've got some high powered crunching boxes to do that. So.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, casual; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.0/10; 6.7s, EN.
EN_-UiWTavWGiY_W000028 · in -18.5 dBFS · gain -1.5 dB · emolia-02480
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.84.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.89 to 0.38 (-0.51), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.23, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.77 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.77 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · en · emolia
k 5d_a -0.509d_b 0.838step_a 0.685step_b 0.246min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00026_S03384track EN_B00026_S03384total 35.8slevel spread 0.9 dBmax seam 0.8 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, energised, slightly relaxed, clear
(measured, fairly steady, almost no disfluency, narration)In the summertime, ducks and geese migrate to the Arctic to build nests and raise their young. So if plants and animals can survive in the far north, what about people?
full caption & clip details
A child masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 12.1s, EN.
EN_B00026_S03384_W000023 · in -17.3 dBFS · gain -2.7 dB · emolia-00752
(fear, doubt, distress·normal-paced, fairly steady, no disfluency, playful)How would you stay warm during the cold, dark winters?
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, doubt, distress; style: playful, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.7/10; 3.6s, EN.
EN_B00026_S03384_W000024 · in -17.0 dBFS · gain -3.0 dB · emolia-00752
(fear, awe, doubt ·measured, moderately variable, almost no disfluency, narration)How would you stay protected from the icy winds and snow storms? How would you find food?
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, awe, doubt; style: narration, storytelling; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.8/10; 6.2s, EN.
EN_B00026_S03384_W000025 · in -17.9 dBFS · gain -2.1 dB · emolia-00752
(normal-paced, fairly steady, almost no disfluency, narration)Before there were stores, fancy jackets, or electricity, these people survived in the frozen north.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 7.8s, EN.
EN_B00026_S03384_W000026 · in -17.9 dBFS · gain -2.1 dB · emolia-00752
(emotional numbness·measured, fairly steady, almost no disfluency, narration)They built houses from driftwood, earth, whale bones, and snow.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness; style: narration, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.4/10; 5.5s, EN.
EN_B00026_S03384_W000027 · in -18.0 dBFS · gain -2.0 dB · emolia-00752
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Emotional Numbness essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.89.
Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.84 to 0.14 (-0.71), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.25, then +0.23, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 31 s · snippets
k 5d_a -0.706d_b 0.890step_a 0.481step_b 0.246min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch2_part3_batch2_part3_track batch2_part3_batch2_part3_total 31.1slevel spread 1.4 dBmax seam 1.1 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth
(brisk, energised, neutral tension, storytelling)That driver, notably the only person on the bus, wearing a seatbelt.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: storytelling, dramatic; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 5.0/10; 4.5s.
batch2_part3_batch2_part3_chunk_100_1_87887 · in -26.8 dBFS · gain +6.8 dB · snippets-01045
(measured, normally alert, slightly relaxed, monologue)We've got to keep the students properly restrained in the event.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.7/10; 4.4s.
batch2_part3_batch2_part3_chunk_100_1_88004 · in -27.1 dBFS · gain +7.1 dB · snippets-01045
(brisk, normally alert, slightly relaxed, formal)While the National Highway Traffic Safety Administration says school buses are the most regulated vehicles on the road.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, dramatic; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.9/10; 6.0s.
batch2_part3_batch2_part3_chunk_100_1_88092 · in -26.0 dBFS · gain +6.0 dB · snippets-01045
(fatigue exhaustion·normal-paced, normally alert, slightly relaxed, monologue)Yeah with that I've daily network stats in the last 24 hours there are 648 transactions on the Swire network.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion; style: monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.6/10; 5.7s.
batch2_part3_batch2_part3_chunk_100_1_88203 · in -26.6 dBFS · gain +6.6 dB · snippets-01045
(emotional numbness, doubt, pain·measured, subdued, slightly relaxed, whispered)use a country's currency like that. (low mumble) Um, then you do have uh, (low mumble) you have the problem that they may for political reasons, (low mumble) um, just
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness, doubt, pain; style: whispered, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.4/10; 9.8s.
batch2_part3_batch2_part3_chunk_100_1_88248 · in -27.4 dBFS · gain +7.4 dB · snippets-01045
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Emotional Numbness essentially absent — 0.01, lower than 99 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.80.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.82 to 0.67 (-0.15), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.22, then +0.16, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.65 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.65 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 55 s · en · emolia
k 5d_a -0.150d_b 0.805step_a 0.333step_b 0.217min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_oLoQX8RA6tItrack EN_oLoQX8RA6tItotal 55.5slevel spread 3.3 dBmax seam 2.1 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, moderate pitch range
(normal-paced, fairly steady, little disfluency, monologue)The next most important step is the Recode, or what in-game is called Accuracy.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.4/10; 4.4s, EN.
EN_oLoQX8RA6tI_W000022 · in -19.1 dBFS · gain -0.9 dB · emolia-00711
(normal-paced, fairly steady, almost no disfluency, monologue)As can be seen on screen, it has high recoil. It starts with vertical kick, then kicks to the right, returns to the center, only to kick upwards again.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.2/10; 7.8s, EN.
EN_oLoQX8RA6tI_W000023 · in -20.0 dBFS · gain +0.0 dB · emolia-00711
(concentration, triumph· normal-paced, fairly steady, some disfluency, monologue)Knowing this pattern makes predicting the recoil easier. Obviously using a grip improves the accuracy, but on this weapon it hardly reduces the recoil. You can see that its vertical kick is a little less, but it's not significant enough to make it a viable attachment. This means we're better off selecting another attachment.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph; style: monologue, formal; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 2.7/10; 18.3s, EN.
EN_oLoQX8RA6tI_W000024 · in -18.4 dBFS · gain -1.6 dB · emolia-00711
(concentration, interest, contentment·measured, steady, some disfluency, monologue)The hipfire spread isn't something to brag about either. We can't put it into a number, but on screen you can see the spread while standing, crouching and proning. While average for the assault rifles, it's smarter to aim down the sights. Keep in mind that this can change depending on our stance, as we can see the spread getting smaller when crouching or proning.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest, contentment; style: monologue, formal; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.2/10; 19.2s, EN.
EN_oLoQX8RA6tI_W000025 · in -18.8 dBFS · gain -1.2 dB · emolia-00711
(normal-paced, steady, almost no disfluency, formal)Movement speed with the Rampart 17 is similar to the other assault rifles at 90%.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.0/10; 5.2s, EN.
EN_oLoQX8RA6tI_W000026 · in -16.7 dBFS · gain -3.3 dB · emolia-00711
This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Anger essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.82.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism barely moves at all, sitting near 0.88 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.20, then +0.19, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 44 s · de · emolia
k 5d_a -0.048d_b 0.816step_a 0.629step_b 0.232min_cos_consec —min_cos_anchor —dataset emolialang despeaker DE_FSY0BnnloUAtrack DE_FSY0BnnloUAtotal 43.9slevel spread 1.3 dBmax seam 1.3 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, some disfluency, average clarity
(brisk, moderately variable, wide pitch range, playful)Bei der Körperhaltung auf dem Motorrad gibt es so die ein oder anderen Missverständnisse, wie wir vielleicht in dem letzten Video schon erfahren haben. Da haben wir nämlich über die sogenannte Faustregel gesprochen.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, conversational; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.4/10; 11.6s, DE.
DE_FSY0BnnloUA_W000000 · in -12.6 dBFS · gain -7.4 dB · emolia-00151
(confusion, thankfulness gratitude·normal-paced, fairly steady, moderate pitch range, casual)Und heute widmen wir uns mal bisschen den Oberkörper und was es da so für, (low mumble) naja, ich sag mal jetzt nicht unbedingt Missverständnisse, aber ein paar so.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion, thankfulness gratitude; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.5/10; 9.8s, DE.
DE_FSY0BnnloUA_W000001 · in -12.3 dBFS · gain -7.7 dB · emolia-00151
(intoxication altered states of consciousness· normal-paced, fairly steady, moderate pitch range, casual)Falschmeinungen im Internet herumkursieren. Schaut es euch an. Ich bin gespannt, ob es der ein oder andere schon wusste. Bis gleich.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.1/10; 7.0s, DE.
DE_FSY0BnnloUA_W000002 · in -13.6 dBFS · gain -6.4 dB · emolia-00151
(fear, confusion· normal-paced, fairly steady, moderate pitch range, casual)Man muss das ganze Thema Körperhaltung und allgemein die Haltung am Motorrad in mehrere Sequenzen aufschlüsseln, weil ich glaube nicht, dass es möglich ist, das Ganze in einem
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, confusion; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.7/10; 8.7s, DE.
DE_FSY0BnnloUA_W000003 · in -12.6 dBFS · gain -7.4 dB · emolia-00151
(brisk, fairly steady, moderate pitch range, dramatic)einigermaßen beizubringen. Und deswegen schlüssel ich das Thema so auf und deswegen unterhalten wir uns heute mal über den Oberkörper.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, conversational; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.9/10; 6.2s, DE.
DE_FSY0BnnloUA_W000004 · in -13.0 dBFS · gain -7.0 dB · emolia-00151
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Relief barely there — 0.19, lower than 81 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.80.
Nothing was asked of the other axis, and in fact Confusion drifts down from 0.98 to 0.77 (-0.21), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.10, then +0.23, then +0.23 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.03 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.03 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · en · emolia
k 5d_a -0.215d_b 0.803step_a 0.641step_b 0.244min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_ex5q3QUL95otrack EN_ex5q3QUL95ototal 39.8slevel spread 4.2 dBmax seam 2.9 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice
(confusion, awe, contemplation · measured, normally alert, slightly relaxed, whispered)If someone says the Buddha has spoken spiritual truths, he slanders the Buddha due to his inability to understand what the Buddha teaches. Subhuti, as to speaking truth, no truth can be spoken, therefore it is called speaking truth.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as confusion, awe, contemplation; style: whispered, narration; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.9/10; 16.1s, EN.
EN_ex5q3QUL95o_W000278 · in -19.9 dBFS · gain -0.1 dB · emolia-02457
(measured, very low-energy, slightly relaxed, whispered)At that time, Subhuti, the wise elder, addressed the Buddha.
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 3.0/10; 4.5s, EN.
EN_ex5q3QUL95o_W000279 · in -18.7 dBFS · gain -1.3 dB · emolia-02457
(doubt, fear· measured, normally alert, slightly relaxed, whispered)Will there be living beings in the future who believe in this Sutra when they hear it?
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, dark, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt, fear; style: whispered, narration; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.3/10; 4.4s, EN.
EN_ex5q3QUL95o_W000280 · in -15.7 dBFS · gain -4.3 dB · emolia-02457
(contemplation, awe, infatuation·slow, very low-energy, relaxed, whispered)The living beings to whom you refer Are neither living beings nor not living beings. Why, Subuti.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contemplation, awe, infatuation; style: whispered, storytelling; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.5/10; 7.9s, EN.
EN_ex5q3QUL95o_W000282 · in -18.5 dBFS · gain -1.5 dB · emolia-02457
(relief, thankfulness gratitude, awe ·measured, normally alert, slightly relaxed, whispered)Blessed Lord, when you attained complete enlightenment, did you feel in your mind that nothing had been acquired?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, thankfulness gratitude, awe; style: whispered, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.5/10; 6.2s, EN.
EN_ex5q3QUL95o_W000285 · in -19.3 dBFS · gain -0.7 dB · emolia-02457
This chain comes from the one-sided rule: only Triumph had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Triumph essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.82.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.98 to 0.19 (-0.79), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.20, then +0.19, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · en · emolia
k 5d_a -0.789d_b 0.817step_a 0.888step_b 0.215min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00027_S03560track EN_B00027_S03560total 31.5slevel spread 4.0 dBmax seam 3.4 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a child somewhat masculine voice
(fatigue exhaustion · normal-paced, energised, neutral tension, casual)And if we have more meeples than five on our card, we have to give it back in the middle and get points for the meeple.
full caption & clip details
A child somewhat masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fatigue exhaustion; style: casual, playful; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.8/10; 6.9s, EN.
EN_B00027_S03560_W000005 · in -19.2 dBFS · gain -0.8 dB · emolia-00764
(intoxication altered states of consciousness·measured, normally alert, fully relaxed, casual)We take, if we take a card, we (low mumble) uhm, push the line in front and get a new card.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, fully relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; slurred, frequent disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness; style: casual, playful; below-average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 7.1s, EN.
EN_B00027_S03560_W000006 · in -20.2 dBFS · gain +0.2 dB · emolia-00764
(normal-paced, normally alert, fully relaxed, casual)So the (ahem) (low mumble) expensive card is only every time the last one.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.6/10; 5.2s, EN.
EN_B00027_S03560_W000007 · in -23.2 dBFS · gain +3.2 dB · emolia-00764
(normal-paced, energised, slightly relaxed, playful)We take a card and we have, (ahem) uhm, seven different characters and seven different buildings.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, casual; average recording, some background noise; genuineness 2.2/6; vocal-burst blend 1.0/10; 6.1s, EN.
EN_B00027_S03560_W000008 · in -19.9 dBFS · gain -0.1 dB · emolia-00764
(normal-paced, normally alert, slightly relaxed, casual)And for each building we have one character. So we take a character and...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.1/10; 5.6s, EN.
EN_B00027_S03560_W000009 · in -20.6 dBFS · gain +0.6 dB · emolia-00764
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.80. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude barely there — 0.17, lower than 83 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.83.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.93 to 0.16 (-0.77), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.24, then +0.23, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.24 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.24 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 27 s · en · emolia
k 5d_a -0.771d_b 0.830step_a 0.543step_b 0.245min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_tAW9zt_jqmutrack EN_tAW9zt_jqmutotal 26.7slevel spread 5.1 dBmax seam 5.1 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(concentration, interest · normal-paced, some disfluency, average clarity, casual)(low mumble) Uhm, we're planning, this is, (ahem) uh, you know, kind of a quasi longitudinal effort, so we'll try to keep a certain number of the fields we used in the original study stable.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, interest; style: casual; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.4/10; 9.3s, EN.
EN_tAW9zt_jqmu_W000154 · in -17.7 dBFS · gain -2.3 dB · emolia-02308
(normal-paced, frequent disfluency, average clarity, casual)But, (ahem) uh, we especially are interested to hear if there's certain kinds of
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 4.3s, EN.
EN_tAW9zt_jqmu_W000155 · in -22.8 dBFS · gain +2.8 dB · emolia-02308
(normal-paced, no disfluency, clear, casual)Data points or metrics you'd like to see tracked over time in this area.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.4/10; 3.5s, EN.
EN_tAW9zt_jqmu_W000156 · in -20.3 dBFS · gain +0.3 dB · emolia-02308
(brisk, almost no disfluency, clear, playful)But again, if you want to ask us any other questions about what we presented that is also welcome to you, do not feel beholden to these questions at all.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, dramatic; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.8/10; 5.5s, EN.
EN_tAW9zt_jqmu_W000157 · in -19.0 dBFS · gain -1.0 dB · emolia-02308
(thankfulness gratitude, relief, embarrassment·normal-paced, some disfluency, average clarity, casual)Good afternoon, thank you for your presentation. I was not able to
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, relief, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.0/10; 3.5s, EN.
EN_tAW9zt_jqmu_W000158 · in -19.3 dBFS · gain -0.7 dB · emolia-02308