Manifest tier. proxy_taillift, rule PXR, T=0.5, step cap 0.25. Population 1,035 chains (12 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 916.
Rule.PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable Source. trajectories_v5lift.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.proxy_taillift__PXR__T0.50__C0.25__INTERNAL — population 1,035 chains (12 h). SHAREABLE variant: 916. Filter.rule=='PXR' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and qmax>=0.5 and cmax<=0.25 This tier was resampled. The manifest's earlier published filter applied the one-sided test to every tier regardless of its actual rule. Tier populations were always counted under the correct rule, so the sizes quoted here were never wrong — but the earlier selection drew from a contaminated pool, of which 94.1 % were not members of this tier. These samples are drawn with the corrected filter, which tests qmax >= T and cmax <= C — per-row quantities computed under this tier's own rule. For a proxy tier no d_a/d_b predicate could be correct, because a proxy-rescued chain may legitimately exceed the cap on the named axis while its smoothness is certified on the proxy axis. Sampled from 1,035 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Thankfulness Gratitude barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.52.
At the same time Doubt goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.22 (lower than 78 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.14, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.62 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.62 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.62, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 21 s · en · emolia
hear it un-normalised (raw levels, max seam 1.0 dB)
k 4d_a -0.514d_b 0.518step_a 0.248step_b 0.241min_cos_consec 0.6181min_cos_anchor 0.6181dataset emolialang enspeaker EN_oA_eQnfPVGctrack EN_oA_eQnfPVGctotal 20.8slevel spread 1.0 dBmax seam 1.0 dBqmax 0.514cmax 0.248cos from recomputed from spkemb_traj
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · average recording, normally alert, slightly relaxed
(slow, steady, frequent disfluency, monologue)If you believe that you can make a business case out of your project.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.5/10; 4.6s, EN.
EN_oA_eQnfPVGc_W000231 · in -19.4 dBFS · gain -0.7 dB · emolia-01486
(measured, fairly steady, frequent disfluency, monologue)(low mumble) Uh, a project for, uh, (low mumble) that is focusing on, (ahem) uh, ecotunism.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.7/10; 4.2s, EN.
EN_oA_eQnfPVGc_W000233 · in -19.6 dBFS · gain -0.4 dB · emolia-01486
(malevolence malice, anger· measured, fairly steady, frequent disfluency, monologue)It's very clear. You have a tourist and the tourist in the end, (low mumble) uh, will pay. They will even pay
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, anger; style: monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.0/10; 5.8s, EN.
EN_oA_eQnfPVGc_W000234 · in -18.7 dBFS · gain -1.3 dB · emolia-01486
(measured, steady, some disfluency, monologue)Before the trip starts, it's one of the easiest business cases because
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.0/10; 5.8s, EN.
EN_oA_eQnfPVGc_W000235 · in -19.6 dBFS · gain -0.5 dB · emolia-01486
This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Interest below average — 0.39, lower than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.59.
At the same time Emotional Numbness goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.21, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 50 s · en · emolia
hear it un-normalised (raw levels, max seam 0.6 dB)
k 4d_a -0.534d_b 0.592step_a 0.223step_b 0.230min_cos_consec 0.7727min_cos_anchor 0.7532dataset emolialang enspeaker EN_3vqBV15JtP8track EN_3vqBV15JtP8total 49.8slevel spread 0.6 dBmax seam 0.6 dBqmax 0.534cmax 0.230cos from recomputed from spkemb_traj
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, average recording, quiet background, slightly relaxed, frequent disfluency, somewhat unclear
(emotional numbness · normal-paced, normally alert, fairly steady, monologue)So this is where the MySQL document store (low mumble) comes, where the SQL is completely now optional.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, didactic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 6.6s, EN.
EN_3vqBV15JtP8_W000045 · in -18.6 dBFS · gain -1.4 dB · emolia-00416
(concentration, malevolence malice·measured, normally alert, steady, didactic)And you don't need to copy your data and to duplicate the data (low mumble) between NoSQL and relational to do (low mumble) what you need to do. And I will show you that today.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, malevolence malice; style: didactic, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.7/10; 10.0s, EN.
EN_3vqBV15JtP8_W000046 · in -19.2 dBFS · gain -0.8 dB · emolia-00416
(concentration · measured, subdued, steady, didactic)Is, uh, (low mumble) storing all the, (ahem) uh, (low mumble) collection, but also all the relational table together in one single machine or duplicate in clusters and, and (low mumble) replicas, if you need that, of course, but all on the same (low mumble) architecture.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.3/10; 15.6s, EN.
EN_3vqBV15JtP8_W000047 · in -19.1 dBFS · gain -0.9 dB · emolia-00416
(interest, concentration · measured, subdued, fairly steady, casual)So it's a solution for everybody. So the business owner are very happy because they have the ACID compliance that, uh, (low mumble) (ahem) some, uh, (low mumble) document store doesn't have. So in, (low mumble) uh, and by default on many, uh, (low mumble) document store, if you plug the power, you may lose your data.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.4/10; 17.2s, EN.
EN_3vqBV15JtP8_W000048 · in -18.9 dBFS · gain -1.1 dB · emolia-00416
This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.64.
At the same time Doubt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.22 (lower than 78 % of clips in this corpus), a change of -0.73. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.12, then +0.16, then +0.15 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · zh · emolia
hear it un-normalised (raw levels, max seam 1.8 dB)
k 5d_a -0.728d_b 0.637step_a 0.223step_b 0.202min_cos_consec 0.7657min_cos_anchor 0.7657dataset emolialang zhspeaker ZH_B00038_S05687track ZH_B00038_S05687total 36.0slevel spread 1.8 dBmax seam 1.8 dBqmax 0.637cmax 0.223cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, no background noise, normally alert, slightly relaxed
This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Shame around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.48.
At the same time Sourness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.56. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.16, then +0.02, then +0.12 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.62 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.62, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 87 s · en · podcast
hear it un-normalised (raw levels, max seam 4.0 dB)
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, normally alert, moderately variable
(sourness, jealousy and envy, bitterness · neutral tension, some disfluency, average clarity, conversational)for all context.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as sourness, jealousy and envy, bitterness; style: conversational, casual; below-average recording, some background noise; genuineness 5.2/6; vocal-burst blend 8.1/10; 26.2s, EN.
498027_00072128 · in -21.5 dBFS · gain +1.5 dB · podcast-00091
(intoxication altered states of consciousness, jealousy and envy, amusement·relaxed, frequent disfluency, average clarity, casual)It's like one of the brown ones, I think.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, jealousy and envy, amusement; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 5.1/10; 27.4s, EN.
498027_00074748 · in -25.5 dBFS · gain +5.5 dB · podcast-01983
(longing, contemplation, infatuation· relaxed, frequent disfluency, somewhat unclear, casual)looks black from far away. (low mumble) Um Yeah, no, I'll see. It like I would like to have a serious relationship. And it's part of my new year's is like I wanna give relationships a shot. But in the past I've never really
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, contemplation, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 3.8/10; 15.7s, EN.
498027_00078376 · in -25.6 dBFS · gain +5.6 dB · podcast-02310
(embarrassment, doubt, longing · relaxed, some disfluency, average clarity, casual)the past someone like I was like super in like with kind of
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, doubt, longing; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.1/10; 3.9s, EN.
498027_00080216 · in -24.6 dBFS · gain +4.6 dB · podcast-02677
(shame, relief, pain·neutral tension, frequent disfluency, average clarity, casual)felt the urge to overcome all my fears and issues. But we'll see. I hope in the new year's like I'm able to delve into that side of myself more. (ahem) Um but besides that, my main thing
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, relief, pain; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.9/10; 12.8s, EN.
498027_00081248 · in -24.7 dBFS · gain +4.7 dB · podcast-02361
Concentration ↓ / Intoxication Altered States of Consciousness ↑proxy_taillift__PXR__T0.50__C0.25__INTERNAL · #5
This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Intoxication Altered States of Consciousness barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.52.
At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.43 (lower than 57 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.20, then +0.07, then +0.16 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 33 s · en · emolia
hear it un-normalised (raw levels, max seam 3.3 dB)
k 5d_a -0.533d_b 0.516step_a 0.250step_b 0.202min_cos_consec 0.8503min_cos_anchor 0.8845dataset emolialang enspeaker EN_WxNLtAuGb1Ytrack EN_WxNLtAuGb1Ytotal 33.4slevel spread 5.8 dBmax seam 3.3 dBqmax 0.516cmax 0.250cos from recomputed from spkemb_traj
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth, balanced body, average recording, normally alert, slightly relaxed, some disfluency
(concentration · measured, fairly steady, slurred, whispered)So let's say if we have a (ahem) 2-pentanol, we can turn this into an OH group, (ahem) uh, using sodium borohydride, which is a reducing agent.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as concentration; style: whispered, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.3/10; 8.5s, EN.
EN_WxNLtAuGb1Y_W000039 · in -27.7 dBFS · gain +7.7 dB · emolia-02475
(emotional numbness·normal-paced, steady, slurred, monologue)But the negative formal charge is not the same as the partial positive charge on boron. They're two different things. This nucleophilic hydride ion
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, whispered; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.6/10; 8.3s, EN.
EN_WxNLtAuGb1Y_W000041 · in -26.1 dBFS · gain +6.1 dB · emolia-02475
(normal-paced, fairly steady, average clarity, casual)And then in the next step, it can grab a hydrogen from H2O+, or we could use water as well. Let's use water in this example.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.0/10; 7.6s, EN.
EN_WxNLtAuGb1Y_W000043 · in -22.8 dBFS · gain +2.8 dB · emolia-02475
(normal-paced, fairly steady, average clarity, whispered)Lithium aluminum hydride works the same way. Sodium borohydride is used to reduce
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, casual; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 1.3/10; 4.8s, EN.
EN_WxNLtAuGb1Y_W000045 · in -24.0 dBFS · gain +4.0 dB · emolia-02475
(normal-paced, fairly steady, slurred, casual)Aldehydes, ketones, it could even reduce acid chlorides.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 2.5/10; 3.6s, EN.
EN_WxNLtAuGb1Y_W000046 · in -21.9 dBFS · gain +1.9 dB · emolia-02475
This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.51.
At the same time Relief goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.64. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.22, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 39 s · de · emolia
k 4d_a -0.644d_b 0.511step_a 0.244step_b 0.245min_cos_consec 0.8752min_cos_anchor 0.8764dataset emolialang despeaker DE_4MqG3TzV9mutrack DE_4MqG3TzV9mutotal 39.3slevel spread 0.7 dBmax seam 0.7 dBqmax 0.511cmax 0.245cos from recomputed from spkemb_traj
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, light breath
(relief, impatience and irritability, contempt · brisk, energised, neutral tension, cartoonish)ihr könnt diese Stadt zur Freiheit führen. Freiheit gegen Dol-Guruk. Sie warten nur auf euch. Aber wir brauchen Geld dafür.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, impatience and irritability, contempt; style: cartoonish, storytelling; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.6/10; 6.1s, DE.
DE_4MqG3TzV9mu_W000045 · in -16.9 dBFS · gain -3.1 dB · emolia-00163
(impatience and irritability, anger, interest· brisk, energised, neutral tension, casual)Ohne Geld funktioniert das nicht. Wir müssen all die Menschen zum Aufstand bewegen. Oder, wie gesagt, ihr brecht ein. Die andere Alternative ist, ihr geht zu Dol-Guruk hin und macht Diplomatie. Erkennt euch nicht, ihr habt euch nicht zu Schulden kommen lassen, ihr seid keine Novadis, also bringt euch nicht um, ja. Ihr seid noch mit, mit, äh, (ahem) heiler Haut, mit weißer, weißer Weste, ja, sagt man so.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger, interest; style: casual, authoritative; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.6/10; 19.9s, DE.
DE_4MqG3TzV9mu_W000046 · in -17.0 dBFS · gain -3.0 dB · emolia-00163
(confusion·normal-paced, normally alert, slightly relaxed, conversational)So, und hier geht er hin und sagt, (ahem) hier, wir haben gehört, du hast Dokument. Was willst du für Dokument? Das ist Diplomatie.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: conversational, casual; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.1/10; 6.8s, DE.
DE_4MqG3TzV9mu_W000047 · in -17.2 dBFS · gain -2.8 dB · emolia-00163
(confusion, doubt· normal-paced, normally alert, slightly relaxed, conversational)Dann seid ihr Freunde von, von, von mir, ihr seid Defendies. Das ist, das sind die drei Tipps, die mir einfallen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, doubt; style: conversational, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 5.9s, DE.
DE_4MqG3TzV9mu_W000048 · in -16.6 dBFS · gain -3.5 dB · emolia-00163
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.66.
At the same time Thankfulness Gratitude goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.17, then +0.25, then +0.21 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 54 s · en · emolia
k 5d_a -0.523d_b 0.657step_a 0.230step_b 0.248min_cos_consec 0.8503min_cos_anchor 0.8703dataset emolialang enspeaker EN_tOV_V0I_GJAtrack EN_tOV_V0I_GJAtotal 53.6slevel spread 2.5 dBmax seam 2.5 dBqmax 0.523cmax 0.248cos from recomputed from spkemb_traj
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, fairly steady, frequent disfluency, moderate pitch range
(thankfulness gratitude · measured, normally alert, slightly relaxed, casual)(low mumble) In my view, (low mumble) it is desirable, (ahem) as the government has announced, that we should have a proper and thoroughgoing and independent review (ahem) of the PPRF. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 1.6/10; 12.2s, EN.
EN_tOV_V0I_GJA_W000039 · in -20.6 dBFS · gain +0.7 dB · emolia-00663
(contemplation, concentration, doubt·normal-paced, normally alert, slightly relaxed, monologue)And in that regard, (low mumble) uhm, I'm hoping that we can think laterally and think about other possible ways in which we might, (low mumble) uhm, (ahem) assess performance and do so in ways that perhaps are less
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, concentration, doubt; style: monologue, casual; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.7/10; 13.4s, EN.
EN_tOV_V0I_GJA_W000040 · in -20.3 dBFS · gain +0.3 dB · emolia-00663
(concentration, emotional numbness, helplessness·measured, subdued, slightly relaxed, casual)(low mumble) Um, intrusive, less stressful for staff, (low mumble) with lower transaction costs, (low mumble) but without any significant reduction and kind of, (low mumble) uh, the incentives to, to, uh, (low mumble) generate good performance.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, emotional numbness, helplessness; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 1.2/10; 12.3s, EN.
EN_tOV_V0I_GJA_W000041 · in -22.9 dBFS · gain +2.9 dB · emolia-00663
(doubt, helplessness, contemplation· measured, very low-energy, neutral tension, conversational)Maybe we can't come up with anything (low mumble) that's going to be a (low mumble) substantive improvement, but I think we should at least have a go. One thing I believe
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, helplessness, contemplation; style: conversational, casual; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 1.7/10; 10.3s, EN.
EN_tOV_V0I_GJA_W000042 · in -20.7 dBFS · gain +0.7 dB · emolia-00663
(infatuation· measured, normally alert, slightly relaxed, casual)The kind of peer review, uh, (low mumble) assessment process that we have in place at the moment.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as infatuation; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.8/10; 4.7s, EN.
EN_tOV_V0I_GJA_W000044 · in -21.2 dBFS · gain +1.2 dB · emolia-00663
This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Elation around average — 0.46, lower than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.52.
At the same time Disgust goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.24, then +0.05, then -0.00 — not a clean run: step 4 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 57 s · en · emolia
k 5d_a -0.531d_b 0.524step_a 0.234step_b 0.238min_cos_consec 0.8020min_cos_anchor 0.7887dataset emolialang enspeaker EN_Ty-HrGJjg1Ytrack EN_Ty-HrGJjg1Ytotal 56.7slevel spread 5.1 dBmax seam 5.1 dBqmax 0.524cmax 0.238cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, brisk, energised, light breath
(slightly relaxed, fairly steady, almost no disfluency, casual)Has anybody called XRP and Cardano tech stocks?
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, authoritative; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.2/10; 3.4s, EN.
EN_Ty-HrGJjg1Y_W000015 · in -17.4 dBFS · gain -2.6 dB · emolia-01133
(impatience and irritability, jealousy and envy, contempt· slightly relaxed, fairly steady, some disfluency, authoritative)Like tech stocks? No. Well, let's hope XRP is not like a stock or a security. The government has bigger problems. Is anybody talking about altcoins as tech investments? By the way, you turn on CNBC today and check out how other tech investments are training, trading, Google, et cetera. Apple flying altcoins are next.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as impatience and irritability, jealousy and envy, contempt; style: authoritative, dramatic; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 2.6/10; 26.2s, EN.
EN_Ty-HrGJjg1Y_W000016 · in -22.5 dBFS · gain +2.5 dB · emolia-01133
(pleasure ecstasy, affection, thankfulness gratitude·neutral tension, moderately variable, almost no disfluency, casual)Hello from India at 1 a.m. Raj, we appreciate you staying up late. SEO from London, hello. Charlie S. Robert. Charlie's here early today. Thank you.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, full; average clarity, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, affection, thankfulness gratitude; style: casual, dramatic; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.9/10; 11.1s, EN.
EN_Ty-HrGJjg1Y_W000017 · in -19.4 dBFS · gain -0.6 dB · emolia-01133
(elation, hope enthusiasm optimism, pleasure ecstasy · neutral tension, moderately variable, some disfluency, casual)Appreciate having everybody on the show, whether you're new to the show, you've been here for a while, you're new to crypto, you're in the right place.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.2/10; 8.2s, EN.
EN_Ty-HrGJjg1Y_W000018 · in -21.2 dBFS · gain +1.2 dB · emolia-01133
(elation, hope enthusiasm optimism, pride·slightly tense, moderately variable, some disfluency, authoritative)Okay. Ash one three LD. Can you talk about the correlation between gold and the dollar? Absolutely.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as elation, hope enthusiasm optimism, pride; style: authoritative, dramatic; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.5/10; 7.2s, EN.
EN_Ty-HrGJjg1Y_W000019 · in -19.7 dBFS · gain -0.3 dB · emolia-01133
This chain comes from the proxy rule: the same two-sided test as above, but because Doubt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Doubt below average — 0.37, lower than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.54.
At the same time Elation goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.46 (lower than 54 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.09 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 57 s · en · emolia
k 4d_a -0.507d_b 0.536step_a 0.231step_b 0.221min_cos_consec -0.0919min_cos_anchor -0.1033dataset emolialang enspeaker EN_0yVpoHamPvEtrack EN_0yVpoHamPvEtotal 56.9slevel spread 1.2 dBmax seam 1.2 dBqmax 0.507cmax 0.231cos from recomputed from spkemb_traj
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(elation, hope enthusiasm optimism, pleasure ecstasy · neutral tension, casual, conversational)To people in internet. So I can say, okay, this is very good because maybe Google (ahem) is liking the site, no?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.1/10; 9.0s, EN.
EN_0yVpoHamPvE_W000007 · in -17.8 dBFS · gain -2.2 dB · emolia-01011
(concentration·slightly relaxed, casual, didactic)So with any kind of changes that you see in the performance report, my recommendation is always to try to drill down and find some examples of this change in
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: casual, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.4/10; 12.7s, EN.
EN_0yVpoHamPvE_W000008 · in -16.6 dBFS · gain -3.4 dB · emolia-01011
(concentration, contemplation, contentment· slightly relaxed, monologue, casual)In a way that is very visible. Uh, (low mumble) so for example, you could look in to see if there are specific pages that have changed in like the ranking or in the number of impressions. (low mumble) Uh, or if there are specific queries where you see this change. So one thing that might be happening is that maybe just more people are searching.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, contemplation, contentment; style: monologue, casual; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.8/10; 19.7s, EN.
EN_0yVpoHamPvE_W000009 · in -16.9 dBFS · gain -3.1 dB · emolia-01011
(doubt· slightly relaxed, casual, monologue)(ahem) Uh, it might also be happening that you're showing for slightly different queries than before, (low mumble) uh, where it's like there may be more, more impressions for those queries, but you're not quite in that competitive range yet.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.4/10; 15.0s, EN.
EN_0yVpoHamPvE_W000010 · in -16.6 dBFS · gain -3.4 dB · emolia-01011
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness below average — 0.39, lower than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.57.
At the same time Contemplation goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.57. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.00, then +0.11, then +0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, fairly steady
(contemplation, doubt, intoxication altered states of consciousness · slow, normally alert, relaxed, casual)impacts or rocks or something like that that maybe that maybe took them out to it's one of the
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, doubt, intoxication altered states of consciousness; style: casual, monologue; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 4.2/10; 6.6s, EN.
430781_00339552 · in -32.5 dBFS · gain +12.5 dB · podcast-01336
(confusion, fear, contemplation · slow, normally alert, relaxed, casual)biggest mass extinctions that we're aware of, but we don't really know, right? 'Cause like who
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, fear, contemplation; style: casual, conversational; average recording, no background noise; genuineness 3.8/6; vocal-burst blend 4.3/10; 7.2s, EN.
430781_00340208 · in -33.2 dBFS · gain +13.2 dB · podcast-01331
(fear, distress, astonishment surprise·normal-paced, normally alert, slightly relaxed, casual)there could have been a mass extinction like that, you know, fifty million years before this, and we just haven't figured it out yet because we don't have the fossil records.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, distress, astonishment surprise; style: casual, monologue; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.7/10; 8.4s, EN.
430781_00340928 · in -35.1 dBFS · gain +15.1 dB · podcast-01323
(measured, normally alert, slightly relaxed, casual)So there you go. That takes us all the way through the development of complex life. (low mumble) Uh reptiles are there,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.1/10; 8.4s, EN.
430781_00341768 · in -30.4 dBFS · gain +10.4 dB · podcast-01322
(emotional numbness, intoxication altered states of consciousness·slow, very low-energy, relaxed, casual)mammals are out there doing their mammary thing. (low mumble) Um birds are being weird. I
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness, intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.1/10; 6.7s, EN.
430781_00342608 · in -31.2 dBFS · gain +11.2 dB · podcast-01332
Infatuation ↓ / Sexual Lust ↑proxy_taillift__PXR__T0.50__C0.25__INTERNAL · #11
This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sexual Lust below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.64.
At the same time Infatuation goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.20, then +0.04, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.10 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, normally alert, light breath
(measured, slightly relaxed, fairly steady, formal)A man who served time on home invasion and firearms charges was in Keona's apartment.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.1/10; 5.6s.
batch190_part2_batch190_part2_chunk_2702_1_2530547 · in -20.2 dBFS · gain +0.2 dB · snippets-00473
(normal-paced, slightly relaxed, fairly steady, storytelling)Lead detective Sean O'Connell says Carmela was one of the first people he looked at.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: storytelling, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.5/10; 5.2s.
batch190_part2_batch190_part2_chunk_2702_1_2530745 · in -21.7 dBFS · gain +1.7 dB · snippets-00473
(emotional numbness, fear, pain·measured, slightly relaxed, fairly steady, narration)saying she knew someone capable of this crime.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear, pain; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 4.9/10; 3.0s.
batch190_part2_batch190_part2_chunk_2702_1_2530826 · in -23.7 dBFS · gain +3.7 dB · snippets-00473
(doubt, confusion, fear · measured, relaxed, moderately variable, conversational)I think somebody has her and they need to bring her home. Why would they (ahem) uh want to (low mumble) uh kidnap her? I don't know.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, confusion, fear; style: conversational, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.3/10; 6.0s.
batch190_part2_batch190_part2_chunk_2702_1_2530835 · in -20.6 dBFS · gain +0.6 dB · snippets-00473
(sexual lust, longing· measured, slightly relaxed, fairly steady, casual)Absolutely. We knew we had something. We could at least say definitively, this is where she went.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, longing; style: casual, monologue; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 4.5/10; 5.4s.
batch190_part2_batch190_part2_chunk_2702_1_2530994 · in -21.3 dBFS · gain +1.3 dB · snippets-00473
This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.56.
At the same time Thankfulness Gratitude goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.58. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.21 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.21, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 35 s · en · emolia
k 4d_a -0.584d_b 0.562step_a 0.218step_b 0.196min_cos_consec 0.2423min_cos_anchor 0.2084dataset emolialang enspeaker EN_T5-vp8X6iM8track EN_T5-vp8X6iM8total 35.5slevel spread 3.1 dBmax seam 2.9 dBqmax 0.562cmax 0.218cos from recomputed from spkemb_traj
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(thankfulness gratitude, awe, astonishment surprise · casual, monologue)I think for them they're like, wow, we didn't realize that this many folks
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, awe, astonishment surprise; style: casual, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.9/10; 4.1s, EN.
EN_T5-vp8X6iM8_W000073 · in -17.0 dBFS · gain -3.0 dB · emolia-02495
(casual, conversational)(low mumble) Uhm, would be interested in that I think more and more that, (low mumble) uh, folks who are interested talk to folks like at Cooper Square or talk to our folks that are about to do it. They get excited like, wow, okay, this is an actual option that could keep us in our home.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.2/10; 16.1s, EN.
EN_T5-vp8X6iM8_W000074 · in -19.9 dBFS · gain -0.1 dB · emolia-02495
(interest· conversational, casual)Now what happens to the state in all of this? I mean, a lot of the models that you're describing are ones where, given half a chance, communities can develop, become their own owners, their own managers. Fantastic.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as interest; style: conversational, casual; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.7/10; 11.2s, EN.
EN_T5-vp8X6iM8_W000075 · in -20.1 dBFS · gain +0.1 dB · emolia-02495
(pain· casual, conversational)But you grew up in public housing. We still need public housing, right?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pain; style: casual, conversational; good recording, quiet background; genuineness 4.7/6; vocal-burst blend 3.9/10; 3.6s, EN.
EN_T5-vp8X6iM8_W000077 · in -19.5 dBFS · gain -0.5 dB · emolia-02495
This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Thankfulness Gratitude below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.63.
At the same time Shame goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.41 (lower than 59 % of clips in this corpus), a change of -0.58. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.05, then +0.14 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, thin, average recording, quiet background, brisk, normally alert, some disfluency, light breath
(shame, anger, sadness · slightly relaxed, moderately variable, clear, monologue)in questo modo riescono ad avere un beneficio nel corpo e nella mente, allontanando il vuoto, l'immaterialità e il senso di alienazione che purtroppo i fatti di cronaca di tutti i giorni drammaticamente ci raccontano.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, anger, sadness; style: monologue, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.0/10; 14.8s, IT.
italy_17_903_1487136_1501984 · in -11.4 dBFS · gain -8.6 dB · eurospeech-01741
(shame, disappointment, awe· slightly relaxed, moderately variable, clear, monologue)i Paesi europei più illuminati di noi, come la Francia e la Germania, considerano i corpi di ballo un valore aggiunto? La Germania ha 50 corpi di ballo e la Francia addirittura 95 e considera i corpi di ballo un arricchimento, non solo culturale ma anche delle potenzialità
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, disappointment, awe; style: monologue, cartoonish; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 4.7/10; 20.0s, IT.
italy_17_903_1514688_1534688 · in -10.9 dBFS · gain -9.1 dB · eurospeech-01741
(awe, anger, interest·neutral tension, moderately variable, clear, cartoonish)Non possiamo ancora una volta rinunciare ad un valore tipico e specifico della nostra nazionalità
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, anger, interest; style: cartoonish, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 5.4/10; 20.0s, IT.
italy_17_903_1534688_1554688 · in -11.3 dBFS · gain -8.7 dB · eurospeech-01741
(concentration, triumph, contemplation·slightly relaxed, fairly steady, average clarity, monologue)nazionalità italica e non possiamo soprattutto fingere che ciò rappresenti una responsabilità delle fondazioni lirico-sinfoniche. Il Governo è stato sordo nell'arco di tutta la legislatura ad un'esigenza fondamentale che la nostra cultura ci chiede,
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, triumph, contemplation; style: monologue, formal; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.6/10; 19.9s, IT.
italy_17_903_1554688_1574623 · in -12.3 dBFS · gain -7.7 dB · eurospeech-01741
(thankfulness gratitude, anger, relief·neutral tension, moderately variable, average clarity, dramatic)ad un obbligo che doveva corrispondere ad un finanziamento delle fondazioni lirico-sinfoniche e ad una visione (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, anger, relief; style: dramatic, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.3/10; 10.3s, IT.
italy_17_903_1574623_1584920 · in -12.0 dBFS · gain -8.0 dB · eurospeech-01741
Jealousy and Envy ↓ / Pain ↑proxy_taillift__PXR__T0.50__C0.25__INTERNAL · #14
This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.55.
At the same time Jealousy and Envy goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 39 s · en · emolia
k 4d_a -0.517d_b 0.549step_a 0.245step_b 0.196min_cos_consec 0.9201min_cos_anchor 0.9183dataset emolialang enspeaker EN_IxFYEAhlc_0track EN_IxFYEAhlc_0total 38.9slevel spread 0.9 dBmax seam 0.9 dBqmax 0.517cmax 0.245cos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, steady
(jealousy and envy · normal-paced, light breath, formal, monologue)They visited Fireclay Works at Stourbridge and his brother John Wood, who had a forge at Wendsbury
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.3s, EN.
EN_IxFYEAhlc_0_W000029 · in -14.1 dBFS · gain -5.9 dB · emolia-00671
(intoxication altered states of consciousness·measured, no audible breath, formal, newsreading)John was working scraps and using pots to do so.In 1763, Charles and John Wood patented their iron making process, which is usually described by historians as potting and stamping.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 15.3s, EN.
EN_IxFYEAhlc_0_W000030 · in -14.7 dBFS · gain -5.3 dB · emolia-00671
(measured, light breath, formal, monologue)This followed one to John Wood alone in 1761
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 4.7s, EN.
EN_IxFYEAhlc_0_W000031 · in -13.8 dBFS · gain -6.2 dB · emolia-00671
(measured, light breath, formal, newsreading)In December 1763, Peter Howe's tobacco firm became bankrupt, as did the Braziers' business of Gabriel Griffiths and Robert Ross
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.2s, EN.
EN_IxFYEAhlc_0_W000032 · in -14.6 dBFS · gain -5.4 dB · emolia-00671
This chain comes from the proxy rule: the same two-sided test as above, but because Helplessness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Helplessness essentially absent — 0.07, lower than 93 % of clips in this corpus — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.69.
At the same time Concentration goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.37 (lower than 63 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.34, then +0.00, then +0.00, then +0.35 — a slow start, with most of the change arriving in the final step.
The largest step is 0.35, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fatigue Exhaustion below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.54.
At the same time Concentration goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.41 (lower than 59 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.19, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 27 s · en · emolia
k 4d_a -0.534d_b 0.541step_a 0.228step_b 0.191min_cos_consec 0.8103min_cos_anchor 0.8453dataset emolialang enspeaker EN_X2-y7r3xsj8track EN_X2-y7r3xsj8total 26.6slevel spread 2.3 dBmax seam 2.3 dBqmax 0.534cmax 0.228cos from recomputed from spkemb_traj
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · slightly dark, balanced body, average recording, quiet background, frequent disfluency, somewhat unclear, fairly narrow pitch
(concentration · slow, very low-energy, relaxed, didactic)And we've tried to do some analysis of these sectors and the businesses there to find how we can, uh, (low mumble) how we can design, (ahem) um, or how.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, neutral openness; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_X2-y7r3xsj8_W000003 · in -18.2 dBFS · gain -1.8 dB · emolia-02297
(slow, normally alert, slightly relaxed, monologue)Business models can be designed for improvement of rural businesses.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.6/10; 5.0s, EN.
EN_X2-y7r3xsj8_W000004 · in -18.7 dBFS · gain -1.3 dB · emolia-02297
(measured, normally alert, slightly relaxed, monologue)Three, uh, (low mumble) cafes with different themes and today our virtual visit.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.5/10; 4.8s, EN.
EN_X2-y7r3xsj8_W000006 · in -16.4 dBFS · gain -3.6 dB · emolia-02297
(measured, normally alert, slightly relaxed, monologue)Is going to be, as we said, to this agro-industrial research hub. And, (low mumble) uhm, Apolline.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.8/10; 5.9s, EN.
EN_X2-y7r3xsj8_W000007 · in -17.4 dBFS · gain -2.6 dB · emolia-02297
Jealousy and Envy ↓ / Elation ↑proxy_taillift__PXR__T0.50__C0.25__INTERNAL · #17
This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Elation around average — 0.46, lower than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.51.
At the same time Jealousy and Envy goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.07, then +0.21 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult feminine voice · slightly bright, fairly smooth, quiet background, moderately variable, some disfluency, wide pitch range
(jealousy and envy, contempt, sourness · normal-paced, normally alert, slightly relaxed, dramatic)approach it this year. Okay? Jimena Lina,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, contempt, sourness; style: dramatic, authoritative; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.0/10; 7.2s, EN.
612337_00192836 · in -24.2 dBFS · gain +4.2 dB · podcast-01855
(fatigue exhaustion, impatience and irritability, disappointment·fast, energised, neutral tension, casual)the event, you will be at Chicago is all the time. Jimena, one thing. Actually, I'm in Chicago, and it's at the 8 of the machine,
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as fatigue exhaustion, impatience and irritability, disappointment; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.2/10; 10.8s, EN.
612337_00193552 · in -23.2 dBFS · gain +3.2 dB · podcast-01848
(fast, energised, neutral tension, casual)We're going to have three links. In one or cola, as we said in my place. Uh dicen Jennifer Olivia, Antonio.
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 5.0/6; vocal-burst blend 7.5/10; 6.6s, EN.
612337_00195208 · in -24.5 dBFS · gain +4.5 dB · podcast-01896
(elation, triumph, pleasure ecstasy·brisk, energised, slightly relaxed, casual)And we're at the ocher of the maiden and three files.
full caption & clip details
A child feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as elation, triumph, pleasure ecstasy; style: casual, playful; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.6/10; 3.1s, EN.
612337_00198648 · in -21.3 dBFS · gain +1.3 dB · podcast-04585
Concentration ↓ / Sexual Lust ↑proxy_taillift__PXR__T0.50__C0.25__INTERNAL · #18
This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sexual Lust essentially absent — 0.03, lower than 97 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.69.
At the same time Concentration goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.37 (lower than 63 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.21, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 29 s · zh · emolia
k 4d_a -0.500d_b 0.686step_a 0.238step_b 0.246min_cos_consec 0.8174min_cos_anchor 0.8320dataset emolialang zhspeaker ZH_B00057_S02582track ZH_B00057_S02582total 29.1slevel spread 0.7 dBmax seam 0.7 dBqmax 0.500cmax 0.246cos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, slightly relaxed, no disfluency
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.52.
At the same time Doubt goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.37 (lower than 63 % of clips in this corpus), a change of -0.56. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.05, then +0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 39 s · zh · emolia
k 4d_a -0.561d_b 0.517step_a 0.233step_b 0.248min_cos_consec 0.7715min_cos_anchor 0.7715dataset emolialang zhspeaker ZH_B00000_S05686track ZH_B00000_S05686total 38.5slevel spread 1.0 dBmax seam 0.5 dBqmax 0.517cmax 0.248cos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, no disfluency
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.6/10; 5.5s, ZH.
ZH_B00000_S05686_W000010 · in -17.8 dBFS · gain -2.2 dB · emolia-03275
Sexual Lust ↓ / Hope Enthusiasm Optimism ↑proxy_taillift__PXR__T0.50__C0.25__INTERNAL · #20
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.65.
At the same time Sexual Lust goes the other way, from 0.76 (higher than 76 % of clips in this corpus) to 0.17 (lower than 83 % of clips in this corpus), a change of -0.59. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.25, then +0.21, then +0.19 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 30 s · en · emolia
k 5d_a -0.591d_b 0.652step_a 0.203step_b 0.247min_cos_consec 0.7165min_cos_anchor 0.7215dataset emolialang enspeaker EN_KNPQuKAGK20track EN_KNPQuKAGK20total 30.5slevel spread 3.1 dBmax seam 3.1 dBqmax 0.591cmax 0.247cos from recomputed from spkemb_traj
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady, average clarity
(some disfluency, casual, monologue)The Squaxin Island tribe has stewarded this land since time immemorial.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.9/10; 4.1s, EN.
EN_KNPQuKAGK20_W000008 · in -21.5 dBFS · gain +1.5 dB · emolia-00902
(some disfluency, casual, playful)We're developed by the recipients of the 2122 OSPI OER Project Grant.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.1/10; 5.8s, EN.
EN_KNPQuKAGK20_W000010 · in -18.5 dBFS · gain -1.5 dB · emolia-00902
(little disfluency, monologue, casual)And this opportunity targets the development or adaptation.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.2/10; 3.8s, EN.
EN_KNPQuKAGK20_W000011 · in -20.7 dBFS · gain +0.7 dB · emolia-00902
(some disfluency, casual, monologue)Of openly licensed resources, especially in content areas that are currently lacking (low mumble) in (low mumble) really high quality standards aligned (ahem) resources.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 12.1s, EN.
EN_KNPQuKAGK20_W000012 · in -21.6 dBFS · gain +1.6 dB · emolia-00902
(some disfluency, casual, monologue)Just to give us a little bit of context and put us all on the same page here.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.8/10; 4.1s, EN.
EN_KNPQuKAGK20_W000013 · in -19.6 dBFS · gain -0.4 dB · emolia-00902