PXR at chain length k=4, all corpora, at the mining floor.
Rule.PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable Source. trajectories_v5.parquet | Family. rule x chain length Sampled from 110,162 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Emotional Numbness ↓ / Disgust ↑k-PXR-k4 · #1
This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it clearly present at 0.65, higher than 65 % of clips in this corpus. That is a total rise of 0.32.
At the same time Emotional Numbness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.12 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · snippets
hear it un-normalised (raw levels, max seam 1.4 dB)
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, slightly relaxed, moderate pitch range
(emotional numbness, contemplation · energised, fairly steady, no disfluency, narration)at the core of Manson's beliefs was the concept of
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as emotional numbness, contemplation; style: narration, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.3/10; 4.0s.
batch88_part2_batch88_part2_chunk_1792_1_1464480 · in -19.8 dBFS · gain -0.2 dB · snippets-01345
(normally alert, steady, almost no disfluency, newsreading)Boko Haram, a militant Islamist group, emerged in Nigeria in the early 2000s and has since gained notoriety for its brutal tactics and a series of high-profile attacks across the region.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.2s.
batch88_part2_batch88_part2_chunk_1792_1_1464548 · in -21.2 dBFS · gain +1.2 dB · snippets-01345
(normally alert, steady, almost no disfluency, narration)The events also highlight the importance of recognizing and addressing the warning signs of radicalization and extremist beliefs within communities.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, newsreading; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.9/10; 8.9s.
batch88_part2_batch88_part2_chunk_1792_1_1464679 · in -21.0 dBFS · gain +1.0 dB · snippets-01345
(normally alert, fairly steady, almost no disfluency, casual)was a Japanese doomsday cult founded by Shoko Asahara in 1984.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 2.2/10; 4.8s.
batch88_part2_batch88_part2_chunk_1792_1_1464772 · in -20.4 dBFS · gain +0.5 dB · snippets-01345
This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Malevolence Malice below average — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.59.
At the same time Fatigue Exhaustion goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.21, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · zh · emolia
hear it un-normalised (raw levels, max seam 2.4 dB)
k 4d_a -0.408d_b 0.589step_a 0.169step_b 0.215min_cos_consec 0.8150min_cos_anchor 0.8401dataset emolialang zhspeaker ZH_B00038_S01623track ZH_B00038_S01623total 24.6slevel spread 2.9 dBmax seam 2.4 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.6/10; 8.8s, ZH.
ZH_B00038_S01623_W000059 · in -19.4 dBFS · gain -0.6 dB · emolia-03655
Shame ↓ / Infatuation ↑k-PXR-k4 · #3
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation clearly present — 0.73, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Shame goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.03, then -0.00 — not a clean run: step 3 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.65 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 38 s · en · emolia
hear it un-normalised (raw levels, max seam 1.4 dB)
k 4d_a -0.278d_b 0.260step_a 0.147step_b 0.228min_cos_consec 0.6452min_cos_anchor 0.6646dataset emolialang enspeaker EN_B00030_S08249track EN_B00030_S08249total 38.3slevel spread 1.4 dBmax seam 1.4 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an elderly masculine voice · no background noise, very low-energy, steady
(shame, sadness, bitterness · measured, slightly relaxed, little disfluency, whispered)I wedded, nor dreaded the curse I had invoked, and its bitterness was not visited upon me at once.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, rough, very full; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as shame, sadness, bitterness; style: whispered, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.6/10; 7.3s, EN.
EN_B00030_S08249_W000333 · in -23.9 dBFS · gain +3.9 dB · emolia-00838
(contentment, awe, contemplation·slow, relaxed, little disfluency, whispered)But once again, in the silence of the night, they came through my lattice the soft sighs which had forsaken me, and they modelled themselves into a familiar sweet voice, saying
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is warm, slightly dark, slightly rough, balanced body; slurred, little disfluency, fairly narrow pitch, minimal breath; affect is mildly negative, neutral stance, neutral openness; reads as contentment, awe, contemplation; style: whispered, ASMR; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 12.4s, EN.
EN_B00030_S08249_W000334 · in -25.3 dBFS · gain +5.3 dB · emolia-00838
(contentment, longing, infatuation· slow, relaxed, no disfluency, whispered)Sleep in peace. For the spirit of love.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, very full; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, submissive, neutral openness; reads as contentment, longing, infatuation; style: whispered, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 7.6/10; 3.3s, EN.
EN_B00030_S08249_W000335 · in -24.2 dBFS · gain +4.2 dB · emolia-00838
(infatuation, sexual lust, awe· slow, relaxed, little disfluency, whispered)In taking to thy passionate heart her who is um, who is Ermengarde, thou art absolved for reasons which shall be made known to thee in heaven of thy vows unto Elenora.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, very full; slurred, little disfluency, narrow pitch range, minimal breath; affect is mildly negative, submissive, neutral openness; reads as infatuation, sexual lust, awe; style: whispered, ASMR; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.6/10; 14.9s, EN.
EN_B00030_S08249_W000336 · in -24.6 dBFS · gain +4.6 dB · emolia-00838
Relief ↓ / Jealousy and Envy ↑k-PXR-k4 · #4
This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Jealousy and Envy clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.
At the same time Relief goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.13, then +0.08 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 47 s · en · podcast
hear it un-normalised (raw levels, max seam 1.4 dB)
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, some disfluency, average clarity
(relief, triumph, pride · normally alert, neutral tension, moderately variable, casual)I enjoyed watching it. I was extremely happy when I was like, Oh, wow. (ahem) Uh finally, I don't have to see the Rangers all over my timeline in the postseason anymore. Like it's done. I've been released. Yep.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, vulnerable; reads as relief, triumph, pride; style: casual, dramatic; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 4.1/10; 14.1s, EN.
946244_00010608 · in -23.4 dBFS · gain +3.4 dB · podcast-04301
(teasing, fear, relief · normally alert, slightly relaxed, fairly steady, casual)And I don't think you're gonna have to see the Rangers go that far this year either.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing, fear, relief; style: casual, conversational; average recording, no background noise; genuineness 3.4/6; vocal-burst blend 3.1/10; 3.6s, EN.
946244_00012024 · in -22.0 dBFS · gain +2.0 dB · podcast-04297
(disappointment, bitterness, interest·very low-energy, slightly relaxed, fairly steady, casual)It's because they're the Rangers. The closest they've gotten in the last few years was that loss to the Kings, in which Alex m Alec Martinez took care of them in that three two victory, and I don't think they'll be able to overcome that for A few more years they had to let Terasenko walk. Not sure if Kane's gonna play. Kane might be retired with all his injuries.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disappointment, bitterness, interest; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.8/10; 19.8s, EN.
946244_00012552 · in -22.3 dBFS · gain +2.3 dB · podcast-06371
(jealousy and envy, emotional numbness, doubt·normally alert, slightly relaxed, fairly steady, conversational)Has there been any update on a Patrick Kane contract? Anything surrounding him? I've seen a lot of speculation on Twitter about him, but nothing's ever been confirmed.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as jealousy and envy, emotional numbness, doubt; style: conversational, casual; good recording, no background noise; mildly explicit content; genuineness 3.4/6; vocal-burst blend 3.7/10; 9.4s, EN.
946244_00014551 · in -22.5 dBFS · gain +2.5 dB · podcast-04293
Emotional Numbness ↓ / Sexual Lust ↑k-PXR-k4 · #5
This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sexual Lust clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.31.
At the same time Emotional Numbness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.23, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 42 s · en · podcast
hear it un-normalised (raw levels, max seam 2.2 dB)
Unchanged across all 4 clips: an adult masculine voice · neutral-bright
(emotional numbness, helplessness, disappointment · measured, very low-energy, slightly relaxed, monologue)not positive. And we're just slipping further and further behind the English Premier League
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness, helplessness, disappointment; style: monologue, whispered; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 3.5/10; 6.8s, EN.
635231_00049160 · in -32.4 dBFS · gain +12.4 dB · podcast-04191
(fatigue exhaustion, emotional numbness, infatuation· measured, subdued, relaxed, monologue)(low mumble) and other leagues in Europe. And, you know, to stay on the kind of Celtic and Rangers comparison again, you know, back in the day, you know,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is slightly warm, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, emotional numbness, infatuation; style: monologue, whispered; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 4.4/10; 9.2s, EN.
635231_00049840 · in -31.7 dBFS · gain +11.7 dB · podcast-04149
(jealousy and envy, sexual lust, infatuation · measured, normally alert, slightly relaxed, narration)Celtic and Rangers used to compare themselves to kind of top kind of (low mumble) um European clubs, you know, your Ajax, Porto, that used to be kind of the level. That was the peer group. And now it's more like your Copenhagen's.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, sexual lust, infatuation; style: narration, whispered; good recording, no background noise; mildly explicit content; genuineness 2.1/6; vocal-burst blend 4.8/10; 12.3s, EN.
635231_00050760 · in -31.3 dBFS · gain +11.3 dB · podcast-04137
(slow, very low-energy, neutral tension, casual)(low mumble) Um, your Club Bruges. And but actually, you know, the m longer this goes on, the the lower that quality of club is. And it's such a slow (contented sigh) sometimes these things happen so slowly.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is slightly warm, neutral-bright, slightly rough, full; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 4.2/10; 13.5s, EN.
635231_00052040 · in -29.1 dBFS · gain +9.1 dB · podcast-03756
Contemplation ↓ / Affection ↑k-PXR-k4 · #6
This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.25.
At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.24, then +0.02 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 29 s · en · emolia
k 4d_a -0.391d_b 0.252step_a 0.159step_b 0.236min_cos_consec 0.8185min_cos_anchor 0.8181dataset emolialang enspeaker EN_wYThp1HpUwutrack EN_wYThp1HpUwutotal 28.7slevel spread 5.8 dBmax seam 4.3 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, slightly relaxed
(contemplation, concentration, interest · normal-paced, normally alert, fairly steady, monologue)The next question I'd like you to think about is what is the problem of the story and how did that problem get solved?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, concentration, interest; style: monologue, ASMR; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.1/10; 6.8s, EN.
EN_wYThp1HpUwu_W000163 · in -21.4 dBFS · gain +1.4 dB · emolia-00925
(measured, normally alert, fairly steady, monologue)When I've asked this question of other students who have heard the story,
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.0/10; 4.5s, EN.
EN_wYThp1HpUwu_W000164 · in -17.0 dBFS · gain -3.0 dB · emolia-00925
(sadness, distress, disappointment·normal-paced, normally alert, steady, formal)Other kids often say that the problem of the story was that it was really difficult to get to the North Pole.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sadness, distress, disappointment; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.5/10; 6.3s, EN.
EN_wYThp1HpUwu_W000165 · in -17.0 dBFS · gain -3.0 dB · emolia-00925
(slow, very low-energy, fairly steady, ASMR)And the way that it was solved was with (low mumble) Matthew Henson's, uh, (low mumble) his, (low mumble) uh, resourcefulness and the rest of the group's teamwork.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, submissive, neutral openness; no dominant emotion; style: ASMR, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.0/10; 10.6s, EN.
EN_wYThp1HpUwu_W000166 · in -15.6 dBFS · gain -4.4 dB · emolia-00925
Emotional Numbness ↓ / Disgust ↑k-PXR-k4 · #7
This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.44.
At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 38 s · en · emolia
k 4d_a -0.298d_b 0.436step_a 0.230step_b 0.233min_cos_consec 0.9601min_cos_anchor 0.9387dataset emolialang enspeaker EN_Ps4Ps7rMTwktrack EN_Ps4Ps7rMTwktotal 37.8slevel spread 1.7 dBmax seam 1.2 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness · steady, formal, newsreading)Mechanical engineering, welding, electrical machinery, control systems, electric circuits, engine room simulators and graphics
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 7.4s, EN.
EN_Ps4Ps7rMTwk_W000094 · in -15.8 dBFS · gain -4.2 dB · emolia-00872
(fairly steady, newsreading, formal)Bronze statues of Fulton and Christopher Columbus represent commerce on the balustrade of the galleries of the main Reading Room in the Thomas Jefferson Building of the Library of Congress on Capitol Hill in Washington, D.C.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 11.6s, EN.
EN_Ps4Ps7rMTwk_W000095 · in -15.3 dBFS · gain -4.7 dB · emolia-00872
(awe· fairly steady, formal, authoritative)They are two of sixteen historical figures, each pair representing one of the eight pillars of civilization
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 5.7s, EN.
EN_Ps4Ps7rMTwk_W000096 · in -14.1 dBFS · gain -5.9 dB · emolia-00872
(fairly steady, newsreading, formal)The Guatemalan government in 1910 erected a bust of Fulton in one of the parks of Guatemala City.In 2006, he was inducted into the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.7s, EN.
EN_Ps4Ps7rMTwk_W000097 · in -15.0 dBFS · gain -5.0 dB · emolia-00872
Triumph ↓ / Intoxication Altered States of Consciousness ↑k-PXR-k4 · #8
This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Intoxication Altered States of Consciousness around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.35.
At the same time Triumph goes the other way, from 0.67 (higher than 67 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.15, then +0.03 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 27 s · en · emolia
k 4d_a -0.283d_b 0.350step_a 0.219step_b 0.160min_cos_consec 0.9312min_cos_anchor 0.9254dataset emolialang enspeaker EN_RxWaYFXy1bItrack EN_RxWaYFXy1bItotal 27.0slevel spread 2.4 dBmax seam 2.4 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(fairly steady, formal, casual)August 16, 1945 The Nakajima Aircraft Company changed its name to Fuji-Sonyo Co., Ltd.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 8.1s, EN.
EN_RxWaYFXy1bI_W000025 · in -14.7 dBFS · gain -5.3 dB · emolia-00596
(steady, formal, monologue)November 6, 1945 – The GHQ defined Fuji-Sonyo as the Zebatsu and decided to disband them
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.2s, EN.
EN_RxWaYFXy1bI_W000026 · in -16.0 dBFS · gain -4.0 dB · emolia-00596
(steady, monologue, formal)May 1950 Fuji-San-Yoko, Limited was disbanded
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.6/10; 4.4s, EN.
EN_RxWaYFXy1bI_W000027 · in -13.6 dBFS · gain -6.4 dB · emolia-00596
(fairly steady, formal, monologue)July 1950 Fuji-Sung Yoko, Ltd. was divided into 12 companies
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.2/10; 5.9s, EN.
EN_RxWaYFXy1bI_W000028 · in -14.2 dBFS · gain -5.8 dB · emolia-00596
Concentration ↓ / Interest ↑k-PXR-k4 · #9
This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Interest clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.
At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.01, then +0.11, then +0.18 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.33 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.33, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, measured, slightly relaxed
(concentration, disappointment, jealousy and envy · subdued, steady, little disfluency, monologue)Jeg synes, at grisene burde have deres egen lov. Efter det her lovforslag har hundene fortsat deres egen lov – jeg synes faktisk, vi skulle fastholde, at de 30 millioner grise også skulle have en lov. På den baggrund kan Enhedslisten ikke støtte lovforslaget, men vi vil meget gerne være med til at arbejde for mere dyrevelfærd.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disappointment, jealousy and envy; style: monologue, formal; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.3/10; 17.1s, DA.
denmark_20181M077_2019-03-28_1000_15950048_15967136 · in -25.1 dBFS · gain +5.1 dB · eurospeech-00360
(concentration, thankfulness gratitude·normally alert, fairly steady, frequent disfluency, monologue)Tak for det. Jeg har bare et enkelt spørgsmål. Ordføreren nævnte, at vi har en alt for stor animalsk produktion i Danmark. Hvor meget skal den reduceres efter Enhedslistens opfattelse? Er det en fjernelse af den eksportorienterede del,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, thankfulness gratitude; style: monologue, didactic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.2/10; 16.2s, DA.
denmark_20181M077_2019-03-28_1000_15967136_15983360 · in -25.9 dBFS · gain +5.8 dB · eurospeech-00360
(triumph, pride, concentration · normally alert, fairly steady, some disfluency, monologue)(ahem) I øjeblikket bruger man 80 pct. af landbrugsarealet til den animalske produktion. Vi går ind for, at man reducerer landbrugsarealet med 500.000 ha, så vi får 200.000 ha mere natur, 100.000 ha mere skov,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, pride, concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.8/10; 13.1s, DA.
denmark_20181M077_2019-03-28_1000_15995920_16008976 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00360
(interest· normally alert, fairly steady, some disfluency, monologue)og gerne omstiller (low mumble) til nogle energiafgrøder (low mumble) i landbruget, og det kunne så fortsat være landbrugsarealer. Vi ser gerne, at der er nogle organogene jorder, 100.000 ha, som tages ud af omdrift, og det kan så stadig væk
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.8/10; 12.8s, DA.
denmark_20181M077_2019-03-28_1000_16008976_16021776 · in -26.2 dBFS · gain +6.2 dB · eurospeech-00360
Concentration ↓ / Infatuation ↑k-PXR-k4 · #10
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.37.
At the same time Concentration goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.10, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 45 s · en · emolia
k 4d_a -0.264d_b 0.369step_a 0.192step_b 0.142min_cos_consec 0.9535min_cos_anchor 0.9551dataset emolialang enspeaker EN_kp2oAu11ts8track EN_kp2oAu11ts8total 45.0slevel spread 1.2 dBmax seam 1.2 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration · steady, almost no disfluency, newsreading, formal)As the same compiler is available for all of the above operating systems, there is no need for recoding to produce identical products for different platforms, except when operating system dependent features are used. Cross compiling is supported with Ming-W.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.8s, EN.
EN_kp2oAu11ts8_W000022 · in -14.8 dBFS · gain -5.2 dB · emolia-02415
(fairly steady, almost no disfluency, authoritative, newsreading)Under Microsoft Windows, Harbor is more stable but less well documented than Clipper, but has multi-platform capability and is more transparent, customizable and can run from a USB flash drive
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.4s, EN.
EN_kp2oAu11ts8_W000023 · in -15.0 dBFS · gain -5.0 dB · emolia-02415
(fairly steady, no disfluency, formal, monologue)Under Linux and Windows Mobile, Clipper source code can be compiled with Harbor with very little adaptation
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.5s, EN.
EN_kp2oAu11ts8_W000024 · in -13.9 dBFS · gain -6.1 dB · emolia-02415
(steady, no disfluency, formal, newsreading)Most software originally written to run on XBase++, Flagship, FoxPro, X Harbor and others dialects can be compiled with Harbor with some adaptation
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.8s, EN.
EN_kp2oAu11ts8_W000025 · in -15.1 dBFS · gain -4.9 dB · emolia-02415
This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Elation clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.26.
At the same time Thankfulness Gratitude goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.11, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.43 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.33 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.43, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · fast, neutral tension, moderately variable, wide pitch range
(thankfulness gratitude, embarrassment, amusement · normally alert, frequent disfluency, somewhat unclear, casual)su opinión. Y si encima todo esto le juntas con gente que no juega en serio, que se lo está pasando, vamos, que empiezan con coña y tal, pues fue muy divertido.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, embarrassment, amusement; style: casual, playful; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 8.8/10; 10.8s, ES.
115744_00675560 · in -22.6 dBFS · gain +2.6 dB · podcast-05381
(relief, disappointment, pride· normally alert, frequent disfluency, somewhat unclear, casual)Pues la vez. Y no sé, impresiones de esto, pues que estuvo divertido. La verdad es que yo no sobreviví. Me llevé a uno por delante. (low mumble) Y no sobreviví.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, very dark, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as relief, disappointment, pride; style: casual, playful; below-average recording, some background noise; genuineness 5.2/6; vocal-burst blend 9.3/10; 10.8s, ES.
115744_00676736 · in -21.3 dBFS · gain +1.3 dB · podcast-05387
(elation, embarrassment, pride ·energised, some disfluency, average clarity, casual)Sí, no, nosotros hicimos una conga y quitándome a mí que yo me aparte un poco del fregado y me metí al asteroide para pillar unos preciosos torpedos de plasma. Torpedos de plasma que...
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation, embarrassment, pride; style: casual, dramatic; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.2/10; 9.2s, ES.
115744_00690656 · in -22.9 dBFS · gain +2.9 dB · podcast-05397
(elation · energised, some disfluency, somewhat unclear, casual)Bueno, pues con esta. Sí, con esta primera fase, pues al final, pues por puntos, se hicieron los cruces para la seconda fase. A 100 (surprised gasp) puntos, pero la lista tenía que ser, como bien hemos dicho antes, una lista de solo tres naves.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation; style: casual, dramatic; below-average recording, some background noise; genuineness 4.8/6; vocal-burst blend 10.0/10; 16.2s, ES.
115744_00692552 · in -22.5 dBFS · gain +2.5 dB · podcast-00668
This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Thankfulness Gratitude clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.36.
At the same time Emotional Numbness goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then -0.07, then +0.20 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 55 s · en · emolia
k 4d_a -0.380d_b 0.363step_a 0.215step_b 0.233min_cos_consec 0.8804min_cos_anchor 0.8548dataset emolialang enspeaker EN_B00040_S03013track EN_B00040_S03013total 55.0slevel spread 1.8 dBmax seam 1.7 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · fairly smooth, balanced body, energised, average clarity, wide pitch range
(normal-paced, slightly relaxed, fairly steady, casual)Caught a slant was able to bring it in for a touchdown, but also the backup quarterback PJ Walker really highlights as a low light. Had a couple interceptions throughout this training camp.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.5/10; 11.6s, EN.
EN_B00040_S03013_W000032 · in -24.0 dBFS · gain +4.0 dB · emolia-01016
(disappointment, shame, fatigue exhaustion· normal-paced, neutral tension, moderately variable, casual)Four passes and poor decisions today and really didn't look to be a backup at like the quality backup that we thought we were getting. Yet again, one day, one practice, just calling it out that he didn't look too great. (low mumble) Um, also I guess a highlight.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; reads as disappointment, shame, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 6.5/10; 16.9s, EN.
EN_B00040_S03013_W000033 · in -24.1 dBFS · gain +4.1 dB · emolia-01016
(disgust, amusement, astonishment surprise·brisk, slightly tense, moderately variable, casual)(ahem) Uh, cheese claypool pancake TJ Edwards. And when he did that, he went over him and said, go to sleep and flex time. That's, that's pretty awesome.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as disgust, amusement, astonishment surprise; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 2.6/6; vocal-burst blend 1.4/10; 8.5s, EN.
EN_B00040_S03013_W000034 · in -22.3 dBFS · gain +2.3 dB · emolia-01016
(thankfulness gratitude, embarrassment, amusement · brisk, neutral tension, moderately variable, casual)(ahem) At the same time, dude, it's practice. Like I appreciate it, but let's not injure our linebacker out here. So I appreciate the frustration. And he was probably frustrated because the offense wasn't doing anything, but still like the tenacity. And that's what he brings as a former tight end that does play wide receiver.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, embarrassment, amusement; style: casual, playful; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.1/10; 17.6s, EN.
EN_B00040_S03013_W000035 · in -22.3 dBFS · gain +2.3 dB · emolia-01016
This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Thankfulness Gratitude around average — 0.46, lower than 54 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.43.
At the same time Fatigue Exhaustion goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.23, then +0.17 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.47 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.47 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.47, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 22 s · en · emolia
k 4d_a -0.464d_b 0.427step_a 0.175step_b 0.229min_cos_consec 0.4674min_cos_anchor 0.4674dataset emolialang enspeaker EN_PpJGxcVW_IEtrack EN_PpJGxcVW_IEtotal 21.8slevel spread 2.4 dBmax seam 1.5 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, moderate pitch range, light breath
(normal-paced, fairly steady, some disfluency, casual)If you've got humans in your image, got one here, quite far away, you can actually...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; no dominant emotion; style: casual, storytelling; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.7/10; 3.8s, EN.
EN_PpJGxcVW_IE_W000069 · in -16.3 dBFS · gain -3.7 dB · emolia-01183
(awe, pain·measured, steady, frequent disfluency, casual)This tool is brilliant. You can actually...
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pain; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.7/10; 3.0s, EN.
EN_PpJGxcVW_IE_W000071 · in -17.3 dBFS · gain -2.7 dB · emolia-01183
(relief·normal-paced, fairly steady, some disfluency, monologue)Take the sky reflection and put it in the water, because it would look quite strange if you didn't have the sky reflection in nice smooth water.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: monologue, didactic; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.7/10; 7.4s, EN.
EN_PpJGxcVW_IE_W000072 · in -18.7 dBFS · gain -1.3 dB · emolia-01183
(normal-paced, fairly steady, little disfluency, monologue)And for landscape photographers again, you'll love this. You can actually add a water blur in. Once again, we have no water, but you can bring this up.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.0/10; 7.1s, EN.
EN_PpJGxcVW_IE_W000073 · in -17.3 dBFS · gain -2.7 dB · emolia-01183
Concentration ↓ / Confusion ↑k-PXR-k4 · #14
This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Confusion below average — 0.34, lower than 66 % of clips in this corpus — and ends with it clearly present at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.40.
At the same time Concentration goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.00, then +0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 35 s · zh · emolia
k 4d_a -0.288d_b 0.397step_a 0.248step_b 0.249min_cos_consec 0.9182min_cos_anchor 0.9020dataset emolialang zhspeaker ZH_B00009_S06684track ZH_B00009_S06684total 35.4slevel spread 2.3 dBmax seam 2.3 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.32.
At the same time Emotional Numbness goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.08, then +0.04, then +0.19 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · en · emolia
k 4d_a -0.262d_b 0.319step_a 0.247step_b 0.190min_cos_consec 0.8294min_cos_anchor 0.8294dataset emolialang enspeaker EN_UmybJnfTuqctrack EN_UmybJnfTuqctotal 32.6slevel spread 3.4 dBmax seam 3.3 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · fairly smooth, slightly relaxed, clear, light breath
(emotional numbness, relief · measured, subdued, steady, whispered)And finally, do not leave cells blank or merge cells.
full caption & clip details
A child feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly cool, dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, relief; style: whispered, monologue; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 4.6s, EN.
EN_UmybJnfTuqc_W000099 · in -19.7 dBFS · gain -0.3 dB · emolia-01972
(normal-paced, normally alert, fairly steady, casual)Tables should be used for organizing data.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.2/10; 3.0s, EN.
EN_UmybJnfTuqc_W000100 · in -16.4 dBFS · gain -3.5 dB · emolia-01972
(normal-paced, normally alert, steady, whispered)Keeping the tables as simple as possible and avoid nesting tables. So tables within tables.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.0/10; 6.9s, EN.
EN_UmybJnfTuqc_W000101 · in -17.9 dBFS · gain -2.1 dB · emolia-01972
(concentration· normal-paced, normally alert, fairly steady, formal)First, you'll notice that the title is part of the table and it has been merged across multiple cells. Secondly, there are also merged cells in the table. This is a better example of the same information. Notice that the title has been taken out of the table.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.7/10; 17.6s, EN.
EN_UmybJnfTuqc_W000103 · in -19.8 dBFS · gain -0.2 dB · emolia-01972
Sourness ↓ / Emotional Numbness ↑k-PXR-k4 · #16
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.45.
At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.17, then +0.18 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 57 s · fr · emolia
k 4d_a -0.321d_b 0.454step_a 0.207step_b 0.183min_cos_consec 0.8746min_cos_anchor 0.8267dataset emolialang frspeaker FR_JThu7Jf0OeYtrack FR_JThu7Jf0OeYtotal 57.5slevel spread 1.8 dBmax seam 1.8 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(sourness, shame, anger · normal-paced, some disfluency, somewhat unclear, monologue)Je vous laisse juger des contenus respectifs des chaînes de ces deux auteurs, et je vous laisse me démontrer que Mademoiselle Wotta est deux fois plus pertinente, ou deux fois plus dissidente, ou deux fois plus cohérente, ou deux fois plus intéressante que, (low mumble) euh, Monsieur Rougeron. Pour le, ah, pour ma part, je n'en suis pas totalement, totalement, (low mumble) euh, convaincu.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, shame, anger; style: monologue, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 3.7/10; 20.0s, FR.
FR_JThu7Jf0OeY_W000071 · in -20.4 dBFS · gain +0.4 dB · emolia-02831
(awe, interest· normal-paced, some disfluency, somewhat unclear, monologue)en histoire, euh, (low mumble) j'avais pour une fois, abondance de, de bien, abondance de choix, donc il paraît que ça ne nuit pas. Donc, nous avons deux femmes contre un homme, les, nous avons Charlie, je sais pas quoi, des Revues du Monde, et (low mumble) euh, l'autre, je sais plus comment elle s'appelle, de, euh, c'est une autre histoire.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, interest; style: monologue, authoritative; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.9/10; 16.9s, FR.
FR_JThu7Jf0OeY_W000072 · in -19.0 dBFS · gain -1.0 dB · emolia-02831
(normal-paced, some disfluency, average clarity, monologue)392 000 abonnés, et pour c'est une autre histoire, 1441 euros mensuels pour 163 900 abonnés, ce qui représente donc 0,2 centimes par abonné pour les revues du monde, 0,9 centimes par abonné
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.9/10; 13.0s, FR.
FR_JThu7Jf0OeY_W000073 · in -20.8 dBFS · gain +0.8 dB · emolia-02831
(emotional numbness·measured, frequent disfluency, somewhat unclear, didactic)pour, c'est une autre histoire, et seulement pour l'homme, 0,1 centime. Donc, un rapport de fois deux.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.8/10; 7.2s, FR.
FR_JThu7Jf0OeY_W000074 · in -19.8 dBFS · gain -0.2 dB · emolia-02831
Embarrassment ↓ / Shame ↑k-PXR-k4 · #17
This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Shame clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.
At the same time Embarrassment goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then -0.02, then +0.05 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, average clarity, moderate pitch range
(embarrassment · slightly relaxed, fairly steady, frequent disfluency, casual)one that jumps out to me immediately is I was talking to Malcolm Gladwell, (low mumble) uh, best-selling author, podcaster, uh
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.8/10; 7.8s, EN.
530226_00065592 · in -26.8 dBFS · gain +6.8 dB · podcast-05038
(contemplation, infatuation, shame·relaxed, fairly steady, frequent disfluency, casual)he I had asked him because I (low mumble) uh, you know, as a as a peer of his, I was curious how he filters his projects. How does he think of what he should and shouldn't do? What is a Malcolm Gladwell project to him? And
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, infatuation, shame; style: casual, conversational; average recording, no background noise; genuineness 4.7/6; vocal-burst blend 8.7/10; 15.0s, EN.
530226_00066644 · in -27.3 dBFS · gain +7.3 dB · podcast-02992
(contemplation, doubt, affection·neutral tension, moderately variable, some disfluency, casual)because you know, Malcolm Gladwell has a brand and he has a voice and he does things that kind of feel distinctively, Malcolm Gladwell. And I was curious how he defined that for himself. And he said he doesn't do that, he doesn't think of himself uh (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, doubt, affection; style: casual, conversational; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.6/10; 12.7s, EN.
530226_00068144 · in -26.2 dBFS · gain +6.2 dB · podcast-02973
(shame, infatuation, concentration·slightly relaxed, fairly steady, some disfluency, casual)as as having one particular kind of voice. He doesn't think of himself as a brand. And then he said, he said this line that I immediately jotted down. He said, self-conceptions are powerfully limiting, which is to say that if you have some particular vision of what you are, some particular definition of what you are. Well, then you're going to close off all of these other opportunities that don't fit that narrow window.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, infatuation, concentration; style: casual, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 5.3/10; 24.2s, EN.
530226_00069416 · in -26.2 dBFS · gain +6.2 dB · podcast-06326
Concentration ↓ / Pride ↑k-PXR-k4 · #18
This chain comes from the proxy rule: the same two-sided test as above, but because Pride is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pride clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.
At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.22, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-bright, balanced body, quiet background
(concentration, bitterness, sadness · measured, energised, neutral tension, cartoonish)Landet har stort sett vært i konflikt og i en krigslignende tilstand. Det har få venner i omverdenen. Jeg mener det er viktig at Norge deltar aktivt i internasjonalt samarbeid for å hindre en total statskollaps i Eritrea,
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, bitterness, sadness; style: cartoonish, storytelling; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.7/10; 15.0s, NO.
norway_9531-1_17121424_17136384 · in -32.5 dBFS · gain +12.5 dB · eurospeech-02334
(disappointment, impatience and irritability, concentration ·normal-paced, normally alert, slightly relaxed, monologue)med de følger det kan få for en allerede hardt prøvet befolkning, og den (ahem) destabilisering det kan bety for den konfliktrammede regionen rundt Afrikas Horn. Det er ikke naturlig med noe omfattende tosidig utviklingssamarbeid med et sånt regime,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, impatience and irritability, concentration; style: monologue, storytelling; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 4.9/10; 14.7s, NO.
norway_9531-1_17136384_17151088 · in -31.9 dBFS · gain +11.9 dB · eurospeech-02334
(pride, shame, triumph·measured, normally alert, slightly relaxed, narration)men når mennesker er i desperat behov for nødhjelp, får vi og det internasjonale samfunn behov for å stille opp. Så får vi støtte politiske prosesser for å bringe utviklinga inn på en bedre vei.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, shame, triumph; style: narration, storytelling; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 4.1/10; 12.9s, NO.
norway_9531-1_17151088_17164032 · in -32.1 dBFS · gain +12.1 dB · eurospeech-02334
(pride ·normal-paced, normally alert, neutral tension, dramatic)en bedre vei. Jeg mener situasjonen i Eritrea også mye viser noen av de dilemmaene vi har i internasjonal politikk. Menneskerettigheter brytes i en rekke land,
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is positive, neutral stance, slightly guarded; reads as pride; style: dramatic, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.5/10; 10.5s, NO.
norway_9531-1_17164032_17174528 · in -31.6 dBFS · gain +11.6 dB · eurospeech-02334
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.29.
At the same time Emotional Numbness goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.04, then +0.22, then +0.11 — not a clean run: step 1 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · en · emolia
k 4d_a -0.489d_b 0.291step_a 0.197step_b 0.219min_cos_consec 0.8962min_cos_anchor 0.9108dataset emolialang enspeaker EN_Cf4SW2WKIVktrack EN_Cf4SW2WKIVktotal 30.2slevel spread 3.0 dBmax seam 3.0 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · slightly cool, neutral-bright, rough, highly aroused, almost no disfluency, wide pitch range
(emotional numbness · normal-paced, slightly relaxed, fairly steady, authoritative)Unified system of separation, retirement and pension.
full caption & clip details
An adult masculine voice; delivery is highly aroused, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, rough, thin; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, dominant, fairly guarded; reads as emotional numbness; style: authoritative, formal; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 3.5s, EN.
EN_Cf4SW2WKIVk_W000449 · in -20.2 dBFS · gain +0.2 dB · emolia-00567
(pride· normal-paced, tense, moderately variable, authoritative)This grants a monthly disability pension in Liu.
full caption & clip details
An adult masculine voice; delivery is highly aroused, normal-paced, tense, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, almost no disfluency, wide pitch range, audible breath; affect is neutral, dominant, fairly guarded; reads as pride; style: authoritative, dramatic; below-average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.1/10; 3.5s, EN.
EN_Cf4SW2WKIVk_W000450 · in -18.5 dBFS · gain -1.5 dB · emolia-00567
(malevolence malice·measured, slightly relaxed, fairly steady, authoritative)It promotes the use of internet, internet, and other ICT to provide opportunities for citizens.
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, rough, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as malevolence malice; style: authoritative, dramatic; below-average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 11.0s, EN.
EN_Cf4SW2WKIVk_W000452 · in -21.6 dBFS · gain +1.6 dB · emolia-00567
(concentration, malevolence malice ·brisk, neutral tension, moderately variable, authoritative)This will provide for a rational and holistic management and development of our country's land and water resources. Hold owners accountable for making these lands productive and sustainable.
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as concentration, malevolence malice; style: authoritative, dramatic; below-average recording, some background noise; genuineness 0.3/6; vocal-burst blend 0.7/10; 11.7s, EN.
EN_Cf4SW2WKIVk_W000453 · in -20.3 dBFS · gain +0.3 dB · emolia-00567
Relief ↓ / Emotional Numbness ↑k-PXR-k4 · #20
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.41.
At the same time Relief goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.13, then +0.05, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 62 s · en · emolia
k 4d_a -0.469d_b 0.414step_a 0.210step_b 0.229min_cos_consec 0.9007min_cos_anchor 0.9007dataset emolialang enspeaker EN_QgwpVXZ5liutrack EN_QgwpVXZ5liutotal 62.2slevel spread 4.1 dBmax seam 4.1 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, slightly relaxed, steady
(relief · subdued, light breath, didactic, monologue)And this delta positive charge remains here and so one alcohol group will leave.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: didactic, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.5/10; 6.7s, EN.
EN_QgwpVXZ5liu_W000203 · in -19.2 dBFS · gain -0.8 dB · emolia-00832
(triumph, sexual lust, infatuation· subdued, light breath, monologue, didactic)Creating one new SiOH linkage. This acid catalyzed reactions normally take place at pH of less than 2.2 and it has a fast protonation step and the silicon becomes electrophilic, uh, (low mumble) after the protonation and therefore, is more susceptible to attack by water and the protonation becomes slower
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, sexual lust, infatuation; style: monologue, didactic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.4/10; 27.6s, EN.
EN_QgwpVXZ5liu_W000204 · in -23.3 dBFS · gain +3.3 dB · emolia-00832
(concentration· subdued, normal breath, didactic, monologue)Now in basic conditions the OH minus group attacks the silane, (low mumble) uh, tetra alkoxide silane and you get this kind of a (low mumble) OH delta minus charge here and
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 16.5s, EN.
EN_QgwpVXZ5liu_W000206 · in -22.0 dBFS · gain +2.0 dB · emolia-00832
(emotional numbness, pain·very low-energy, normal breath, didactic, monologue)So, this also gets a OR delta minus chart and then this OR minus leaves and you are left with a new
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.0/10; 11.0s, EN.
EN_QgwpVXZ5liu_W000207 · in -21.9 dBFS · gain +1.9 dB · emolia-00832