PXR: 2 chains from each of the 12 scarcest ordered emotion pairs (supply 1-3 chains each).
Rule.PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable Source. trajectories_v5.parquet | Family. the scarcest cells -- the ones the owner said matter most Sampled from 338,690 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Sadness ↓ / Distress ↑rare-PXR-pairs · #1
This chain comes from the proxy rule: the same two-sided test as above, but because Distress is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Distress around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.50.
At the same time Sadness goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.43 (lower than 57 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.50 — a slow start, with most of the change arriving in the final step.
The largest step is 0.50, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 36 s · en · emolia
hear it un-normalised (raw levels, max seam 0.5 dB)
k 3d_a -0.456d_b 0.499step_a 0.456step_b 0.499min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_hFWfzPwHW_Etrack EN_hFWfzPwHW_Etotal 36.0slevel spread 0.5 dBmax seam 0.5 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, thin, no background noise, normal-paced, normally alert, slightly relaxed, clear
(fairly steady, almost no disfluency, formal, newsreading)This included 32% of Latinos, 29% of African Americans, and almost nobody with disabilities.A 2015 populist poll in the United Kingdom found broad public support for assisted dying.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 14.7s, EN.
EN_hFWfzPwHW_E_W000180 · in -16.0 dBFS · gain -4.0 dB · emolia-02446
(emotional numbness· fairly steady, no disfluency, formal, newsreading)82% of people supported the introduction of assisted dying laws, including 86% of people with disabilities.One concern is that euthanasia might undermine filial responsibility.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.7s, EN.
EN_hFWfzPwHW_E_W000181 · in -15.5 dBFS · gain -4.5 dB · emolia-02446
(distress, disgust·steady, almost no disfluency, formal, newsreading)In some countries, adult children of impoverished parents are legally entitled to support payments under filial responsibility laws
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, disgust; style: formal, newsreading; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 8.4s, EN.
EN_hFWfzPwHW_E_W000182 · in -15.7 dBFS · gain -4.3 dB · emolia-02446
Distress ↓ / Sadness ↑rare-PXR-pairs · #2
This chain comes from the proxy rule: the same two-sided test as above, but because Sadness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sadness around average — 0.43, lower than 57 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.46.
At the same time Distress goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.50, then -0.50, then +0.46 — not a clean run: step 3 moves back the other way by 0.50 before the chain recovers.
The largest step is 0.50, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.34 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.34 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 45 s · fr · emolia
hear it un-normalised (raw levels, max seam 5.8 dB)
k 5d_a -0.520d_b 0.460step_a 0.540step_b 0.497min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_GvSCBxz9394track FR_GvSCBxz9394total 44.9slevel spread 6.0 dBmax seam 5.8 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, fairly steady, moderate pitch range, light breath
(distress, disgust, confusion · fast, slightly relaxed, some disfluency, monologue)Malgré que tu as l'autorisation de son, de son chef, deux fois il peut te dire, (ahem) euh, qu'est-ce qui prouve que c'est, c'est, c'est signé par moi chef. Peut-être que tu l'as falsifié.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, disgust, confusion; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.9/10; 7.5s, FR.
FR_GvSCBxz9394_W000147 · in -18.5 dBFS · gain -1.5 dB · emolia-02762
(impatience and irritability, anger, fear· fast, slightly relaxed, some disfluency, dramatic)(ahem) euh, pourquoi, tu veux, donc, ils sont toujours prêts parce que c'est des jeunes qui, (ahem) euh, qui cherchent toujours à se battre, ceux qui attaquent, qui attaquent un peu ces policiers.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, anger, fear; style: dramatic, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.7/10; 6.6s, FR.
FR_GvSCBxz9394_W000148 · in -20.0 dBFS · gain -0.0 dB · emolia-02762
(disappointment, distress, helplessness· fast, neutral tension, some disfluency, monologue)Donc, puisque je n'arrivais pas, je n'arrive pas, si on n'arrive pas à trouver, à capter la réalité directement, je pense que (ahem) chercher à reconstruire l'histoire et en gardant un peu cette partie de, de réalité, en laissant la liberté, quelque chose de plus sorti, je pense que c'est, c'est aussi intéressant.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, distress, helplessness; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 10.0/10; 15.5s, FR.
FR_GvSCBxz9394_W000149 · in -19.8 dBFS · gain -0.2 dB · emolia-02762
(thankfulness gratitude, affection, relief·normal-paced, slightly relaxed, frequent disfluency, casual)Pour laquelle je voulais te remercier, c'est que je trouve que t'es, enfin.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as thankfulness gratitude, affection, relief; style: casual, conversational; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 2.1/10; 4.1s, FR.
FR_GvSCBxz9394_W000151 · in -13.9 dBFS · gain -6.0 dB · emolia-02762
(normal-paced, slightly relaxed, some disfluency, casual)des documentaires sur la, la jeunesse, il y en a un qui sont valorisants et qui nous montrent que les, mais des, la jeunesse de n'importe quel pays que, qu'en fait les jeunes sont pas ce qu'on dit en permanence qu'ils sont.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.5/10; 10.6s, FR.
FR_GvSCBxz9394_W000152 · in -16.6 dBFS · gain -3.4 dB · emolia-02762
Elation ↓ / Amusement ↑rare-PXR-pairs · #3
This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.64.
At the same time Elation goes the other way, from 0.69 (higher than 69 % of clips in this corpus) to 0.96 (higher than 96 % of clips in this corpus), a change of +0.28. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.58, then +0.00, then +0.06 — most of the change happening immediately, then levelling off.
The largest step is 0.58, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.57 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.57, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 36 s · es · podcast
hear it un-normalised (raw levels, max seam 1.0 dB)
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fast, moderately variable, some disfluency, wide pitch range
(normally alert, slightly relaxed, average clarity, casual)Aparece en la película y de hecho la parte más pesada que aparece en el festival, si recuerdan, eran Riff
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, dramatic; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 6.3/10; 6.6s, ES.
21265_00298640 · in -22.7 dBFS · gain +2.7 dB · podcast-05115
(disgust, teasing, embarrassment·energised, neutral tension, somewhat unclear, casual)y los debutantes B8. Exacto. De hecho, (surprised gasp) lo que sale B8 es como ahí un compilado de imagen, ¿no? Se ve que era muy fuerte para el video.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disgust, teasing, embarrassment; style: casual, playful; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 10.0/10; 11.8s, ES.
21265_00299312 · in -21.7 dBFS · gain +1.7 dB · podcast-05143
(relief, jealousy and envy, affection·normally alert, neutral tension, somewhat unclear, casual)claro, exactamente en el video, apadrinados por el gran papo napolitano. A mí me gusta mucho que se encuentra en YouTube un fragmento grabado por un seguidor, no sé, que de Ricardo Lloro diciendo Parca Sangriento. No, no dejan tocar más, dice
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as relief, jealousy and envy, affection; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 10.0/10; 12.2s, ES.
21265_00300696 · in -22.7 dBFS · gain +2.7 dB · podcast-05114
(amusement, teasing, pleasure ecstasy·energised, neutral tension, average clarity, casual)esto de Barro. Parca sangrienta de los hippies que se mueren. Histórica. Exactamente.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 3.6/10; 5.5s, ES.
21265_00301912 · in -22.2 dBFS · gain +2.2 dB · podcast-00370
Distress ↓ / Bitterness ↑rare-PXR-pairs · #4
This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Bitterness clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Distress goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.07, then +0.10 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 55 s · el · eurospeech
hear it un-normalised (raw levels, max seam 2.3 dB)
k 4d_a -0.538d_b 0.258step_a 0.479step_b 0.100min_cos_consec —min_cos_anchor —dataset eurospeechlang elspeaker greece_olomeleia-20210128btrack greece_olomeleia-20210128btotal 54.5slevel spread 2.7 dBmax seam 2.3 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · slightly cool, wide pitch range
(distress, anger, disgust · brisk, normally alert, slightly relaxed, authoritative)Είναι μια δύσκολη περίοδος, αλλά μέσα από τις δυσκολίες θα βγούμε πιο ενωμένοι, ισχυρότεροι και πιο προσηλωμένοι στον στόχο και καθώς ο εμβολιασμός θα προχωρά, θα διαπιστώνουμε όλοι ότι το εμβόλιο είναι ασφαλές
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as distress, anger, disgust; style: authoritative, dramatic; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 2.3/10; 13.0s, EL.
greece_olomeleia-20210128b_3745776_3758768 · in -19.5 dBFS · gain -0.5 dB · eurospeech-00778
(malevolence malice, disgust, impatience and irritability·normal-paced, normally alert, neutral tension, authoritative)και όλοι θα συνειδητοποιήσουν πως αυτό είναι η μόνη λύση. Το εμβόλιο είναι ο μόνος τρόπος για να ξαναπάρουμε τη ζωή στα χέρια μας. Σας ευχαριστώ.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, disgust, impatience and irritability; style: authoritative, dramatic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.3/10; 11.0s, EL.
greece_olomeleia-20210128b_3758768_3769760 · in -20.4 dBFS · gain +0.4 dB · eurospeech-00778
(contempt, anger, sadness·measured, very low-energy, neutral tension, monologue)αν νομίζετε ότι στη φάση που βρίσκεται σήμερα αυτή η πολλαπλή κρίση, πραγματικά η συνταγή
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is slightly cool, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, anger, sadness; style: monologue, dramatic; below-average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.8/10; 12.8s, EL.
greece_olomeleia-20210128b_3813968_3826720 · in -22.2 dBFS · gain +2.2 dB · eurospeech-00778
(bitterness, contempt, malevolence malice· measured, energised, neutral tension, dramatic)είναι να δυναμιτιστεί το πολιτικό κλίμα, να υπάρχουν κινήσεις αυταρχισμού, καταστολής, συρρίκνωσης δημοκρατικών δικαιωμάτων, εκφοβισμού και τρομοκρατίας,
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, fairly guarded; reads as bitterness, contempt, malevolence malice; style: dramatic, cartoonish; below-average recording, some background noise; genuineness 3.3/6; vocal-burst blend 4.0/10; 17.4s, EL.
greece_olomeleia-20210128b_3826720_3844080 · in -19.9 dBFS · gain -0.1 dB · eurospeech-00778
Distress ↓ / Bitterness ↑rare-PXR-pairs · #5
This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Bitterness clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.32.
At the same time Distress goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 79 s · en · podcast
hear it un-normalised (raw levels, max seam 2.0 dB)
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, neutral tension, frequent disfluency, somewhat unclear
(distress, helplessness, sadness · measured, subdued, fairly steady, casual)that I mean the Raiders are gonna struggle. And the NFC West, when Geno Smith was in it, was also a really tough division because he had the Niners, he had the Rams. All these teams in that division. It just it was a struggle. And the Raiders are gonna have to s just pull out some random wins this season if they really have a chance.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as distress, helplessness, sadness; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 4.4/10; 22.5s, EN.
628153_00097216 · in -22.1 dBFS · gain +2.0 dB · podcast-03021
(triumph, relief, jealousy and envy·normal-paced, subdued, moderately variable, casual)But it was good for them. They got a bait they got a good win away against the Patriots. Nothing really discreet. Geno Smith looked okay. If he didn't get sacked four times, he would have been I mean, I would have talked about him having really a good game. But nothing really discreet, (ahem) nothing really crazy there. But you want to talk about a game that I did not expect to be as good as it was. The Steelers Jets game, wow.
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as triumph, relief, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 8.6/10; 28.0s, EN.
628153_00099464 · in -20.0 dBFS · gain +0.0 dB · podcast-06300
(bitterness, triumph, jealousy and envy ·measured, energised, fairly steady, casual)The Pittsburgh Steelers and the New York Jets really put on at that point, probably one of the best games of the week. And honestly, it could be the game of the week up until that point. And then the Bills and Ravens happened. But anyway, Aaron Rodgers and Justin Fields put on an absolute quarterback duel. Really impressive, honestly, from both teams.
full caption & clip details
A young adult masculine voice; delivery is energised, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as bitterness, triumph, jealousy and envy; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 7.2/10; 27.9s, EN.
628153_00102256 · in -20.8 dBFS · gain +0.8 dB · podcast-06301
Anger ↓ / Distress ↑rare-PXR-pairs · #6
This chain comes from the proxy rule: the same two-sided test as above, but because Distress is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Distress around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.52.
At the same time Anger goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.96 (higher than 96 % of clips in this corpus), a change of +0.26. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.00, then +0.52 — a slow start, with most of the change arriving in the final step.
The largest step is 0.52, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · zh · emolia
k 5d_a 0.259d_b 0.524step_a 0.095step_b 0.524min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00015_S06570track ZH_B00015_S06570total 32.2slevel spread 4.3 dBmax seam 3.4 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-bright
An adult masculine voice; delivery is highly aroused, fast, tense, moderately variable; timbre is slightly cool, neutral-bright, very rough, thin; very clear, almost no disfluency, very wide pitch range, normal breath; affect is negative, dominant, guarded; reads as thankfulness gratitude, impatience and irritability, malevolence malice; style: cartoonish, ranting; poor recording, quiet background; genuineness 2.4/6; vocal-burst blend 3.1/10; 7.5s, ZH.
ZH_B00015_S06570_W000016 · in -18.5 dBFS · gain -1.5 dB · emolia-03422
(distress, impatience and irritability, anger· fast, energised, neutral tension, dramatic)那荷鲁晓夫大声就问他一下,这个是谁写的,你们谁写的,谁站出来,然后没有人回应,也没有人敢回应。
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as distress, impatience and irritability, anger; style: dramatic, cartoonish; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.1/10; 8.7s, ZH.
ZH_B00015_S06570_W000017 · in -21.4 dBFS · gain +1.4 dB · emolia-03422
Anger ↓ / Distress ↑rare-PXR-pairs · #7
This chain comes from the proxy rule: the same two-sided test as above, but because Distress is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Distress around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.47.
At the same time Anger goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.47 — a slow start, with most of the change arriving in the final step.
The largest step is 0.47, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 20 s · snippets
k 3d_a -0.259d_b 0.469step_a 0.157step_b 0.469min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch41_part1_batch41_parttrack batch41_part1_batch41_parttotal 19.9slevel spread 1.4 dBmax seam 1.1 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, fairly steady, moderate pitch range
(anger · normal-paced, some disfluency, average clarity, casual)the world the financial world and the world of people is changing the whole time. History doesn't repeat itself whereas in physics history repeats itself all the time. You can do the same experiment over and over again.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger; style: casual, narration; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 3.9/10; 9.6s.
batch41_part1_batch41_part1_chunk_1364_1_1217771 · in -24.6 dBFS · gain +4.6 dB · snippets-01104
(emotional numbness, disgust, sourness·measured, no disfluency, somewhat unclear, monologue)Some of my coworkers were self-satisfied, complacent, with a lazy state of mind.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, slightly dark, fairly smooth, balanced body; somewhat unclear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness, disgust, sourness; style: monologue, casual; good recording, no background noise; explicit content; genuineness 1.4/6; vocal-burst blend 4.2/10; 4.5s.
batch41_part1_batch41_part1_chunk_1364_1_1217790 · in -23.5 dBFS · gain +3.5 dB · snippets-01104
(distress· measured, almost no disfluency, average clarity, narration)making a social call and I remember she said she said the Fed has got to lower interest rates.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as distress; style: narration, monologue; good recording, no background noise; mildly explicit content; genuineness 1.9/6; vocal-burst blend 3.0/10; 5.5s.
batch41_part1_batch41_part1_chunk_1364_1_1217818 · in -23.2 dBFS · gain +3.2 dB · snippets-01104
Pleasure Ecstasy ↓ / Elation ↑rare-PXR-pairs · #8
This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Elation at the very top of the corpus — 0.97, higher than 97 % of clips in this corpus — and works its way down to clearly present at 0.69, higher than 69 % of clips in this corpus. That is a total fall of 0.28.
At the same time Pleasure Ecstasy goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are -0.09, then -0.19 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 31 s · en · emolia
k 3d_a -0.324d_b -0.283step_a 0.222step_b 0.192min_cos_consec 0.7233min_cos_anchor 0.7410dataset emolialang enspeaker EN_B00040_S05893track EN_B00040_S05893total 30.7slevel spread 1.4 dBmax seam 1.4 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth, balanced body, average clarity, light breath
(pleasure ecstasy, hope enthusiasm optimism, elation · brisk, energised, neutral tension, casual)I do love a laptop for its portability, but there's just something about sitting down on a desktop and diving into work that just feels more...
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.7/10; 6.9s, EN.
EN_B00040_S05893_W000006 · in -19.3 dBFS · gain -0.7 dB · emolia-01016
(fear, disappointment, pain· brisk, energised, neutral tension, casual)Formal. It feels more substantial. And that's why I built this disaster behind me. I wanted to have like the best system I could possibly have to sit down and tackle like more ambitious videos. I never really took advantage of this system. I think it's because I don't like sitting back here. I built this office because I wanted privacy, but I work alone in this space now. I don't need privacy.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is negative, slightly dominant, slightly guarded; reads as fear, disappointment, pain; style: casual, storytelling; average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 6.3/10; 20.3s, EN.
EN_B00040_S05893_W000007 · in -20.7 dBFS · gain +0.7 dB · emolia-01016
(normal-paced, normally alert, slightly relaxed, casual)So today is the day I relocate the tower setup.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.3/10; 3.2s, EN.
EN_B00040_S05893_W000008 · in -20.1 dBFS · gain +0.1 dB · emolia-01016
Pleasure Ecstasy ↓ / Elation ↑rare-PXR-pairs · #9
This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Elation at the very top of the corpus — 0.98, higher than 98 % of clips in this corpus — and works its way down to clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total fall of 0.26.
At the same time Pleasure Ecstasy goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are -0.12, then -0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.31 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.28 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.31, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult masculine voice · slightly cool, slightly rough, some background noise, normal-paced
(pleasure ecstasy, hope enthusiasm optimism, elation · normally alert, relaxed, fairly steady, casual)Falando isso, galera. É... (low mumble) Pra quem não tá ir por outra plataforma, procura lá no YouTube o nosso canal, o canal PolentaVerso, se você quer ver uns caras que não sabe jogar e é metido a gravar vídeo de gameplay, fazer stream jogo, vídeo de zoeira, tem ali o Benevolente lançou a nova linha Gameplay Selvagem,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is slightly cool, very dark, slightly rough, slightly thin; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pleasure ecstasy, hope enthusiasm optimism, elation; style: casual; below-average recording, some background noise; genuineness 5.4/6; vocal-burst blend 10.0/10; 20.6s, PT.
885628_00051824 · in -27.0 dBFS · gain +7.0 dB · podcast-00288
(disgust, confusion, thankfulness gratitude· normally alert, neutral tension, moderately variable, casual)Você pode xingar a gente, é o que a gente gosta, né? Se você não quiser dar o seu like, se inscreve no canal pra xingar a gente. Isso
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, confusion, thankfulness gratitude; style: casual, conversational; below-average recording, some background noise; genuineness 6.0/6; vocal-burst blend 5.8/10; 11.8s, PT.
885628_00055220 · in -27.7 dBFS · gain +7.7 dB · podcast-02629
(energised, fully relaxed, moderately variable, casual)parte mais específica que eu gostei foi a Saifa transportando o Castelvania.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, fully relaxed, moderately variable; timbre is slightly cool, very dark, slightly rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; poor recording, some background noise; genuineness 3.6/6; vocal-burst blend 2.2/10; 5.6s, PT.
885628_00060216 · in -29.7 dBFS · gain +9.7 dB · podcast-00419
Disappointment ↓ / Sadness ↑rare-PXR-pairs · #10
This chain comes from the proxy rule: the same two-sided test as above, but because Sadness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sadness around average — 0.43, lower than 57 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.45.
At the same time Disappointment goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.45 — a slow start, with most of the change arriving in the final step.
The largest step is 0.45, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 61 s · en · emolia
k 4d_a -0.508d_b 0.453step_a 0.508step_b 0.453min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_3_5JjrnMwXUtrack EN_3_5JjrnMwXUtotal 61.1slevel spread 1.5 dBmax seam 1.5 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(almost no disfluency, no audible breath, newsreading, formal)Manning recruited chemists from industry, universities, and government to help study mustard gas poisoning, investigate and mass-produce new toxic chemicals, and develop gas masks and other treatments.A center for chemical weapons research was established at American University in Washington, D.C. to house researchers.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 20.6s, EN.
EN_3_5JjrnMwXU_W000018 · in -16.3 dBFS · gain -3.7 dB · emolia-02468
(no disfluency, light breath, newsreading, formal)The US military paid to convert classrooms into laboratories. Within a year of setting up the center, the number of scientists and technicians employed there would increase from 272 to over 1,000.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.4s, EN.
EN_3_5JjrnMwXU_W000019 · in -14.8 dBFS · gain -5.2 dB · emolia-02468
(no disfluency, light breath, formal, newsreading)Industrial plants were established in nearby cities to synthesize toxic chemicals for use in research and armaments
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.9s, EN.
EN_3_5JjrnMwXU_W000020 · in -16.3 dBFS · gain -3.7 dB · emolia-02468
(no disfluency, no audible breath, newsreading, formal)Shells were filled with toxic gas in Edgewood, Maryland. Women were employed to produce gas masks in Long Island City.On 5 July 1917 General John J. Pershing oversaw the creation of a new military unit dealing with gas, the Gas Service Section.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.7s, EN.
EN_3_5JjrnMwXU_W000021 · in -15.3 dBFS · gain -4.7 dB · emolia-02468
Disappointment ↓ / Sadness ↑rare-PXR-pairs · #11
This chain comes from the proxy rule: the same two-sided test as above, but because Sadness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sadness around average — 0.43, lower than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.51.
At the same time Disappointment goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.51 — a slow start, with most of the change arriving in the final step.
The largest step is 0.51, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 50 s · fr · emolia
k 4d_a -0.597d_b 0.509step_a 0.453step_b 0.509min_cos_consec —min_cos_anchor —dataset emolialang frspeaker FR_XDvJOyjUUSctrack FR_XDvJOyjUUSctotal 49.6slevel spread 2.0 dBmax seam 2.0 dB
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
(disappointment, shame, bitterness · normal-paced, some disfluency, somewhat unclear, monologue)Effectivement, je suis, je valide et je m'attends également à ce que les femmes quand elles sortent de leur zone de confort et quand je les encourage à tenter de réparer quelque chose avant de le jeter ou à tenter de nettoyer ou de changer leur disque dur avant de jeter l'ordinateur ou à tenter de dévisser une prise électrique.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, shame, bitterness; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.3/10; 19.9s, FR.
FR_XDvJOyjUUSc_W000047 · in -19.2 dBFS · gain -0.8 dB · emolia-02669
(impatience and irritability, disgust, malevolence malice· normal-paced, almost no disfluency, clear, monologue)avant d'appeler quelqu'un en disant, bah, tu coupes le courant au disjoncteur, une prise électrique qui a deux vis et deux couleurs. Tu ne peux pas te tromper. C'est impossible.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, disgust, malevolence malice; style: monologue, didactic; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.3/10; 8.8s, FR.
FR_XDvJOyjUUSc_W000048 · in -19.1 dBFS · gain -0.9 dB · emolia-02669
(relief· normal-paced, some disfluency, average clarity, monologue)impossible. Il y a un fil d'une couleur, un fil de l'autre couleur, tu dois vérifier qu'il y a bien le contact. Et bien, 9 fois sur 10, elles me disent non, je ne préfère pas, j'ai peur.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: monologue, authoritative; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.8/10; 8.4s, FR.
FR_XDvJOyjUUSc_W000049 · in -18.6 dBFS · gain -1.4 dB · emolia-02669
(sadness, distress·measured, some disfluency, somewhat unclear, monologue)à mon explication sur l'incantation numéro 5, hein, c'est à dire que, euh, (low mumble) les gens qui ont besoin d'un arbitrage externe, il faut les mettre sous tutelle, sous curatelle, ou leur retirer le droit de vote. Parce que si vous n'êtes pas à même
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, distress; style: monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 5.6/10; 11.9s, FR.
FR_XDvJOyjUUSc_W000050 · in -20.6 dBFS · gain +0.6 dB · emolia-02669
Teasing ↓ / Sexual Lust ↑rare-PXR-pairs · #12
This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sexual Lust clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.30.
At the same time Teasing goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.08 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 24 s · en · emolia
k 3d_a -0.604d_b 0.299step_a 0.604step_b 0.214min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00008_S01060track EN_B00008_S01060total 24.3slevel spread 1.3 dBmax seam 1.3 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, normally alert, fairly steady, some disfluency, average clarity
(teasing, pride, triumph · brisk, neutral tension, moderate pitch range, casual)Or a good, what does Patton say, a good plan executed now is better than a great plan executed next week. Say, they're saying the same thing.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as teasing, pride, triumph; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 6.8/10; 6.9s, EN.
EN_B00008_S01060_W000316 · in -16.9 dBFS · gain -3.1 dB · emolia-00409
(hope enthusiasm optimism·normal-paced, neutral tension, moderate pitch range, casual)There's usually more than one way to obtain results. I like that. I wanna focus on that because
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as hope enthusiasm optimism; style: casual, conversational; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 7.2/10; 5.1s, EN.
EN_B00008_S01060_W000317 · in -18.1 dBFS · gain -1.9 dB · emolia-00409
(sexual lust, concentration·brisk, slightly relaxed, wide pitch range, casual)Again, I'm not arguing with you about six and one half dozen the other. I'm not arguing with you about should we bring six vehicles or five vehicles. I'm not arguing with you if we should, if we should invest
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sexual lust, concentration; style: casual, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 5.8/10; 12.0s, EN.
EN_B00008_S01060_W000318 · in -16.8 dBFS · gain -3.2 dB · emolia-00409
Teasing ↓ / Sexual Lust ↑rare-PXR-pairs · #13
This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sexual Lust below average — 0.34, lower than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.57.
At the same time Teasing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.63. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.17, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · quiet background
(teasing, hope enthusiasm optimism, amusement · normal-paced, normally alert, neutral tension, casual)Mm-hmm. So I don't know if you ever watched a ECW match. (ahem) Um, but pretty much ECW, all the matches were extreme. So you can bring anything. You can be tables, chairs, you can pin anyone anywhere, outside the arena, inside the arena. Pretty much as long as you have a referee, anything goes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as teasing, hope enthusiasm optimism, amusement; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 7.2/10; 22.8s, EN.
458131_00316196 · in -23.2 dBFS · gain +3.2 dB · podcast-04224
(relief, jealousy and envy, longing·measured, normally alert, neutral tension, casual)I'm glad that Finn is coming as the demon then, because I I always like Finn Balor. When we seen them when we went to that raw,
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly dominant, slightly vulnerable; reads as relief, jealousy and envy, longing; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.4/10; 10.2s, EN.
458131_00318656 · in -22.1 dBFS · gain +2.1 dB · podcast-05666
(intoxication altered states of consciousness, hope enthusiasm optimism, elation·slow, very low-energy, slightly relaxed, didactic)And then more WWE news that I realized too on October 1st, which will be a Friday or yeah, October 1st, which will be next Friday, is the first night of the WWE draft. And it's rumored, well, it's already been announced, spoiled, that we're gonna have Drew McIntyre.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, hope enthusiasm optimism, elation; style: didactic, whispered; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.4/10; 28.7s, EN.
458131_00320424 · in -21.2 dBFS · gain +1.2 dB · podcast-04222
(sexual lust·measured, normally alert, slightly relaxed, whispered)Roman Reigns belongs on SmackDown and Drew McIntyre stays on Raw. So that may be the first moving on the draft. It may be announced that Drew McIntyre goes to SmackDown.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as sexual lust; style: whispered, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.1/10; 14.5s, EN.
458131_00324640 · in -20.9 dBFS · gain +0.9 dB · podcast-04073
Teasing ↓ / Amusement ↑rare-PXR-pairs · #14
This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Amusement at the very top of the corpus — 0.98, higher than 98 % of clips in this corpus — and works its way down to clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total fall of 0.27.
At the same time Teasing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.63. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are -0.01, then -0.25 — a slow start, with most of the change arriving in the final step.
The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 25 s · en · emolia
k 3d_a -0.627d_b -0.265step_a 0.542step_b 0.251min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00008_S07273track EN_B00008_S07273total 25.2slevel spread 3.4 dBmax seam 3.4 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, normal-paced, normally alert, light breath
(teasing, amusement, intoxication altered states of consciousness · relaxed, moderately variable, frequent disfluency, casual)Sounds like they did. There's like that one, (low mumble) uh, Olympian that clearly got raped by like a chairman in China. I (ahem) don't know about chairman, I don't know what the term is. Something. Some funny term. Someone most likely sat in emperors.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, intoxication altered states of consciousness; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 3.1/10; 16.1s, EN.
EN_B00008_S07273_W000073 · in -19.7 dBFS · gain -0.3 dB · emolia-00414
(sexual lust, amusement, astonishment surprise·slightly relaxed, fairly steady, some disfluency, casual)And then she had to write a letter. Do you remember this? No. There was a girl, what was she, a gymnast?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, amusement, astonishment surprise; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 2.9/10; 5.5s, EN.
EN_B00008_S07273_W000074 · in -23.1 dBFS · gain +3.1 dB · emolia-00414
(relaxed, fairly steady, frequent disfluency, conversational)I couldn't maintain a website.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, no background noise; genuineness 4.5/6; vocal-burst blend 0.9/10; 3.2s, EN.
EN_B00008_S07273_W000075 · in -23.0 dBFS · gain +3.0 dB · emolia-00414
Teasing ↓ / Amusement ↑rare-PXR-pairs · #15
This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Amusement clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.
At the same time Teasing goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.99 (higher than 99 % of clips in this corpus), a change of +0.26. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.26, then +0.03, then -0.01, then +0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
The largest step is 0.26, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.22 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.27 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.22, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, quiet background
(normal-paced, normally alert, slightly relaxed, casual)So, average electricity usage rates per kilowatt hour in Queensland (low mumble) as of 2022, April 2022, 19.97 cents per kilowatt hour. Okay. So write that down.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.9/10; 13.8s, EN.
895130_00208632 · in -23.4 dBFS · gain +3.4 dB · podcast-01155
(sexual lust, malevolence malice, intoxication altered states of consciousness· normal-paced, energised, neutral tension, casual)I was to buy an electric car tomorrow, I'd give a fuck. I'd probably buy a Tesla.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, dark, very rough, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as sexual lust, malevolence malice, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; explicit content; genuineness 5.4/6; vocal-burst blend 4.9/10; 3.2s, EN.
895130_00212888 · in -24.3 dBFS · gain +4.3 dB · podcast-03237
(sexual lust, intoxication altered states of consciousness, amusement·brisk, energised, neutral tension, casual)remember Mr. Wong from fucking Hyundai. He's he's tipping his little feet in it. Elon's just he's big first. Elon has jumped, like he's pulled his pants down,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, intoxication altered states of consciousness, amusement; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 6.2/10; 7.9s, EN.
895130_00213696 · in -19.6 dBFS · gain -0.5 dB · podcast-03240
(amusement, bitterness, impatience and irritability· brisk, energised, neutral tension, casual)he's built a fucking rocket, and he's like, you know what? I'm gonna make well he built the electric car first, but very
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, bitterness, impatience and irritability; style: casual, conversational; average recording, quiet background; explicit content; genuineness 5.5/6; vocal-burst blend 6.1/10; 4.2s, EN.
895130_00214496 · in -26.2 dBFS · gain +6.2 dB · podcast-03256
(amusement, teasing, astonishment surprise·normal-paced, normally alert, neutral tension, casual)got his deep electric cars, and then he's like, you know what, I'm gonna get full dick, I'm gonna build a rocket. (chuckle) Like, does Hyundai have a rocket?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as amusement, teasing, astonishment surprise; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 1.9/10; 6.8s, EN.
895130_00214960 · in -23.1 dBFS · gain +3.1 dB · podcast-03246
Amusement ↓ / Teasing ↑rare-PXR-pairs · #16
This chain comes from the proxy rule: the same two-sided test as above, but because Teasing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Teasing below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.63.
At the same time Amusement goes the other way, from 0.71 (higher than 71 % of clips in this corpus) to 0.97 (higher than 97 % of clips in this corpus), a change of +0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.62, then -0.10, then +0.04, then +0.07 — not a clean run: step 2 moves back the other way by 0.10 before the chain recovers.
The largest step is 0.62, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.61 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.61 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 26 s · en · emolia
k 5d_a 0.266d_b 0.632step_a 0.154step_b 0.620min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_gXYkXkBeIOEtrack EN_gXYkXkBeIOEtotal 25.5slevel spread 3.1 dBmax seam 3.1 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly cool, quiet background
(normal-paced, normally alert, slightly relaxed, casual)In the cards and in the description below for those who want to see.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, dark, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.5/10; 3.6s, EN.
EN_gXYkXkBeIOE_W000205 · in -18.6 dBFS · gain -1.4 dB · emolia-02590
(teasing, embarrassment, contempt·slow, highly aroused, relaxed, casual)My little interview with, with the inventor of the infamous, I mean, famous,
full caption & clip details
An elderly somewhat masculine voice; delivery is highly aroused, slow, relaxed, variable; timbre is slightly cool, dark, rough, thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, submissive, vulnerable; reads as teasing, embarrassment, contempt; style: casual; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.2/10; 6.1s, EN.
EN_gXYkXkBeIOE_W000206 · in -20.5 dBFS · gain +0.5 dB · emolia-02590
(disgust, contempt, emotional numbness·normal-paced, normally alert, slightly relaxed, casual)Mermaid Lyndon Monophan. I didn't even say infamous. No, no, she's famous.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, dark, slightly rough, thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, contempt, emotional numbness; style: casual, playful; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.3/10; 4.7s, EN.
EN_gXYkXkBeIOE_W000207 · in -20.4 dBFS · gain +0.4 dB · emolia-02590
(confusion, embarrassment, doubt· normal-paced, normally alert, relaxed, casual)Alright, (low mumble) uhm, what else? What else? Yeah, besides how you became a mermaid, what is it?
full caption & clip details
A child somewhat masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as confusion, embarrassment, doubt; style: casual, conversational; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.7/10; 6.0s, EN.
EN_gXYkXkBeIOE_W000208 · in -17.4 dBFS · gain -2.6 dB · emolia-02590
(teasing, interest, amusement·brisk, energised, neutral tension, casual)What is it you hope to be when you get older as a mermaid? Do you plan to do mermaid shows like Liz?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as teasing, interest, amusement; style: casual, dramatic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.0/10; 4.6s, EN.
EN_gXYkXkBeIOE_W000209 · in -18.3 dBFS · gain -1.7 dB · emolia-02590
Amusement ↓ / Teasing ↑rare-PXR-pairs · #17
This chain comes from the proxy rule: the same two-sided test as above, but because Teasing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Teasing below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.
At the same time Amusement goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.99 (virtually no clip in this corpus scores higher), a change of +0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.55, then +0.08 — most of the change happening immediately, then levelling off.
The largest step is 0.55, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(slightly relaxed, fairly steady, frequent disfluency, casual)It's on the Amazon, where it has like at least about 300 slaves working around the country.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 3.2/10; 7.0s, EN.
407387_00180536 · in -27.9 dBFS · gain +7.9 dB · podcast-02266
(teasing, amusement· slightly relaxed, fairly steady, some disfluency, casual)But okay, let's move on before we turn into venture capitalists, like live.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as teasing, amusement; style: casual, playful; average recording, quiet background; explicit content; genuineness 4.9/6; vocal-burst blend 2.0/10; 4.3s, EN.
407387_00181360 · in -28.5 dBFS · gain +8.5 dB · podcast-02300
(teasing, amusement, embarrassment·relaxed, moderately variable, frequent disfluency, casual)Okay. (chuckle) What the good error? Do you really think that Putin could be misinformed about how the Russian economy is being crippled by sanctions? Not something that's easy to hide, and it's in fact not happening yet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, embarrassment; style: casual, playful; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 2.7/10; 14.7s, EN.
407387_00182648 · in -27.0 dBFS · gain +7.0 dB · podcast-02265
Elation ↓ / Teasing ↑rare-PXR-pairs · #18
This chain comes from the proxy rule: the same two-sided test as above, but because Teasing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Teasing below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.
At the same time Elation goes the other way, from 0.71 (higher than 71 % of clips in this corpus) to 0.97 (higher than 97 % of clips in this corpus), a change of +0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.60, then -0.03, then +0.07 — not a clean run: step 2 moves back the other way by 0.03 before the chain recovers.
The largest step is 0.60, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.57 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.52 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.57, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, some disfluency
(slightly relaxed, fairly steady, average clarity, conversational)one for Hugo John Senna Bim Bam Ayrton Senna.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, playful; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 0.6/10; 3.9s, EN.
912890_00090536 · in -19.5 dBFS · gain -0.5 dB · podcast-04141
(astonishment surprise, confusion, teasing· slightly relaxed, moderately variable, average clarity, casual)is three of that sur la cupidole d'ailleurs. La
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as astonishment surprise, confusion, teasing; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 2.3/10; 6.7s, EN.
912890_00103848 · in -18.1 dBFS · gain -1.9 dB · podcast-04131
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, amusement, affection; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 2.5/10; 6.2s, EN.
912890_00124456 · in -16.5 dBFS · gain -3.5 dB · podcast-00701
(teasing, amusement, intoxication altered states of consciousness·relaxed, moderately variable, somewhat unclear, casual)Crick crac hop, it's the brick you fit my slop. That's it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, slightly guarded; reads as teasing, amusement, intoxication altered states of consciousness; style: casual, playful; below-average recording, some background noise; mildly explicit content; genuineness 5.7/6; vocal-burst blend 3.6/10; 22.5s, EN.
912890_00126120 · in -15.9 dBFS · gain -4.1 dB · podcast-03762
Elation ↓ / Teasing ↑rare-PXR-pairs · #19
This chain comes from the proxy rule: the same two-sided test as above, but because Teasing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Teasing below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.
At the same time Elation goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.99 (higher than 99 % of clips in this corpus), a change of +0.27. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.60, then +0.04 — a plateau around step 1, where it barely moves.
The largest step is 0.60, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.28 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.21 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.28, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, moderately variable, light breath
(normal-paced, neutral tension, frequent disfluency, casual)Information después de Modex, este vuele direct to PLAD. Y
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 5.0/10; 7.8s, EN.
957936_00077048 · in -26.9 dBFS · gain +6.9 dB · podcast-03744
(malevolence malice, contempt, triumph·fast, slightly relaxed, some disfluency, monologue)cruce PLAD a 9 mil pies. Y ya le dieron la restriction of 9 mil pies, and you have no problem, pues puede.
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, contempt, triumph; style: monologue, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.5/10; 9.2s, EN.
957936_00077820 · in -27.0 dBFS · gain +7.0 dB · podcast-03747
(intoxication altered states of consciousness, teasing, interest·brisk, neutral tension, some disfluency, casual)Obviamente estamos hablando de que con esa previsional, pues ya el piloto va a hacer lo necessario for that it cumple ya con esos (ahem) regimes ofensos. (low mumble) Sin embargo, bueno, sí, específicamente, tanto que sí me va a mandar al cuerno del compañero de Centro México, con todos los
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness, teasing, interest; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 21.2s, EN.
957936_00078768 · in -23.0 dBFS · gain +3.0 dB · podcast-03746
(teasing, elation, pleasure ecstasy· brisk, neutral tension, some disfluency, casual)después de Models se vaya a Play and Cruce. No, mejor te lo paso de una vez anda, como ves, exactamente. But si es cosa de como que concientizar al piloto de que bueno, espera que have this. Nuevamente, pues como lo comentamos, basando la experiencia, si ya sabes que espera que te acordemos.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as teasing, elation, pleasure ecstasy; style: casual, conversational; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 20.5s, EN.
957936_00081528 · in -22.8 dBFS · gain +2.8 dB · podcast-03740
Helplessness ↓ / Shame ↑rare-PXR-pairs · #20
This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Shame clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.25.
At the same time Helplessness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.41 (lower than 59 % of clips in this corpus), a change of -0.58. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.03, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 42 s · italian · mls
k 3d_a -0.576d_b 0.253step_a 0.576step_b 0.223min_cos_consec —min_cos_anchor —dataset mlslang italianspeaker 4974track 4974|dongesualdo_14_verga_total 42.1slevel spread 0.6 dBmax seam 0.6 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged feminine voice · neutral-toned, slightly dark, balanced body, quiet background, measured
(helplessness, relief, contentment · normally alert, slightly relaxed, fairly steady, narration)godendosi il fresco e la libertà della campagna ascoltando i lamenti interminabili e i discorsi sconclusionati dei suoi mezzaiuoli alla
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as helplessness, relief, contentment; style: narration, monologue; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 2.1/10; 11.7s, ITALIAN.
4974_3829_000599 · in -27.3 dBFS · gain +7.3 dB · mls-00071
(contentment, anger, relief ·subdued, neutral tension, moderately variable, monologue)alla moglie che l'aria della campagna faceva star peggio soleva dire per consolarla qui almeno non hai paura d'acchiappare il colèra finché non si tratta di colèra il resto è nulla
full caption & clip details
An elderly somewhat feminine voice; delivery is subdued, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as contentment, anger, relief; style: monologue, whispered; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.6/10; 17.5s, ITALIAN.
4974_3829_000886 · in -26.7 dBFS · gain +6.7 dB · mls-00071
(shame, contentment ·normally alert, slightly relaxed, fairly steady, monologue)lì egli era al sicuro dal colèra come un re nel suo regno guardato di notte e di giorno a ogni contadino aveva procurato il suo bravo schioppo
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, contentment; style: monologue, whispered; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.0/10; 12.6s, ITALIAN.
4974_3829_000091 · in -26.8 dBFS · gain +6.8 dB · mls-00071