vp-PXR-k3

vprof_vc voice profiles, rule PXR, k=3, T=0.25 C=0.25. One profile is one cloned voice by construction, so there is no speaker-identity risk; there is also no time axis, so the order is chosen rather than observed.

Rule. PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes of a/b for emotions that are not directly rampable
Source. vprof_vc (mined by gridsel/vpgrid, PROVISIONAL)  |  Family. voice-profile grid (vprof_vc): one cloned voice, chains CONSTRUCTED not discovered
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Affection ↓  /  Amusementvp-PXR-k3 · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.62, then +0.02 — most of the change happening immediately, then levelling off.

The largest step is 0.62, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 24 s · en · vprof_vc

hear it un-normalised (raw levels, max seam 2.9 dB)
k 3d_a -0.404d_b 0.642step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker anime_016track anime_016total 24.5slevel spread 2.9 dBmax seam 2.9 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(affection, longing, sadness · measured, no disfluency, clear, narration) I cannot bear to see you walk away, my heart aches at the thought. Four shots, perhaps, is the only way I know to keep you near me forever.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as affection, longing, sadness; style: narration, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.4/10; 6.4s, EN.
anime_016__E__Affection__B__en.c027.k1 · in -22.5 dBFS · gain +2.5 dB · vprof_vc-00000
(pleasure ecstasy, amusement, teasing · normal-paced, no disfluency, clear, playful) The secret mischief didn't just cloud my pixie duties. It sprinkled bad luck all over my little light too.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pleasure ecstasy, amusement, teasing; style: playful, narration; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.0/10; 5.4s, EN.
anime_016__C__sprightly-pixie__en.c043.k3 · in -22.3 dBFS · gain +2.3 dB · vprof_vc-00000
(amusement, teasing, pleasure ecstasy · normal-paced, little disfluency, very clear, playful) (breathy giggle) Der jetzige Zeitpunkt war perfekt, um die frisch gebackenen Kekse zu präsentieren. Er konnte seine Ergebnisse mit großer Zuversicht vorstellen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, little disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as amusement, teasing, pleasure ecstasy; style: playful, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.5/10; 12.4s, DE.
anime_016__X__ga_amuse_laugh__de.c009.k0 · in -25.2 dBFS · gain +5.2 dB · vprof_vc-00000
Affection ↓  /  Angervp-PXR-k3 · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.40.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 36 s · en · vprof_vc

hear it un-normalised (raw levels, max seam 2.1 dB)
k 3d_a -0.404d_b 0.404step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker anime_016track anime_016total 35.8slevel spread 2.1 dBmax seam 2.1 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-bright, balanced body
(affection, longing, helplessness · slow, lethargic, relaxed, casual) He feels like love and being loved, sharing and giving it, somehow fueled the cancer. (coughing) It's heartbreaking to think that something so beautiful could be twisted like this.
full caption & clip details
A young adult masculine voice; delivery is lethargic, slow, relaxed, moderately variable; timbre is slightly warm, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, neutral openness; reads as affection, longing, helplessness; style: casual; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.6/10; 14.8s, EN.
anime_016__X__ga_sad_cry__en.c039.k3 · in -24.6 dBFS · gain +4.6 dB · vprof_vc-00000
(awe, disgust, sexual lust · measured, normally alert, slightly relaxed, monologue) Diese Assassinen, die für die gegnerische Rücklinie gedacht waren, finden die Magier einfach zu zerbrechlich. Ihre niedrigeren Lebenspunkte bedeuten, dass die Tanks ganz leicht an ihnen vorbeikommen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, disgust, sexual lust; style: monologue, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.1/10; 13.9s, DE.
anime_016__X__tears__de.c023.k0 · in -22.5 dBFS · gain +2.5 dB · vprof_vc-00000
(anger, bitterness, impatience and irritability · normal-paced, normally alert, slightly relaxed, conversational) You dragged Moses and Aaron back before Pharaoh, and I swear, you told them to bow down to your god! You think we'll just obey your pointless deity now?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as anger, bitterness, impatience and irritability; style: conversational; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.4/10; 6.8s, EN.
anime_016__E__Bitterness__D__en.c038.k2 · in -22.9 dBFS · gain +2.9 dB · vprof_vc-00000
Affection ↓  /  Contemplationvp-PXR-k3 · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.36.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 27 s · en · vprof_vc

hear it un-normalised (raw levels, max seam 2.2 dB)
k 3d_a -0.252d_b 0.363step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker anime_088track anime_088total 26.6slevel spread 3.6 dBmax seam 2.2 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, no background noise, slightly relaxed, moderate pitch range
(affection, jealousy and envy, disappointment · slow, subdued, moderately variable, conversational) (contented sigh) She just doesn't know what happened to her mom. My mom, who always had the perfect thing to say, whose laugh could lift you up, she was the heart of this whole family.
full caption & clip details
An adult masculine voice; delivery is subdued, slow, slightly relaxed, moderately variable; timbre is slightly warm, neutral-bright, fairly smooth, full; average clarity, almost no disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly submissive, fairly guarded; reads as affection, jealousy and envy, disappointment; style: conversational; good recording, no background noise; mildly explicit content; genuineness 0.8/6; vocal-burst blend 6.9/10; 8.4s, EN.
anime_088__V__FOCS__very_high__en.c020.k3 · in -25.6 dBFS · gain +5.6 dB · vprof_vc-00008
(awe, contentment, relief · slow, very low-energy, moderately variable, monologue) Here, the mountains meet the sea, running right out to the water. (smack one s lips) The Great Ocean Road snakes along them in a long, winding path.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is slightly warm, neutral-bright, fairly smooth, very thin; clear, little disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, fairly guarded; reads as awe, contentment, relief; style: monologue, ASMR; good recording, no background noise; mildly explicit content; genuineness 0.7/6; vocal-burst blend 2.9/10; 8.9s, EN.
anime_088__V__EXPL__moderately_high__en.c039.k1 · in -23.4 dBFS · gain +3.4 dB · vprof_vc-00008
(contemplation, awe, longing · normal-paced, normally alert, fairly steady, narration) Manchmal reicht es schon, jemanden durch ein Fenster zu beobachten, um einen Einblick in seine wahren Gedanken zu bekommen. Es ist wie ein Fenster in seine innere Welt.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, awe, longing; style: narration, formal; very good recording, no background noise; genuineness 0.4/6; vocal-burst blend 5.7/10; 9.1s, DE.
anime_088__V__DARC__extremely_low__de.c047.k2 · in -22.0 dBFS · gain +2.0 dB · vprof_vc-00008
Affection ↓  /  Contemptvp-PXR-k3 · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Contempt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contempt clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.39.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 33 s · de · vprof_vc

hear it un-normalised (raw levels, max seam 0.4 dB)
k 3d_a -0.404d_b 0.387step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang despeaker anime_088track anime_088total 33.0slevel spread 0.5 dBmax seam 0.4 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, good recording, measured
(affection, contentment, thankfulness gratitude · very low-energy, slightly relaxed, fairly steady, monologue) Durch kleine Schritte mit viel Liebe und Geduld blüht der Maltipoo zu einem toleranten Familienfreund auf. (drinking noises) Das lässt ihn zu einem so wunderbaren, vertrauensvollen Begleiter werden.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, contentment, thankfulness gratitude; style: monologue, casual; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 7.6/10; 11.3s, DE.
anime_088__V__R_MIXD__very_high__de.c012.k3 · in -25.0 dBFS · gain +5.0 dB · vprof_vc-00008
(disgust, sexual lust, teasing · very low-energy, fully relaxed, moderately variable, casual) (contented sigh) A pastry made with a very light yeast dough, topped with cream and various grated or crumbled cheeses. (person whistling to get attention) It's a sweet and savory combination you have to try.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, fully relaxed, moderately variable; timbre is slightly warm, neutral-bright, fairly smooth, full; clear, little disfluency, fairly narrow pitch, minimal breath; affect is mildly negative, slightly submissive, neutral openness; reads as disgust, sexual lust, teasing; style: casual, whispered; good recording, no background noise; mildly explicit content; genuineness 0.8/6; vocal-burst blend 0.6/10; 9.7s, EN.
anime_088__V__EMPH__extremely_low__en.c006.k3 · in -24.6 dBFS · gain +4.6 dB · vprof_vc-00008
(contempt, anger, malevolence malice · normally alert, slightly relaxed, fairly steady, storytelling) Du wirst weder deine Bohnen, noch dein Abendessen, noch deinen Darm ruinieren, indem du einen Weg statt des anderen wählst. (nervous giggle) Ehrlich gesagt, macht es zwischen diesen beiden Ansätzen keinen wirklichen Unterschied.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, fairly smooth, very full; clear, no disfluency, moderate pitch range, light breath; affect is negative, neutral stance, fairly guarded; reads as contempt, anger, malevolence malice; style: storytelling, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 5.6/10; 11.8s, DE.
anime_088__V__REGS__extremely_low__de.c010.k2 · in -24.5 dBFS · gain +4.5 dB · vprof_vc-00008
Affection ↓  /  Amusementvp-PXR-k3 · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.43, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

The largest step is 0.43, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 30 s · en · vprof_vc

hear it un-normalised (raw levels, max seam 1.9 dB)
k 3d_a -0.404d_b 0.639step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0039track emolia_c0039total 29.7slevel spread 1.9 dBmax seam 1.9 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(affection, contentment, jealousy and envy · normal-paced, little disfluency, average clarity, narration) Even with Augustus' illness, even knowing how brief this joy is, I feel such profound peace just being with him. (guffaw) Our love, I know that will never fade.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, contentment, jealousy and envy; style: narration, conversational; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.7/10; 6.7s, EN.
emolia_c0039__E__Relief__C__en.c037.k0 · in -22.9 dBFS · gain +2.9 dB · vprof_vc-00016
(fatigue exhaustion, relief, thankfulness gratitude · measured, no disfluency, clear, narration) Send the fishing crew out near the beach for a full seven days. Honestly, I'm sick of waiting for a decent catch this season.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, relief, thankfulness gratitude; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.9/10; 6.5s, EN.
emolia_c0039__E__Sourness__B__en.c042.k3 · in -22.9 dBFS · gain +2.9 dB · vprof_vc-00016
(amusement, teasing, hope enthusiasm optimism · normal-paced, some disfluency, average clarity, playful) Im Ernst, wenn wir uns heute bei der Heiligen Kommunion vor Christus verneigen, könnt ihr vielleicht einen kleinen Vorgeschmack auf die Ewigkeit bekommen. Es ist fast wie ein winziger, himmlischer Geschmack, nicht wahr?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as amusement, teasing, hope enthusiasm optimism; style: playful, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 2.5/10; 16.2s, DE.
emolia_c0039__E__Teasing__A__de.c044.k3 · in -24.8 dBFS · gain +4.8 dB · vprof_vc-00016
Affection ↓  /  Disappointmentvp-PXR-k3 · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Disappointment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disappointment below average — 0.39, lower than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.61.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.60, then +0.01 — most of the change happening immediately, then levelling off.

The largest step is 0.60, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 31 s · en · vprof_vc

k 3d_a -0.404d_b 0.605step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0039track emolia_c0039total 31.2slevel spread 1.7 dBmax seam 1.7 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, clear
(affection, infatuation, pain · measured, moderately variable, almost no disfluency, narration) Perhaps this isn't love when I say you are dearest to me; love is knowing you are the blade I use to cut into my own soul.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, minimal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as affection, infatuation, pain; style: narration; good recording, no background noise; explicit content; genuineness 0.7/6; vocal-burst blend 1.2/10; 6.9s, EN.
emolia_c0039__V__ARSH__very_high__en.c025.k1 · in -22.8 dBFS · gain +2.8 dB · vprof_vc-00016
(bitterness, jealousy and envy, disappointment · normal-paced, fairly steady, no disfluency, narration) Ich wünschte mir nur, ich hätte einen Pastor in meinem Leben, jemanden so absolut engagiert, der immer zu seiner Kirche eilt und seine Glocke für alle läuten lässt. Es ist wahnsinnig, wie erfüllt sein Leben mit solch einem Sinn ist.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, jealousy and envy, disappointment; style: narration, ranting; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 5.7/10; 10.3s, DE.
emolia_c0039__E__Jealousy_and_Envy__D__de.c043.k1 · in -24.4 dBFS · gain +4.4 dB · vprof_vc-00016
(disappointment, disgust, bitterness · normal-paced, fairly steady, no disfluency, narration) All diese Arbeit über den Mangel an Frauen, besonders mit den Problemen der abgebrochenen Feten, die Nilanjana Ray dokumentiert hat, macht mich einfach völlig ausgelaugt. Diese Ultraschalluntersuchung in Jamshedpur hat mir wirklich zu schaffen gemacht.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, disgust, bitterness; style: narration, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 5.0/10; 13.7s, DE.
emolia_c0039__E__Fatigue_Exhaustion__D__de.c016.k1 · in -23.4 dBFS · gain +3.4 dB · vprof_vc-00016
Affection ↓  /  Amusementvp-PXR-k3 · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.64 — a slow start, with most of the change arriving in the final step.

The largest step is 0.64, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 41 s · en · vprof_vc

k 3d_a -0.252d_b 0.643step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0238track emolia_c0238total 41.0slevel spread 2.6 dBmax seam 2.6 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice
(affection, contentment, pleasure ecstasy · normal-paced, normally alert, slightly relaxed, conversational) It's just... they are so loving, so gentle, and so utterly sweet. It’s almost too much to ask for perfection, and sometimes it's just heartbreaking.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, contentment, pleasure ecstasy; style: conversational, narration; good recording, no background noise; mildly explicit content; genuineness 1.5/6; vocal-burst blend 2.6/10; 8.7s, EN.
emolia_c0238__E__Disappointment__D__en.c029.k0 · in -25.1 dBFS · gain +5.2 dB · vprof_vc-00024
(sexual lust, fatigue exhaustion, pleasure ecstasy · slow, very low-energy, relaxed, casual) (contented sigh) Your fat is shrinking, I can feel it, because you're burning so much more fuel now. (normal breathing) Don't stop, or it could all just come right back.
full caption & clip details
A young adult strongly feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, slightly bright, slightly rough, slightly thin; slurred, frequent disfluency, narrow pitch range, minimal breath; affect is mildly negative, submissive, neutral openness; reads as sexual lust, fatigue exhaustion, pleasure ecstasy; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 1.4/6; vocal-burst blend 2.6/10; 13.1s, EN.
emolia_c0238__X__fear_scream__en.c037.k3 · in -27.7 dBFS · gain +7.7 dB · vprof_vc-00024
(amusement, teasing, pleasure ecstasy · normal-paced, energised, neutral tension, casual) Seriously, we gotta put padding and canvas on those walls so no rogue bullets go bouncing around like crazy. Imagine the mess, it would be a total disaster!
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, volatile; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; good recording, quiet background; mildly explicit content; genuineness 1.5/6; vocal-burst blend 0.3/10; 18.9s, EN.
emolia_c0238__X__ga_amuse_laugh__en.c031.k3 · in -27.1 dBFS · gain +7.1 dB · vprof_vc-00024
Affection ↓  /  Angervp-PXR-k3 · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.39.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 26 s · en · vprof_vc

k 3d_a -0.404d_b 0.389step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0238track emolia_c0238total 25.7slevel spread 0.2 dBmax seam 0.2 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, clear
(affection, contentment, infatuation · normal-paced, steady, little disfluency, ASMR) It's just... they are so loving, so gentle, and so utterly sweet. It’s almost too much to ask for perfection, and sometimes it's just heartbreaking.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, contentment, infatuation; style: ASMR, narration; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.7/10; 8.7s, EN.
emolia_c0238__E__Disappointment__D__en.c029.k2 · in -25.2 dBFS · gain +5.2 dB · vprof_vc-00024
(contemplation, fear, confusion · normal-paced, fairly steady, no disfluency, formal) Es ist frustrierend, wie wir Muskelkrämpfe einordnen, wenn man nicht nur bedenkt, wann sie auftreten oder wie schmerzhaft sie sind. Wir müssen sie sogar nach der Art, wie sie sich zeigen, unterteilen, was ziemlich umständlich ist.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, fear, confusion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.9/10; 10.1s, DE.
emolia_c0238__E__Disappointment__B__de.c045.k0 · in -25.4 dBFS · gain +5.4 dB · vprof_vc-00024
(anger, impatience and irritability, contempt · brisk, moderately variable, no disfluency, storytelling) You think I'm stupid? (displeased grunt) The whole loan costs less if you actually help finance this damn property! Don't make me ask again.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as anger, impatience and irritability, contempt; style: storytelling, conversational; good recording, no background noise; mildly explicit content; genuineness 0.7/6; vocal-burst blend 0.6/10; 6.5s, EN.
emolia_c0238__E__Anger__D__en.c004.k1 · in -25.3 dBFS · gain +5.3 dB · vprof_vc-00024
Affection ↓  /  Amusementvp-PXR-k3 · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.59, then +0.05 — most of the change happening immediately, then levelling off.

The largest step is 0.59, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 26 s · en · vprof_vc

k 3d_a -0.404d_b 0.641step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0323track emolia_c0323total 26.3slevel spread 3.6 dBmax seam 3.6 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, moderately variable, wide pitch range
(affection, infatuation, malevolence malice · brisk, energised, slightly relaxed, conversational) Still, I wouldn't hesitate to fight him for you. (spitting) I would do anything to have your love.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as affection, infatuation, malevolence malice; style: conversational, storytelling; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.4/10; 5.2s, EN.
emolia_c0323__V__ATCK__very_high__en.c038.k3 · in -25.0 dBFS · gain +5.0 dB · vprof_vc-00032
(awe, astonishment surprise, elation · brisk, energised, neutral tension, casual) This Eastern Market is bursting with wonder, it’s amazing! (cackle) I just asked Ronan, have you noticed anything extraordinary around here recently?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as awe, astonishment surprise, elation; style: casual, storytelling; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 4.1/10; 2.9s, EN.
emolia_c0323__E__Triumph__A__en.c004.k1 · in -22.8 dBFS · gain +2.8 dB · vprof_vc-00032
(amusement, teasing, pleasure ecstasy · slow, very low-energy, neutral tension, casual) If the investment really pays off, we could afford some real extravagance. (swallows) Think about what we could buy with a larger return, sweetheart. (breathy giggle)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; very slurred, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, vulnerable; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 18.0s, EN.
emolia_c0323__P__explicit__en.c007.k2 · in -26.4 dBFS · gain +6.4 dB · vprof_vc-00032
Amusement ↓  /  Angervp-PXR-k3 · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.28.

At the same time Amusement goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.64. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 30 s · en · vprof_vc

k 3d_a -0.642d_b 0.282step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0323track emolia_c0323total 29.8slevel spread 2.5 dBmax seam 2.5 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, neutral tension, wide pitch range
(amusement, pleasure ecstasy, disgust · slow, very low-energy, moderately variable, casual) Seriously, this whole thing is proper dodgy innit. It's like the absolute worst, proper rubbish, innit. (surprised gasp)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, vulnerable; reads as amusement, pleasure ecstasy, disgust; style: casual, playful; good recording, no background noise; mildly explicit content; genuineness 1.8/6; vocal-burst blend 1.3/10; 11.5s, EN.
emolia_c0323__V__AROU__moderately_high__en.c012.k1 · in -27.0 dBFS · gain +7.0 dB · vprof_vc-00032
(pain, embarrassment, jealousy and envy · normal-paced, normally alert, moderately variable, casual) Look at Tantawi, a minister for so many years under Mubarak, now leading the scaf; it's infuriating. (displeased grunt) The whole military brass, those relics of the old regime, still have their grip.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as pain, embarrassment, jealousy and envy; style: casual, storytelling; good recording, no background noise; mildly explicit content; genuineness 1.9/6; vocal-burst blend 0.5/10; 8.5s, EN.
emolia_c0323__E__Jealousy_and_Envy__A__en.c032.k0 · in -24.6 dBFS · gain +4.6 dB · vprof_vc-00032
(anger, disappointment, impatience and irritability · brisk, energised, fairly steady, casual) We had to shut down every single computer, every data storage unit, every connection, and every listening device the MfS and AfNS were using. It was becoming utterly maddening how they kept operating.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as anger, disappointment, impatience and irritability; style: casual, ranting; good recording, no background noise; mildly explicit content; genuineness 0.5/6; vocal-burst blend 1.5/10; 9.5s, EN.
emolia_c0323__E__Impatience_and_Irritability__D__en.c018.k0 · in -24.9 dBFS · gain +4.9 dB · vprof_vc-00032
Affection ↓  /  Astonishment Surprisevp-PXR-k3 · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.48.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 22 s · en · vprof_vc

k 3d_a -0.404d_b 0.475step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0382track emolia_c0382total 22.1slevel spread 1.3 dBmax seam 1.3 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, full, good recording, no background noise, moderately variable, no disfluency, clear
(affection, contentment, pleasure ecstasy · brisk, energised, slightly relaxed, storytelling) I'd love for you to join us, sweetheart. The IMDb rating plugin is just for our registered community members, and we cherish having you here with us.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, no disfluency, wide pitch range, minimal breath; affect is positive, slightly dominant, neutral openness; reads as affection, contentment, pleasure ecstasy; style: storytelling, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 8.1s, EN.
emolia_c0382__E__Affection__A__en.c026.k2 · in -21.0 dBFS · gain +1.0 dB · vprof_vc-00040
(relief, fear, helplessness · measured, normally alert, slightly relaxed, monologue) (breathy giggle) When you break down mediation, it’s so satisfying because there's always that opening, the substance, and then a clear wrap-up. (breathy giggle) That predictable structure just brings such a calm sense of completion.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as relief, fear, helplessness; style: monologue, narration; good recording, no background noise; mildly explicit content; genuineness 0.7/6; vocal-burst blend 0.4/10; 7.7s, EN.
emolia_c0382__E__Contentment__A__en.c033.k2 · in -21.6 dBFS · gain +1.6 dB · vprof_vc-00040
(astonishment surprise, awe, impatience and irritability · normal-paced, normally alert, neutral tension, storytelling) No way, you're seriously saying that happened? Like, no fucking way that actually went down, dude. That's wild, for real.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, awe, impatience and irritability; style: storytelling, narration; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.3/10; 6.1s, EN.
emolia_c0382__C__undead__en.c025.k1 · in -20.4 dBFS · gain +0.4 dB · vprof_vc-00040
Affection ↓  /  Disappointmentvp-PXR-k3 · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Disappointment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disappointment below average — 0.39, lower than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.61.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.61 — a slow start, with most of the change arriving in the final step.

The largest step is 0.61, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 27 s · en · vprof_vc

k 3d_a -0.404d_b 0.605step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0382track emolia_c0382total 27.0slevel spread 3.4 dBmax seam 2.1 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-bright, fairly smooth
(affection, contentment, pleasure ecstasy · normal-paced, normally alert, slightly relaxed, formal) They are so sweet, truly docile and cheerful, but I wish they were a bit more... lively. Still, their purrs and cuddles make them the beloved, peaceful members of our family.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, contentment, pleasure ecstasy; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 8.0s, EN.
emolia_c0382__E__Disappointment__D__en.c029.k3 · in -20.3 dBFS · gain +0.3 dB · vprof_vc-00040
(longing, contentment, awe · normal-paced, normally alert, slightly relaxed, monologue) We visited the charming city of Ljubljana near the Ljubljanica River. It was a beautiful place, like a postcard from Slovenia.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as longing, contentment, awe; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.9/10; 7.1s, EN.
emolia_c0382__E__Elation__D__en.c000.k2 · in -21.6 dBFS · gain +1.6 dB · vprof_vc-00040
(disappointment, distress, helplessness · brisk, energised, neutral tension, casual) The file didn't copy right to the drive; it's missing parts! (contented sigh) (contented sigh) I can't even open it, it's totally corrupted!
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is negative, submissive, vulnerable; reads as disappointment, distress, helplessness; style: casual, playful; below-average recording, quiet background; genuineness 1.1/6; vocal-burst blend 1.0/10; 11.6s, EN.
emolia_c0382__X__ga_fear_scream__en.c020.k2 · in -23.8 dBFS · gain +3.8 dB · vprof_vc-00040
Affection ↓  /  Angervp-PXR-k3 · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.33.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 23 s · en · vprof_vc

k 3d_a -0.404d_b 0.333step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0645track emolia_c0645total 22.9slevel spread 1.0 dBmax seam 0.6 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult somewhat feminine voice · moderately variable, wide pitch range
(affection, infatuation, contentment · slow, very low-energy, neutral tension) (contented sigh) Marriage, a blessed arrangement, a dream come true, truly. (heavy breathing) And love, true love, it will follow you forever. So treasure your love, every single moment.
full caption & clip details
An adult somewhat feminine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is slightly warm, neutral-bright, slightly rough, full; somewhat unclear, little disfluency, wide pitch range, minimal breath; affect is negative, submissive, vulnerable; reads as affection, infatuation, contentment; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.3/10; 10.8s, EN.
emolia_c0645__C__seasoned-merchant__en.c013.k0 · in -24.6 dBFS · gain +4.6 dB · vprof_vc-00048
(relief, contentment, pleasure ecstasy · brisk, energised, slightly relaxed, playful) I am so ready for my Feierabend, I just want to kick off my shoes and relax. It is time to finally unwind after this whole Tag.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as relief, contentment, pleasure ecstasy; style: playful, ranting; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.7/10; 6.5s, EN.
emolia_c0645__E__Contemplation__B__en.c030.k3 · in -24.1 dBFS · gain +4.1 dB · vprof_vc-00048
(anger, impatience and irritability, contempt · brisk, energised, slightly tense, playful) How dare you think they'll just magically accept them? (yawn) After this ordeal, they'll see what you've done.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, full; average clarity, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as anger, impatience and irritability, contempt; style: playful, dramatic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.6/10; 5.3s, EN.
emolia_c0645__E__Anger__A__en.c038.k1 · in -23.6 dBFS · gain +3.6 dB · vprof_vc-00048
Affection ↓  /  Contemplationvp-PXR-k3 · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.31.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 33 s · en · vprof_vc

k 3d_a -0.404d_b 0.309step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0645track emolia_c0645total 32.6slevel spread 0.6 dBmax seam 0.3 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult somewhat feminine voice · neutral-bright, neutral tension, moderately variable, wide pitch range
(affection, infatuation, contentment · slow, lethargic, little disfluency) (contented sigh) Marriage, a blessed arrangement, a dream come true, truly. (heavy breathing) And love, true love, it will follow you forever. So treasure your love, every single moment.
full caption & clip details
An adult somewhat feminine voice; delivery is lethargic, slow, neutral tension, moderately variable; timbre is slightly warm, neutral-bright, slightly rough, full; somewhat unclear, little disfluency, wide pitch range, minimal breath; affect is negative, submissive, vulnerable; reads as affection, infatuation, contentment; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.4/10; 10.8s, EN.
emolia_c0645__C__seasoned-merchant__en.c013.k2 · in -24.5 dBFS · gain +4.5 dB · vprof_vc-00048
(jealousy and envy, awe, helplessness · slow, very low-energy, some disfluency, casual) (contented sigh) Oh, those clasped hands in a dream... it just screams of turmoil for the Imam, or the very leader of this land. Such knotted fingers speak of deep, gnawing complications in their world.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is negative, neutral stance, vulnerable; reads as jealousy and envy, awe, helplessness; style: casual; average recording, quiet background; genuineness 0.1/6; vocal-burst blend 1.0/10; 14.1s, EN.
emolia_c0645__X__pain_groan__en.c022.k2 · in -24.3 dBFS · gain +4.3 dB · vprof_vc-00048
(contemplation, doubt, distress · brisk, energised, almost no disfluency, monologue) After all this time, can I truly feel anything new from Him? (yawn) Or is it just the same old, hollow ache inside?
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, vulnerable; reads as contemplation, doubt, distress; style: monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.3/10; 7.5s, EN.
emolia_c0645__E__Bitterness__C__en.c006.k0 · in -23.9 dBFS · gain +3.9 dB · vprof_vc-00048
Affection ↓  /  Disappointmentvp-PXR-k3 · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Disappointment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disappointment below average — 0.39, lower than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.60.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.60 — a slow start, with most of the change arriving in the final step.

The largest step is 0.60, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 37 s · en · vprof_vc

k 3d_a -0.404d_b 0.605step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0758track emolia_c0758total 37.0slevel spread 2.1 dBmax seam 1.6 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice
(affection, pleasure ecstasy, contentment · normal-paced, normally alert, slightly relaxed, casual) Seeing her smile when she gets something she truly wants just fills my heart. (breathy giggle) It feels like the simplest joy to bring her happiness.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, pleasure ecstasy, contentment; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 2.6/6; vocal-burst blend 3.7/10; 6.4s, EN.
emolia_c0758__E__Affection__A__en.c010.k1 · in -24.0 dBFS · gain +4.0 dB · vprof_vc-00056
(pleasure ecstasy, amusement, teasing · slow, energised, relaxed, casual) (breathy giggle) Bee Bee See, wie B B C, ist der Name dieser Rundfunkanstalt. Der Rundfunk sendet Programme von jenseits der kontinentalen Grenzen.
full caption & clip details
A young adult masculine voice; delivery is energised, slow, relaxed, volatile; timbre is neutral-toned, slightly dark, gravelly, balanced body; very slurred, frequent disfluency, wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, amusement, teasing; style: casual, playful; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 1.0/10; 14.3s, DE.
emolia_c0758__X__ga_amuse_laugh__de.c033.k0 · in -24.6 dBFS · gain +4.6 dB · vprof_vc-00056
(disappointment, fatigue exhaustion, distress · measured, very low-energy, neutral tension, casual) The whole year total sheets are now available in the updated template, and I am so terrified because I think I broke the whole thing. Please tell me this works, or I am ruined.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, vulnerable; reads as disappointment, fatigue exhaustion, distress; style: casual; below-average recording, quiet background; mildly explicit content; genuineness 0.8/6; vocal-burst blend 1.1/10; 16.0s, EN.
emolia_c0758__X__ga_fear_scream__en.c022.k3 · in -26.2 dBFS · gain +6.2 dB · vprof_vc-00056
Affection ↓  /  Distressvp-PXR-k3 · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Distress is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Distress around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.56.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.56 — a slow start, with most of the change arriving in the final step.

The largest step is 0.56, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 26 s · en · vprof_vc

k 3d_a -0.404d_b 0.558step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0758track emolia_c0758total 25.7slevel spread 5.2 dBmax seam 4.9 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, good recording
(affection, infatuation, longing · brisk, normally alert, slightly relaxed, conversational) I know and love you so much, and if you flutter away, my sparkle will fade completely. Then all the brightest bits of my world would just disappear!
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, infatuation, longing; style: conversational, storytelling; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.8/10; 6.4s, EN.
emolia_c0758__C__sprightly-pixie__en.c026.k0 · in -21.0 dBFS · gain +1.0 dB · vprof_vc-00056
(awe, elation, astonishment surprise · normal-paced, normally alert, slightly relaxed, formal) Es ist absolut aufregend zu sehen, dass globale Verantwortung in diesen Zeiten endlich die Aufmerksamkeit bekommt, die sie verdient. (breathy giggle) Wir können diesen wichtigen Wandel endlich vorantreiben, oder?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, elation, astonishment surprise; style: formal, didactic; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 6.2s, DE.
emolia_c0758__E__Elation__C__de.c007.k3 · in -21.3 dBFS · gain +1.3 dB · vprof_vc-00056
(distress, fear, helplessness · slow, very low-energy, neutral tension, dramatic) Oh Gott, wenn sich das furchtbar anfühlt, spül es sofort mit Wasser aus. Bitte, mach das sofort, bevor es schlimmer wird.
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, slow, neutral tension, variable; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; very slurred, frequent disfluency, wide pitch range, audible breath; affect is negative, submissive, vulnerable; reads as distress, fear, helplessness; style: dramatic, monologue; good recording, quiet background; mildly explicit content; genuineness 1.8/6; vocal-burst blend 2.9/10; 12.8s, DE.
emolia_c0758__X__ga_fear_scream__de.c019.k2 · in -26.2 dBFS · gain +6.2 dB · vprof_vc-00056
Affection ↓  /  Amusementvp-PXR-k3 · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.64.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.60, then +0.04 — most of the change happening immediately, then levelling off.

The largest step is 0.60, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 36 s · en · vprof_vc

k 3d_a -0.404d_b 0.637step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0955track emolia_c0955total 36.0slevel spread 3.8 dBmax seam 3.1 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice
(affection, awe, infatuation · brisk, energised, slightly tense, casual) Because you shine with a light I can't ignore, every part of you, past and present, is utterly captivating to me. You deserve all the love in the world, always and forever.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly tense, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as affection, awe, infatuation; style: casual, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.7/6; vocal-burst blend 2.2/10; 8.1s, EN.
emolia_c0955__E__Infatuation__B__en.c017.k3 · in -22.5 dBFS · gain +2.5 dB · vprof_vc-00064
(sourness, elation, teasing · brisk, energised, slightly relaxed, playful) When he leases the property to tenants, oh, the sheer joy! (ahem) (chuckle) To pass the tax, in fair proportion, as a delightful operating expense feels simply divine.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, elation, teasing; style: playful, casual; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.0/10; 9.1s, EN.
emolia_c0955__E__Pleasure_Ecstasy__B__en.c000.k0 · in -25.6 dBFS · gain +5.6 dB · vprof_vc-00064
(amusement, teasing, emotional numbness · measured, very low-energy, neutral tension, casual) These are the common drug classes for treating legionellosis, with some examples following. (affirmative grunt) Erythromycin is often the preferred choice, depending on the patient's specific condition. (yawn)
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, neutral tension, volatile; timbre is slightly cool, slightly dark, slightly rough, very thin; very slurred, some disfluency, wide pitch range, minimal breath; affect is neutral, slightly submissive, neutral openness; reads as amusement, teasing, emotional numbness; style: casual, playful; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.0/10; 18.5s, EN.
emolia_c0955__V__AGEV__extremely_low__en.c036.k1 · in -26.3 dBFS · gain +6.3 dB · vprof_vc-00064
Affection ↓  /  Bitternessvp-PXR-k3 · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Bitterness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.35.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 42 s · en · vprof_vc

k 3d_a -0.404d_b 0.346step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c0955track emolia_c0955total 41.5slevel spread 2.4 dBmax seam 1.9 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, fairly steady, light breath
(affection, awe, infatuation · brisk, energised, slightly tense, casual) Because you shine with a light I can't ignore, every part of you, past and present, is utterly captivating to me. You deserve all the love in the world, always and forever.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly tense, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as affection, awe, infatuation; style: casual, dramatic; good recording, no background noise; mildly explicit content; genuineness 0.9/6; vocal-burst blend 2.3/10; 8.1s, EN.
emolia_c0955__E__Infatuation__B__en.c017.k2 · in -23.1 dBFS · gain +3.1 dB · vprof_vc-00064
(sadness, helplessness, disgust · measured, normally alert, slightly relaxed, narration) Ich sehne mich danach, (wistful sigh) dass sie die jungen wirklich beschützen, die Last des Wachstums ihrer Sucht tragen und eine echte Chance auf Heilung bieten.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sadness, helplessness, disgust; style: narration, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 5.6/10; 7.8s, DE.
emolia_c0955__E__Longing__D__de.c040.k3 · in -23.6 dBFS · gain +3.6 dB · vprof_vc-00064
(bitterness, contempt, disgust · normal-paced, normally alert, slightly relaxed, cartoonish) Du erinnerst dich an John und Dotty, das Paar, das Pat Walsh uns gebracht hat und das sich schließlich verlobt hat? Die waren ja ein tolles Paar, nicht wahr?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, contempt, disgust; style: cartoonish, conversational; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.8/10; 25.3s, DE.
emolia_c0955__V__ATCK__very_high__de.c022.k0 · in -25.6 dBFS · gain +5.5 dB · vprof_vc-00064
Affection ↓  /  Bitternessvp-PXR-k3 · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Bitterness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.44.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 41 s · en · vprof_vc

k 3d_a -0.404d_b 0.438step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c1028track emolia_c1028total 41.1slevel spread 1.1 dBmax seam 0.7 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average clarity, light breath
(affection, contentment, pleasure ecstasy · normal-paced, normally alert, relaxed, casual) Because that necklace comes from a heart full of love, it shows a special blessing. Someone who wears it will surely be the most wonderful wife.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, contentment, pleasure ecstasy; style: casual, monologue; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 4.7/10; 9.1s, EN.
emolia_c1028__V__ATCK__extremely_low__en.c030.k2 · in -23.0 dBFS · gain +3.0 dB · vprof_vc-00072
(pain, infatuation, confusion · normal-paced, normally alert, neutral tension, casual) This E-mini Dow futures thing, like this weird buzzing in my head, it's like a tiny slice of the whole market. It feels so (ahem) strangely real right now.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pain, infatuation, confusion; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 3.8/6; vocal-burst blend 5.3/10; 6.8s, EN.
emolia_c1028__E__Intoxication_Altered_States_of_Consciousness__D__en.c023.k0 · in -22.3 dBFS · gain +2.3 dB · vprof_vc-00072
(bitterness, sadness, longing · measured, subdued, slightly relaxed, monologue) Mit derselben Farbe und demselben Grundanstrich, wie es sich auf meiner Haut anfühlte, passte ich den Farbton der Schachtel an die Wand an, sodass es genau zu meinem Geschmack passte. Jeder Farbton fühlte sich an wie ein Versprechen, das ich konsumieren wollte.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, vulnerable; reads as bitterness, sadness, longing; style: monologue, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.9/10; 24.9s, DE.
emolia_c1028__E__Sexual_Lust__C__de.c031.k2 · in -21.9 dBFS · gain +1.9 dB · vprof_vc-00072
Affection ↓  /  Hope Enthusiasm Optimismvp-PXR-k3 · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.28.

At the same time Affection goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 22 s · en · vprof_vc

k 3d_a -0.252d_b 0.283step_a step_b min_cos_consec min_cos_anchor dataset vprof_vclang enspeaker emolia_c1028track emolia_c1028total 21.9slevel spread 5.5 dBmax seam 5.5 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(affection, longing, contentment · neutral tension, moderately variable, little disfluency, storytelling) Those people, they just... (guffaw) they became family to me. I felt such a connection with the nurses, like they were my closest friends for so long.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as affection, longing, contentment; style: storytelling, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.5/10; 7.6s, EN.
emolia_c1028__E__Infatuation__C__en.c001.k1 · in -22.7 dBFS · gain +2.7 dB · vprof_vc-00072
(awe, pleasure ecstasy, contentment · slightly relaxed, fairly steady, almost no disfluency, casual) Now that that old shadow of pride has lifted, we can truly see the magnificent spark of the Divine within him. Because of that, God can work through him, and every beautiful plan unfolds just as it should.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pleasure ecstasy, contentment; style: casual, narration; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.5/10; 8.2s, EN.
emolia_c1028__E__Hope_Enthusiasm_Optimism__D__en.c015.k3 · in -22.9 dBFS · gain +3.0 dB · vprof_vc-00072
(hope enthusiasm optimism, elation, disgust · slightly relaxed, fairly steady, some disfluency, casual) Dude, I'm seriously chuffed about this gig, like stoked to bits. It's gonna be sick, no cap, I'm totally hyped for this whole thing.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, elation, disgust; style: casual, storytelling; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.3/10; 5.8s, EN.
emolia_c1028__E__Sexual_Lust__D__en.c005.k0 · in -17.4 dBFS · gain -2.6 dB · vprof_vc-00072