vprof_vc voice profiles, rule VN1, k=3, T=0.25 C=0.25. One profile is one cloned voice by construction, so there is no speaker-identity risk; there is also no time axis, so the order is chosen rather than observed.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. vprof_vc (mined by gridsel/vpgrid, PROVISIONAL) | Family. voice-profile grid (vprof_vc): one cloned voice, chains CONSTRUCTED not discovered
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
AGEV — perceived age ↑vp-VN1-k3 · #1
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 18 s · de · vprof_vc
hear it un-normalised (raw levels, max seam 2.7 dB)
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker anime_016track anime_016total 17.7slevel spread 2.7 dBmax seam 2.7 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, good recording, no background noise, slightly relaxed, no disfluency
(astonishment surprise, amusement, pain · slow, energised, moderately variable, playful)Im Ernst? (mournful wail) Hochintensives Intervalltraining in deinen Therapieplan aufzunehmen... Ich hatte mir etwas anderes für mein Training erhofft.
full caption & clip details
A child masculine voice; delivery is energised, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, gravelly, thin; very clear, no disfluency, very wide pitch range, breathless; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, amusement, pain; style: playful, cartoonish; good recording, no background noise; contains vocal bursts: Surprised Gasp; genuineness 1.3/6; vocal-burst blend 4.4/10; 1.8s, DE.
anime_016__E__Disappointment__A__de.c010.k0 · in -20.8 dBFS · gain +0.8 dB · vprof_vc-00000
(shame, thankfulness gratitude, awe·normal-paced, normally alert, fairly steady, narration)While I appreciate God's way of letting Scripture challenge my beliefs, this constant scrutiny just weighs so heavily on my spirit. It feels like an endless, unwelcome cross to carry.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, thankfulness gratitude, awe; style: narration, formal; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.2/10; 10.0s, EN.
anime_016__E__Contempt__C__en.c024.k1 · in -23.5 dBFS · gain +3.5 dB · vprof_vc-00000
(normal-paced, normally alert, fairly steady, narration)Alpha hoch Beta minus die Quadratwurzel von zwei mal den Kehrwert von Alpha gleich Beta Quadrat
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.7/10; 5.6s, DE.
anime_016__E__Contempt__C__de.c040.k3 · in -22.6 dBFS · gain +2.6 dB · vprof_vc-00000
AROU — arousal / activation ↑vp-VN1-k3 · #2
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 31 s · de · vprof_vc
hear it un-normalised (raw levels, max seam 3.7 dB)
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker anime_016track anime_016total 31.1slevel spread 3.7 dBmax seam 3.7 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · fairly smooth, balanced body, good recording
(longing, disappointment, sadness · slow, very low-energy, relaxed, monologue)Ich kann einfach nicht glauben, dass das schon wieder passiert. (contented sigh) (gurgling) Diese falsche Gastgeberin, die für Martin Gebhardt arbeitet, ist ihm so nah.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, fairly guarded; reads as longing, disappointment, sadness; style: monologue, whispered; good recording, quiet background; explicit content; genuineness 2.0/6; vocal-burst blend 8.0/10; 11.7s, DE.
anime_016__X__sad_crying__de.c014.k1 · in -25.8 dBFS · gain +5.8 dB · vprof_vc-00000
(disgust, sadness, sourness·measured, normally alert, slightly relaxed, monologue)Auf einem Felsen in der Bucht der Inseln brütet eine ganze Kolonie dieser ekelhaften Silbermöwen. (yawn) Und dann hast du den australischen Sturmvogel, den Petrel und diesen seltenen, schwarzgesichteten Möwen, der dort hockt.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, sadness, sourness; style: monologue, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.7/10; 12.1s, DE.
anime_016__E__Disgust__B__de.c007.k0 · in -22.1 dBFS · gain +2.1 dB · vprof_vc-00000
(normal-paced, normally alert, slightly relaxed)Matilde Zambudio, the Councilor for Urban Planning, has commissioned an opinion! This could let the Las Teresitas plan proceed, even if the old general plan is gone.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 7.0s, EN.
anime_016__E__Elation__B__en.c023.k0 · in -23.0 dBFS · gain +3.0 dB · vprof_vc-00000
AGEV — perceived age ↑vp-VN1-k3 · #3
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 29 s · de · vprof_vc
hear it un-normalised (raw levels, max seam 7.6 dB)
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker anime_088track anime_088total 29.4slevel spread 8.8 dBmax seam 7.6 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an infant masculine voice
(distress, fear, helplessness · slow, very low-energy, very tense, casual)Das lässt sie Leben aus drei zusätzlichen Dimensionen gleichzeitig spüren. (sharp inhale) Sie können sogar in diese anderen Ebenen eintreten.
full caption & clip details
An infant masculine voice; delivery is very low-energy, slow, very tense, highly volatile; timbre is cool, very dark, very rough, thin; slurred, heavy disfluency, narrow pitch range, breathless; affect is negative, very submissive, very vulnerable; reads as distress, fear, helplessness; style: casual, dramatic; below-average recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.7/10; 12.7s, DE.
anime_088__V__BRGT__extremely_low__de.c039.k2 · in -31.5 dBFS · gain +11.5 dB · vprof_vc-00008
(astonishment surprise, fear · slow, normally alert, slightly relaxed, narration)He took another Studio head, not grasping the whole thing, because Studio showed the prince's mad love was stronger than all the Kilamor and Megardor. (exhausted groan) It happened right at the decisive moment.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, minimal breath; affect is negative, neutral stance, fairly guarded; reads as astonishment surprise, fear; style: narration, monologue; good recording, no background noise; explicit content; genuineness 0.1/6; vocal-burst blend 3.9/10; 7.7s, EN.
anime_088__V__METL__moderately_low__en.c009.k0 · in -23.9 dBFS · gain +3.9 dB · vprof_vc-00008
(relief, fatigue exhaustion, impatience and irritability·measured, subdued, slightly relaxed)We need to reschedule das Meeting because of this sudden Zug. Can you check the availability for next week's slot, bitte?
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, fairly guarded; reads as relief, fatigue exhaustion, impatience and irritability; good recording, quiet background; explicit content; genuineness 1.2/6; vocal-burst blend 2.5/10; 8.7s, EN.
anime_088__V__R_ORAL__very_high__en.c017.k3 · in -22.6 dBFS · gain +2.6 dB · vprof_vc-00008
AROU — arousal / activation ↑vp-VN1-k3 · #4
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 33 s · en · vprof_vc
hear it un-normalised (raw levels, max seam 1.5 dB)
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker anime_088track anime_088total 33.4slevel spread 1.5 dBmax seam 1.5 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice
(awe, astonishment surprise, longing · slow, very low-energy, relaxed, ASMR)It struck him then, while in creative writing class, as he put together his initial stories. (sniff) (contented sigh) That's when he saw it.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, very dark, gravelly, very thin; slurred, little disfluency, narrow pitch range, normal breath; affect is negative, submissive, vulnerable; reads as awe, astonishment surprise, longing; style: ASMR, whispered; below-average recording, quiet background; mildly explicit content; genuineness 1.0/6; vocal-burst blend 2.9/10; 10.3s, EN.
anime_088__V__R_CHST__very_high__en.c039.k1 · in -24.1 dBFS · gain +4.1 dB · vprof_vc-00008
(interest, awe, astonishment surprise · slow, very low-energy, slightly relaxed, monologue)Sometimes, in intraday trading, a whole trend—its start, its peak, even when it turns around—can happen in just one trading session. It moves that fast.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is slightly warm, neutral-bright, fairly smooth, very thin; clear, some disfluency, moderate pitch range, minimal breath; affect is negative, neutral stance, slightly vulnerable; reads as interest, awe, astonishment surprise; style: monologue, dramatic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.5/10; 10.1s, EN.
anime_088__V__ESTH__moderately_low__en.c020.k3 · in -25.1 dBFS · gain +5.1 dB · vprof_vc-00008
(disgust, infatuation·normal-paced, normally alert, slightly relaxed, monologue)And every single star from those brilliant days, they signed off: Klaus Toppmöller, Ronnie Hellström, Seppl Pirrung, Hannes Bongartz, Wolfgang Wolf, Jürgen Groh, Reinhard Meier, Werner Melzer, Peter Schwarz, and of course, Hans-Peter Briegel. (deep breath) It was a monumental undertaking, you know.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as disgust, infatuation; style: monologue, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 5.2/10; 12.6s, EN.
anime_088__V__R_MIXD__moderately_high__en.c012.k3 · in -23.7 dBFS · gain +3.7 dB · vprof_vc-00008
AGEV — perceived age ↑vp-VN1-k3 · #5
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.01, lower than 99 % of clips in this corpus — and ends with it around average at 0.51, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 21 s · en · vprof_vc
hear it un-normalised (raw levels, max seam 1.2 dB)
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0039track emolia_c0039total 21.3slevel spread 1.2 dBmax seam 1.2 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, slightly relaxed, light breath
(teasing, infatuation, elation · very slow, normally alert, moderately variable, playful)Imagine the sheer joy as they used these masks to invite young girls to share in the spring's sweet celebration. Those hopeful moments of camaraderie beneath the cherry blossoms felt truly magical.
full caption & clip details
A child masculine voice; delivery is normally alert, very slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as teasing, infatuation, elation; style: playful, casual; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 6.0/10; 1.3s, EN.
emolia_c0039__E__Hope_Enthusiasm_Optimism__A__en.c031.k3 · in -23.1 dBFS · gain +3.1 dB · vprof_vc-00016
(brisk, energised, fairly steady)On the third of April, at three thirty in the afternoon, we will commence the long review period. This will last for exactly seven days, beginning that afternoon.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; no dominant emotion; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 8.7s, EN.
emolia_c0039__V__BKGN__extremely_low__en.c007.k2 · in -23.0 dBFS · gain +3.0 dB · vprof_vc-00016
(awe, sourness, disgust·normal-paced, normally alert, fairly steady, ranting)Schau dir diesen Concolor an, diese zarte Rosette aus blassgrünen Blättern, die Hochblätter mit dem sanftesten Orange schimmern und Spitzen wie eingefangenes Sonnenlicht haben. (childlike giggle) Es ist einfach atemberaubend schön.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, sourness, disgust; style: ranting, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 5.0/10; 11.0s, DE.
emolia_c0039__E__Infatuation__A__de.c000.k3 · in -24.2 dBFS · gain +4.2 dB · vprof_vc-00016
AROU — arousal / activation ↑vp-VN1-k3 · #6
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 38 s · de · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c0039track emolia_c0039total 38.4slevel spread 3.4 dBmax seam 1.7 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth
(disappointment, teasing, pain · measured, very low-energy, neutral tension, monologue)(contented sigh) Jesus sagte, ein Prophet wird nicht immer dort respektiert, wo er herkommt. (contented sigh) (wolf whistle) Wegen dieses Mangels an Glauben hat er dort nicht viele Wunder gewirkt. (contented sigh)
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, teasing, pain; style: monologue, casual; below-average recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.8/10; 14.5s, DE.
emolia_c0039__V__AROU__extremely_low__de.c011.k1 · in -26.2 dBFS · gain +6.2 dB · vprof_vc-00016
(sourness, jealousy and envy, anger·normal-paced, normally alert, slightly relaxed, monologue)Wir müssen uns wirklich genau um schwangere Frauen kümmern, die Anämie haben. Es gibt ein paar sehr spezifische Dinge, die wir bei der Behandlung beachten müssen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, jealousy and envy, anger; style: monologue, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 5.5/10; 13.2s, DE.
emolia_c0039__V__ATCK__moderately_low__de.c014.k1 · in -24.5 dBFS · gain +4.5 dB · vprof_vc-00016
(infatuation, pleasure ecstasy, contentment·brisk, energised, slightly relaxed, conversational)You know, when you manage to make me feel this way, this totally charged up... it's just so ridiculously playful. It’s like you’re deliberately pushing my buttons, and I love it.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, little disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as infatuation, pleasure ecstasy, contentment; style: conversational, playful; good recording, no background noise; mildly explicit content; genuineness 0.1/6; vocal-burst blend 2.4/10; 10.4s, EN.
emolia_c0039__E__Teasing__A__en.c031.k0 · in -22.8 dBFS · gain +2.8 dB · vprof_vc-00016
AGEV — perceived age ↑vp-VN1-k3 · #7
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 19 s · en · vprof_vc
k 3d_a 0.500d_b 0.500step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0238track emolia_c0238total 19.0slevel spread 1.5 dBmax seam 1.5 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child strongly feminine voice · fairly smooth, no background noise, slightly relaxed
(amusement, elation, teasing · brisk, energised, moderately variable, playful)Chaplin, bless him, played the jaunty tramp again! (hiccups) He found himself quite smitten with a blind little flower girl in that chilly city.
full caption & clip details
A child strongly feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; crisply articulate, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as amusement, elation, teasing; style: playful, casual; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 6.0/10; 0.7s, EN.
emolia_c0238__C__sprightly-pixie__en.c035.k0 · in -23.9 dBFS · gain +3.9 dB · vprof_vc-00024
(longing, infatuation, sadness·measured, very low-energy, moderately variable, monologue)By chapter thirty-seven, I just couldn't… I couldn't even feel anything when L cuffed Light. All that time, I just kept believing him, thinking he was the one.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, little disfluency, moderate pitch range, minimal breath; affect is mildly negative, submissive, vulnerable; reads as longing, infatuation, sadness; style: monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.0/10; 10.5s, EN.
emolia_c0238__E__Emotional_Numbness__D__en.c045.k2 · in -25.4 dBFS · gain +5.4 dB · vprof_vc-00024
(emotional numbness, intoxication altered states of consciousness, confusion·normal-paced, normally alert, steady, casual)h t t p colon slash slash p dot y slash file name dot txt. that is p dot y slash file name dot txt.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, intoxication altered states of consciousness, confusion; style: casual, narration; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.3/10; 7.4s, EN.
emolia_c0238__E__Concentration__D__en.c046.k1 · in -24.4 dBFS · gain +4.4 dB · vprof_vc-00024
AROU — arousal / activation ↑vp-VN1-k3 · #8
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 29 s · de · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c0238track emolia_c0238total 29.4slevel spread 2.0 dBmax seam 1.1 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged somewhat feminine voice · fairly smooth, balanced body
(longing, helplessness, distress · slow, very low-energy, relaxed, monologue)Sie versuchen, die Dinge wieder in Ordnung zu bringen, Renate und Karlheinz, ich weiß, aber wie er Maria ansieht... (contented sigh) (normal breathing) Oh Gott, es ist wie ein Messer in meiner Brust.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, vulnerable; reads as longing, helplessness, distress; style: monologue; below-average recording, quiet background; mildly explicit content; genuineness 1.8/6; vocal-burst blend 1.7/10; 14.0s, DE.
emolia_c0238__X__fear_scream__de.c032.k0 · in -25.7 dBFS · gain +5.7 dB · vprof_vc-00024
(emotional numbness·normal-paced, normally alert, slightly relaxed, didactic)Das dritte Quartal begann am ersten Juli. Wir erwarten, dass sich die Prognosen gegen Ende des dritten Quartals stabilisieren.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: didactic, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.1/10; 8.1s, DE.
emolia_c0238__E__Embarrassment__A__de.c045.k0 · in -24.7 dBFS · gain +4.7 dB · vprof_vc-00024
(fatigue exhaustion, embarrassment, confusion· normal-paced, normally alert, neutral tension, casual)My (wistful sigh) arms are already aching from shuffling this pile, so I just need to grab two facedown cards, one after the other, and then finally one face up.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, embarrassment, confusion; style: casual, storytelling; good recording, no background noise; mildly explicit content; genuineness 2.4/6; vocal-burst blend 1.6/10; 6.9s, EN.
emolia_c0238__E__Fatigue_Exhaustion__A__en.c018.k1 · in -23.6 dBFS · gain +3.6 dB · vprof_vc-00024
AGEV — perceived age ↑vp-VN1-k3 · #9
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 33 s · en · vprof_vc
k 3d_a 0.500d_b 0.500step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0323track emolia_c0323total 33.4slevel spread 4.8 dBmax seam 4.8 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an infant masculine voice
(pain, helplessness, distress · slow, highly aroused, relaxed, casual)We first (wistful sigh) calculate the basal metabolic rate, based on your personal details. (resonant hum) Then, we can figure out the energy your performance will require.
full caption & clip details
An infant masculine voice; delivery is highly aroused, slow, relaxed, highly volatile; timbre is cool, bright, gravelly, slightly thin; very slurred, frequent disfluency, very wide pitch range, breathless; affect is elated, slightly dominant, very vulnerable; reads as pain, helplessness, distress; style: casual, dramatic; below-average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 18.8s, EN.
emolia_c0323__V__AGEV__extremely_low__en.c005.k2 · in -28.5 dBFS · gain +8.5 dB · vprof_vc-00032
(astonishment surprise, amusement, disgust·brisk, energised, neutral tension, storytelling)So, when we talk about SARMs, you gotta mention they zero in on muscle tissue, which is pretty wild. Plus, they seriously boost building muscle and help you pack on some weight, right?
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as astonishment surprise, amusement, disgust; style: storytelling, dramatic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.3/10; 5.6s, EN.
emolia_c0323__E__Teasing__C__en.c026.k3 · in -23.7 dBFS · gain +3.7 dB · vprof_vc-00032
(contemplation, longing, sadness·normal-paced, normally alert, slightly relaxed)Honestly, (slow breathing) thinking about that epiclesis part of Communion just drains me; it’s so much to process that God gives presence, not our actions.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, longing, sadness; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.9/10; 8.7s, EN.
emolia_c0323__E__Fatigue_Exhaustion__C__en.c007.k2 · in -25.4 dBFS · gain +5.4 dB · vprof_vc-00032
AROU — arousal / activation ↑vp-VN1-k3 · #10
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 35 s · en · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0323track emolia_c0323total 35.0slevel spread 1.2 dBmax seam 1.2 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright
(infatuation, fatigue exhaustion, helplessness · slow, very low-energy, relaxed, casual)(surprised gasp) This never-ending road... (wolf (low mumble) whistle) it just makes everything else feel so heavy right now. (contented sigh) The things going on lately, they've really unsettled me. (breathy giggle)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, variable; timbre is slightly cool, neutral-bright, gravelly, slightly thin; slurred, frequent disfluency, wide pitch range, normal breath; affect is negative, submissive, vulnerable; reads as infatuation, fatigue exhaustion, helplessness; style: casual; below-average recording, quiet background; mildly explicit content; genuineness 1.0/6; vocal-burst blend 1.1/10; 16.3s, EN.
emolia_c0323__V__AROU__extremely_low__en.c013.k3 · in -25.7 dBFS · gain +5.7 dB · vprof_vc-00032
(sadness, distress, disgust· slow, very low-energy, neutral tension, monologue)This medication helps build up the linings in my throat. (wistful sigh) It always leaves a sour, metallic taste lingering in my mouth.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, almost no disfluency, wide pitch range, minimal breath; affect is negative, submissive, vulnerable; reads as sadness, distress, disgust; style: monologue; good recording, no background noise; mildly explicit content; genuineness 0.6/6; vocal-burst blend 0.8/10; 8.6s, EN.
emolia_c0323__E__Sourness__D__en.c015.k3 · in -24.5 dBFS · gain +4.5 dB · vprof_vc-00032
(awe, fear, distress ·normal-paced, normally alert, slightly relaxed, casual)Look at this harbor, it's where the whole fishing fleet ties up, but the sheer scale of the maritime trade... (gulps) the mountains of cement, the endless containers, the steel... it's overwhelming.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as awe, fear, distress; style: casual, monologue; good recording, no background noise; mildly explicit content; genuineness 0.6/6; vocal-burst blend 2.4/10; 9.7s, EN.
emolia_c0323__E__Fear__B__en.c038.k0 · in -24.5 dBFS · gain +4.5 dB · vprof_vc-00032
AGEV — perceived age ↑vp-VN1-k3 · #11
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 21 s · de · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c0382track emolia_c0382total 21.3slevel spread 1.1 dBmax seam 1.1 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child strongly feminine voice
(pain, helplessness, fear · normal-paced, highly aroused, fully relaxed, dramatic)Eine Wahl in einem Traum, hm? (scream) Das ist nur der Schatten von zwei schrecklichen kleinen Gedanken, die sich bekämpfen, jeder will, dass du zustimmst, aber nicht ganz!
full caption & clip details
A child strongly feminine voice; delivery is highly aroused, normal-paced, fully relaxed, volatile; timbre is cool, bright, gravelly, thin; slurred, no disfluency, very wide pitch range, heavy breath; affect is positive, slightly dominant, neutral openness; reads as pain, helplessness, fear; style: dramatic, cartoonish; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 2.6/10; 3.4s, DE.
emolia_c0382__C__screeching-impish-crone__de.c011.k0 · in -21.4 dBFS · gain +1.4 dB · vprof_vc-00040
(teasing, astonishment surprise, doubt·slow, lethargic, relaxed, monologue)Oh, diese Verse aus Sure vier, Vers vierunddreißig. Kannst du nicht glauben, was die jetzt sagen?
full caption & clip details
A middle-aged somewhat feminine voice; delivery is lethargic, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, very wide pitch range, normal breath; affect is mildly negative, slightly submissive, neutral openness; reads as teasing, astonishment surprise, doubt; style: monologue, playful; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.2/10; 11.4s, DE.
emolia_c0382__X__pain_groan__de.c018.k0 · in -22.4 dBFS · gain +2.4 dB · vprof_vc-00040
(pleasure ecstasy, elation, awe·brisk, normally alert, slightly relaxed, casual)The music swelled as the garrison band marched into the square, filling the air with such a joyful sound. (cackle) It felt like everything was finally right in this beautiful moment.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pleasure ecstasy, elation, awe; style: casual, conversational; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.7/10; 6.1s, EN.
emolia_c0382__E__Contentment__C__en.c002.k3 · in -21.5 dBFS · gain +1.5 dB · vprof_vc-00040
AROU — arousal / activation ↑vp-VN1-k3 · #12
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 25 s · en · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0382track emolia_c0382total 25.1slevel spread 4.2 dBmax seam 4.2 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright, fairly smooth
(fatigue exhaustion, jealousy and envy, distress · measured, very low-energy, relaxed, conversational)(contented sigh) It's so much to keep track of, monitoring my daily movement... trying to see if this is just fatigue or something more serious. (contented sigh)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as fatigue exhaustion, jealousy and envy, distress; style: conversational, casual; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.3/10; 12.1s, EN.
emolia_c0382__X__pain_groan__en.c031.k2 · in -24.3 dBFS · gain +4.3 dB · vprof_vc-00040
(emotional numbness, fatigue exhaustion, helplessness· measured, very low-energy, slightly relaxed, monologue)Staring at that overlapping letter writing feature in their software is just draining. My brain is fried trying to figure out that contradiction for them.
full caption & clip details
An adult feminine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is slightly warm, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as emotional numbness, fatigue exhaustion, helplessness; style: monologue, ASMR; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 6.4/10; 5.0s, EN.
emolia_c0382__E__Fatigue_Exhaustion__A__en.c044.k1 · in -20.1 dBFS · gain +0.1 dB · vprof_vc-00040
(awe·normal-paced, normally alert, slightly relaxed, narration)Ehrlich gesagt, ein Grund, bei den Live-Dingen vorbeizuschauen, ist, weil du einfach ein winziges, hektisches Projektil in dein Zuhause einlädst. (chuckle) Außerdem, wer will sich mit einem Haustier herumschlagen, das plötzlich entscheidet, ein Grillen sei ein Wrestling-Gegner?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as awe; style: narration, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.2/10; 7.7s, DE.
emolia_c0382__E__Amusement__C__de.c033.k3 · in -21.0 dBFS · gain +1.0 dB · vprof_vc-00040
AGEV — perceived age ↑vp-VN1-k3 · #13
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 17 s · de · vprof_vc
k 3d_a 0.500d_b 0.500step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c0645track emolia_c0645total 16.6slevel spread 6.3 dBmax seam 6.3 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · brisk
(distress, helplessness, fear · frantic, very tense, volatile, dramatic)Hört her, ihr Narren! (scream) Schaut auf meinen Blog, gefüllt mit Nippes und Wahrheiten über mehr als dreißig Sorten dieser Gebirgspudel-Kreuzungen. Seht die Wunder und erkennt die Lügen, die in ihnen allen stecken!
full caption & clip details
A child masculine voice; delivery is frantic, brisk, very tense, volatile; timbre is cool, bright, gravelly, thin; slurred, no disfluency, very wide pitch range, heavy breath; affect is deeply negative, slightly dominant, guarded; reads as distress, helplessness, fear; style: dramatic, ranting; below-average recording, quiet background; contains vocal bursts: Ahem, Surprised Gasp; genuineness 1.4/6; vocal-burst blend 1.3/10; 3.0s, DE.
emolia_c0645__C__screeching-impish-crone__de.c003.k0 · in -21.0 dBFS · gain +1.0 dB · vprof_vc-00048
(impatience and irritability, contempt, malevolence malice·highly aroused, slightly tense, moderately variable, ranting)After the Ottomans fell, the Allies promised a Kurdish state, yet Kemal's forces and Britain both schemed to control us. They simply couldn't stand letting the Kurds have what was rightfully theirs.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, very rough, thin; clear, almost no disfluency, very wide pitch range, normal breath; affect is negative, very dominant, guarded; reads as impatience and irritability, contempt, malevolence malice; style: ranting, casual; average recording, quiet background; mildly explicit content; genuineness 1.2/6; vocal-burst blend 2.4/10; 9.9s, EN.
emolia_c0645__E__Anger__A__en.c045.k2 · in -18.2 dBFS · gain -1.8 dB · vprof_vc-00048
(pain, contentment, sadness·normally alert, slightly relaxed, fairly steady, authoritative)Manche Stücke ziehen einfach an dir vorbei, wie eine sanfte Sommerbrise. Andere bauen sich aber zu etwas gewaltigem auf und fordern deine volle Aufmerksamkeit.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pain, contentment, sadness; style: authoritative, dramatic; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.7/10; 3.4s, DE.
emolia_c0645__E__Awe__A__de.c031.k2 · in -24.5 dBFS · gain +4.5 dB · vprof_vc-00048
AROU — arousal / activation ↑vp-VN1-k3 · #14
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 42 s · en · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0645track emolia_c0645total 42.5slevel spread 1.9 dBmax seam 1.9 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice
(pain, fatigue exhaustion, sexual lust · slow, lethargic, relaxed, casual)(exhausted groan) (low mumble) (ahem) And then, the teacher gets blamed, because some kid's needs outside the books are tripping up the whole lesson. (exhausted groan) It's like the easiest target to throw the whole problem at. (resonant hum)
full caption & clip details
An adult masculine voice; delivery is lethargic, slow, relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, slightly guarded; reads as pain, fatigue exhaustion, sexual lust; style: casual; below-average recording, quiet background; mildly explicit content; genuineness 1.9/6; vocal-burst blend 0.0/10; 15.7s, EN.
emolia_c0645__X__whimpering__en.c011.k1 · in -25.1 dBFS · gain +5.1 dB · vprof_vc-00048
(jealousy and envy, disappointment, longing· slow, very low-energy, neutral tension, casual)If he even let you use the laundry room, and he even leaves his things there, at least you have one spot you can actually clean. (contented sigh) It's so tiny, but it's something, and that's all I have left.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, neutral tension, variable; timbre is slightly cool, neutral-bright, gravelly, thin; slurred, some disfluency, wide pitch range, normal breath; affect is negative, submissive, vulnerable; reads as jealousy and envy, disappointment, longing; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 1.1/6; vocal-burst blend 0.9/10; 18.1s, EN.
emolia_c0645__X__ga_sad_cry__en.c031.k3 · in -25.3 dBFS · gain +5.3 dB · vprof_vc-00048
(normal-paced, normally alert, slightly relaxed, didactic)Et cetera wird E-T-C-E-T-E-R-A ausgesprochen. Das bedeutet und so weiter, zum Beispiel Äpfel Bananen Orangen und so weiter.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.3s, DE.
emolia_c0645__E__Contemplation__B__de.c002.k1 · in -23.4 dBFS · gain +3.4 dB · vprof_vc-00048
AGEV — perceived age ↑vp-VN1-k3 · #15
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 19 s · de · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c0758track emolia_c0758total 19.2slevel spread 3.3 dBmax seam 2.6 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, balanced body
(pleasure ecstasy, amusement, elation · normal-paced, energised, slightly relaxed, dramatic)Diese alberne Vorstellung von einer Wahl? (scream) Hah! Das ist nur das winselnde Gezeter zweier kleiner Meinungen, die um das kratzen, worauf sie sich wirklich einigen. Ein zerbrochener kleiner Wunsch, das ist er!
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, slightly relaxed, volatile; timbre is neutral-toned, bright, slightly rough, balanced body; average clarity, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, amusement, elation; style: dramatic, conversational; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.8/10; 4.3s, DE.
emolia_c0758__C__screeching-impish-crone__de.c011.k0 · in -25.5 dBFS · gain +5.5 dB · vprof_vc-00056
(pain, fatigue exhaustion, helplessness·slow, very low-energy, relaxed, monologue)Er hat mich einfach von meinem Fahrrad in dieses Maisfeld gestoßen. (sharp whistle) Mein Bein tut so weh vom Sturz.
full caption & clip details
An adult somewhat feminine voice; delivery is very low-energy, slow, relaxed, variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as pain, fatigue exhaustion, helplessness; style: monologue, conversational; good recording, no background noise; mildly explicit content; genuineness 2.3/6; vocal-burst blend 1.8/10; 7.5s, DE.
emolia_c0758__X__pain_groan__de.c013.k0 · in -24.8 dBFS · gain +4.8 dB · vprof_vc-00056
(relief, disappointment, fear·normal-paced, normally alert, slightly relaxed, conversational)Du fühlst es für den Jack Russell, oder das Gefühl geht dich einfach durch. Es ist unmöglich, so einen Hund nicht zu bemerken.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, disappointment, fear; style: conversational, monologue; good recording, quiet background; mildly explicit content; genuineness 1.5/6; vocal-burst blend 3.6/10; 7.1s, DE.
emolia_c0758__E__Emotional_Numbness__D__de.c027.k3 · in -22.2 dBFS · gain +2.2 dB · vprof_vc-00056
AROU — arousal / activation ↑vp-VN1-k3 · #16
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 32 s · de · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c0758track emolia_c0758total 31.5slevel spread 4.6 dBmax seam 3.8 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult somewhat feminine voice · neutral-toned, fairly smooth, balanced body, quiet background
(fatigue exhaustion, sadness, helplessness · slow, very low-energy, relaxed, casual)Wir müssen zweimal im Monat ins Krankenhaus für diese Coronavirus-Kontrollen. (contented sigh) (sharp whistle) Das ist so mühsam, nur endloses Testen und Warten.
full caption & clip details
An adult somewhat feminine voice; delivery is very low-energy, slow, relaxed, variable; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, vulnerable; reads as fatigue exhaustion, sadness, helplessness; style: casual, monologue; below-average recording, quiet background; mildly explicit content; genuineness 2.0/6; vocal-burst blend 1.3/10; 15.3s, DE.
emolia_c0758__X__pain_groan__de.c003.k0 · in -27.0 dBFS · gain +7.0 dB · vprof_vc-00056
(longing, fear, sourness· slow, very low-energy, slightly relaxed, monologue)Oh nein, wenn es zu eng wird, stecke ich für immer in diesem Ding fest! (low mumble) (ahem) (gurgling) Ernsthaft, ich will dieses Zwicken nicht spüren.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, fear, sourness; style: monologue, didactic; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.0/10; 8.7s, DE.
emolia_c0758__X__ga_fear_scream__de.c011.k0 · in -23.2 dBFS · gain +3.2 dB · vprof_vc-00056
(doubt, jealousy and envy, longing ·normal-paced, normally alert, neutral tension, monologue)Erzähl mir, was dir gefällt, nehme ich an, aber ich will das Ganze irgendwann einfach hinter mir haben. Ehrlich gesagt, fühle ich nichts für das Ende.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, jealousy and envy, longing; style: monologue, conversational; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.4/10; 7.3s, DE.
emolia_c0758__E__Emotional_Numbness__D__de.c046.k1 · in -22.4 dBFS · gain +2.4 dB · vprof_vc-00056
AGEV — perceived age ↑vp-VN1-k3 · #17
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.01, lower than 99 % of clips in this corpus — and ends with it around average at 0.51, higher than 51 % of clips in this corpus. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 25 s · en · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0955track emolia_c0955total 24.7slevel spread 0.7 dBmax seam 0.7 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert
(contentment, jealousy and envy, fatigue exhaustion · brisk, neutral tension, moderately variable, conversational)This shea butter, with its vitamin E, just soothes my tired lower legs. (cackle) It feels like a real balm after a long day on my feet.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, very thin; average clarity, little disfluency, wide pitch range, minimal breath; affect is mildly negative, slightly dominant, neutral openness; reads as contentment, jealousy and envy, fatigue exhaustion; style: conversational, casual; good recording, no background noise; mildly explicit content; genuineness 0.6/6; vocal-burst blend 1.9/10; 8.5s, EN.
emolia_c0955__E__Relief__B__en.c004.k0 · in -24.2 dBFS · gain +4.2 dB · vprof_vc-00064
(shame, disappointment, distress·normal-paced, slightly relaxed, fairly steady, narration)I felt such raw exposure when he played; it was like watching someone strip bare their soul without even trying. (fearful gasp) There was a devastating, beautiful vulnerability to it that left me utterly ashamed of my own guardedness.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as shame, disappointment, distress; style: narration, formal; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.4/10; 8.3s, EN.
emolia_c0955__E__Shame__B__en.c004.k3 · in -23.8 dBFS · gain +3.8 dB · vprof_vc-00064
(jealousy and envy, sourness, doubt· normal-paced, slightly relaxed, fairly steady, formal)Everyone else is probably banking on those bankers dropping rates and sending the asset price soaring. (childlike giggle) Honestly, I'm just waiting to see if they actually have the nerve to pull it off.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, sourness, doubt; style: formal, narration; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.0/10; 7.6s, EN.
emolia_c0955__E__Teasing__C__en.c008.k0 · in -24.6 dBFS · gain +4.6 dB · vprof_vc-00064
AROU — arousal / activation ↑vp-VN1-k3 · #18
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 37 s · en · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang enspeaker emolia_c0955track emolia_c0955total 36.7slevel spread 1.3 dBmax seam 1.1 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth
(pain, fatigue exhaustion, distress · slow, lethargic, relaxed)(ahem) (low mumble) Because of how much this place meant, (wolf whistle) it went from being almost forgotten to a real draw for tourists coming to Egypt.
full caption & clip details
A young adult masculine voice; delivery is lethargic, slow, relaxed, variable; timbre is slightly cool, dark, fairly smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, breathless; affect is mildly negative, submissive, vulnerable; reads as pain, fatigue exhaustion, distress; below-average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.7/10; 17.4s, EN.
emolia_c0955__V__AROU__extremely_low__en.c006.k0 · in -27.0 dBFS · gain +7.0 dB · vprof_vc-00064
(awe·fast, energised, slightly relaxed)The concrete supports dictate the facade's look; there's a central gable section bordered by slimmer wings that curve softly outward, (soft whistle) never quite reaching the peak's height.
full caption & clip details
An adult masculine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, minimal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as awe; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 10.8s, EN.
emolia_c0955__V__AGEV__very_high__en.c003.k2 · in -25.9 dBFS · gain +5.9 dB · vprof_vc-00064
(disappointment, helplessness, sadness·normal-paced, energised, slightly relaxed, ranting)Fünfunddreißig Jahre, und die Hotels schrumpfen einfach weiter, verlieren fast einen Prozent jedes einzelne Jahr. Es ist erschreckend, wie schnell sie verschwinden.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, helplessness, sadness; style: ranting, dramatic; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 4.8/10; 8.2s, DE.
emolia_c0955__E__Fear__A__de.c020.k3 · in -25.6 dBFS · gain +5.6 dB · vprof_vc-00064
AGEV — perceived age ↑vp-VN1-k3 · #19
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 11 s · de · vprof_vc
k 3d_a 0.500d_b 0.500step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c1028track emolia_c1028total 11.3slevel spread 2.3 dBmax seam 2.0 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an infant masculine voice
(pain, distress, fear · very slow, frantic, very tense, dramatic)Lass sie einfach stehen, bis der Spieler diese Pfeile selbst abzieht. Lass sie nicht einfach da stehen, es wird lächerlich, zu warten.
full caption & clip details
An infant masculine voice; delivery is frantic, very slow, very tense, highly volatile; timbre is slightly cool, dark, gravelly, thin; slurred, no disfluency, very wide pitch range, breathless; affect is deeply negative, very dominant, guarded; reads as pain, distress, fear; style: dramatic, cartoonish; poor recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.0/10; 2.7s, DE.
emolia_c1028__E__Impatience_and_Irritability__C__de.c030.k1 · in -20.6 dBFS · gain +0.6 dB · vprof_vc-00072
(astonishment surprise·normal-paced, normally alert, slightly relaxed, conversational)Aber ehrlich gesagt, es gibt erstaunliche Wege, diesen Lebensstil zu meistern! Drei echte Möglichkeiten für unglaubliches Wachstum warten darauf, entdeckt zu werden.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: conversational, formal; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.0/10; 3.4s, DE.
emolia_c1028__E__Hope_Enthusiasm_Optimism__B__de.c025.k2 · in -22.6 dBFS · gain +2.6 dB · vprof_vc-00072
(jealousy and envy, sourness, malevolence malice·brisk, normally alert, slightly relaxed, casual)The rest of the butchers and chefs butchered at their homes, spilling out into yards all over the city. It was all a blur, a haze from too much swill.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy, sourness, malevolence malice; style: casual, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 3.8/10; 4.9s, EN.
emolia_c1028__E__Intoxication_Altered_States_of_Consciousness__D__en.c016.k3 · in -22.8 dBFS · gain +2.8 dB · vprof_vc-00072
AROU — arousal / activation ↑vp-VN1-k3 · #20
This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.
The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.
Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.
3 clips · 26 s · de · vprof_vc
k 3d_a 0.499d_b 0.499step_a —step_b —min_cos_consec —min_cos_anchor —dataset vprof_vclang despeaker emolia_c1028track emolia_c1028total 26.3slevel spread 1.7 dBmax seam 1.0 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · neutral-toned, neutral-bright, balanced body
(fatigue exhaustion, pain, sadness · slow, lethargic, relaxed, monologue)(contented sigh) All die mühsame Arbeit, die Rückzahlung an einen Aktionär... (wistful (wistful sigh) sigh) das ist doch nur der Tageskurs minus der Kaufpreis, plus Dividenden, minus Bankgebühren. Diese letzte Zahl zu sehen, bringt so einen schweren Schmerz.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly negative, submissive, vulnerable; reads as fatigue exhaustion, pain, sadness; style: monologue, narration; average recording, quiet background; mildly explicit content; genuineness 1.6/6; vocal-burst blend 3.0/10; 15.7s, DE.
emolia_c1028__E__Sadness__C__de.c039.k0 · in -22.6 dBFS · gain +2.6 dB · vprof_vc-00072
(relief, teasing, doubt· slow, very low-energy, slightly relaxed, casual)Diese These, ach, sie singt! (chuckle) Sie erfasst den Kern des Aufsatzes so perfekt und legt meine Haltung zu dieser ganzen Angelegenheit mit solch glorreicher Klarheit offen.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as relief, teasing, doubt; style: casual, playful; average recording, no background noise; contains vocal bursts: Ahem, Exhausted Groan; genuineness 3.6/6; vocal-burst blend 2.5/10; 1.7s, DE.
emolia_c1028__E__Pleasure_Ecstasy__B__de.c005.k0 · in -23.5 dBFS · gain +3.5 dB · vprof_vc-00072
(contempt, fear, anger·brisk, normally alert, slightly relaxed, ranting)Du musst sicherstellen, dass die Ports achtzig und vierzig-dreiundvierzig in deinem Netzwerk offen sind. Das lässt den Webverkehr ohne Probleme durchkommen.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, fear, anger; style: ranting, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.2/10; 8.7s, DE.
emolia_c1028__V__ARSH__extremely_low__de.c016.k1 · in -24.3 dBFS · gain +4.3 dB · vprof_vc-00072