voicenet__VN1__T0.80__C0.25__INTERNAL

Manifest tier. voicenet, rule VN1, T=0.8, step cap 0.25. Population 20,740 chains (296 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 13,139.

Rule. VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C
Source. trajectories_v5.parquet  |  Family. the tier owner's manifest tiers -- the exact subsets used for training
Manifest tier. voicenet__VN1__T0.80__C0.25__INTERNAL — population 20,740 chains (296 h). SHAREABLE variant: 13,139.
Filter. rule=='VN1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.8 and step_b<=0.25
Sampled from 136 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
VFLX — vocal flexibility / inflectionvoicenet__VN1__T0.80__C0.25__INTERNAL · #1

This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with vocal flexibility / inflection (VFLX) high — 0.86, higher than 86 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.05, lower than 95 % of clips in this corpus. That is a total fall of 0.81.

It takes 5 clips to get there. Clip to clip the moves are -0.23, then -0.23, then -0.17, then -0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.33 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 73 s · en · podcast

hear it un-normalised (raw levels, max seam 0.8 dB)
k 5d_a -0.809d_b -0.809step_a 0.235step_b 0.235min_cos_consec 0.3325min_cos_anchor 0.1661dataset podcastlang enspeaker 936961track 936961total 72.5slevel spread 0.8 dBmax seam 0.8 dBcos from orange-id (speaker identity)
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, fairly steady, average clarity, moderate pitch range
(interest, infatuation, sexual lust · normal-paced, normally alert, slightly relaxed, casual) depending on where they are. Yeah. What is that (ahem) um conversation like for you? (ahem) Um, I mean, I think I uh in my experience, my clients (ahem) uh who get their fee adjusted down are are always super happy about it. What are those (ahem) um sessions like for you where you as a clinician bring in the fee structure to new to like change it, to change it up to increase the rate?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, infatuation, sexual lust; style: casual, conversational; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 8.9/10; 23.3s, EN.
936961_00146328 · in -18.7 dBFS · gain -1.3 dB · podcast-03267
(normal-paced, normally alert, slightly relaxed, casual) You know, I've had two experiences where it it jumped (ahem) um sizably.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 3.0/10; 4.6s, EN.
936961_00148688 · in -19.5 dBFS · gain -0.5 dB · podcast-03278
(relief, doubt, infatuation · normal-paced, subdued, relaxed, casual) (low mumble) Totally get it, but you know, they've shown it works better if it's at this certain rate. That's why it's set this way. (ahem) Um, and then you know what, they readjusted their life and their finances and found a way to keep coming. (wistful sigh) Wow. Okay. What was it like for you having as a clinician having that conversation?
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, doubt, infatuation; style: casual, ASMR; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.7/10; 29.3s, EN.
936961_00150496 · in -18.9 dBFS · gain -1.1 dB · podcast-05968
(relief, sadness, shame · normal-paced, normally alert, relaxed, casual) That one was tougher. (low mumble) Um I, you know, issues of unfairness came into the room.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, sadness, shame; style: casual, conversational; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.6/10; 7.2s, EN.
936961_00153456 · in -19.5 dBFS · gain -0.5 dB · podcast-03261
(contemplation, shame, concentration · measured, normally alert, relaxed, monologue) know, I I had to be clear about my role in protecting the therapeutic frame for this person.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, submissive, vulnerable; reads as contemplation, shame, concentration; style: monologue, casual; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 1.1/10; 7.5s, EN.
936961_00154336 · in -18.7 dBFS · gain -1.3 dB · podcast-03266
AROU — arousal / activationvoicenet__VN1__T0.80__C0.25__INTERNAL · #2

This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with arousal / activation (AROU) high — 0.89, higher than 89 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.08, lower than 92 % of clips in this corpus. That is a total fall of 0.81.

It takes 5 clips to get there. Clip to clip the moves are -0.22, then -0.13, then -0.24, then -0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.23 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.23 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 31 s · en · emolia

hear it un-normalised (raw levels, max seam 4.9 dB)
k 5d_a -0.809d_b -0.809step_a 0.236step_b 0.236min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_n9JytJhukNEtrack EN_n9JytJhukNEtotal 31.0slevel spread 6.1 dBmax seam 4.9 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice
(impatience and irritability, contempt, anger · brisk, energised, neutral tension, casual) So Rusty's just gonna turn around and fucking book it and write it as strong.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, contempt, anger; style: casual, conversational; average recording, some background noise; genuineness 4.6/6; vocal-burst blend 1.4/10; 3.5s, EN.
EN_n9JytJhukNE_W000800 · in -19.2 dBFS · gain -0.8 dB · emolia-02183
(impatience and irritability, helplessness, doubt · normal-paced, normally alert, slightly relaxed, conversational) If he's gonna run that direction, then I won't effort on the thorn whip.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as impatience and irritability, helplessness, doubt; style: conversational, casual; average recording, some background noise; genuineness 5.5/6; vocal-burst blend 0.7/10; 4.2s, EN.
EN_n9JytJhukNE_W000801 · in -16.1 dBFS · gain -3.9 dB · emolia-02183
(confusion, intoxication altered states of consciousness, doubt · normal-paced, normally alert, slightly relaxed, casual) Yeah, I wouldn't ever on that. I wouldn't even attempt to do it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as confusion, intoxication altered states of consciousness, doubt; style: casual, conversational; average recording, some background noise; genuineness 6.0/6; vocal-burst blend 3.4/10; 4.0s, EN.
EN_n9JytJhukNE_W000802 · in -21.1 dBFS · gain +1.1 dB · emolia-02183
(fatigue exhaustion · measured, normally alert, neutral tension, casual) So now the hallway starts doing a (ahem) real quick left, right, left, right, twisting and then you end up in another area.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as fatigue exhaustion; style: casual, conversational; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.8/10; 8.9s, EN.
EN_n9JytJhukNE_W000804 · in -19.6 dBFS · gain -0.3 dB · emolia-02183
(intoxication altered states of consciousness · measured, normally alert, slightly relaxed, casual) Where you can go left and then it like turns into an L shape or you can go right and it goes down to a long hallway.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.0/10; 9.9s, EN.
EN_n9JytJhukNE_W000805 · in -22.2 dBFS · gain +2.2 dB · emolia-02183
S_STRY — style: storytellingvoicenet__VN1__T0.80__C0.25__INTERNAL · #3

This is a VoiceNet dimension, not an emotion: style: storytelling (S_STRY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with style: storytelling (S_STRY) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.83.

It takes 5 clips to get there. Clip to clip the moves are -0.24, then -0.20, then -0.16, then -0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 71 s · en · podcast

hear it un-normalised (raw levels, max seam 2.4 dB)
k 5d_a -0.831d_b -0.831step_a 0.245step_b 0.245min_cos_consec 0.8512min_cos_anchor 0.8066dataset podcastlang enspeaker 28646track 28646total 71.2slevel spread 4.0 dBmax seam 2.4 dBcos from orange-id (speaker identity)
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · slightly rough, quiet background, measured
(contempt, impatience and irritability, anger · energised, slightly tense, moderately variable, authoritative) these twelve Jesus sent out charging them, go nowhere among the Gentiles, go nowhere among the Gentiles,
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, full; average clarity, almost no disfluency, wide pitch range, audible breath; affect is positive, slightly dominant, fairly guarded; reads as contempt, impatience and irritability, anger; style: authoritative, storytelling; below-average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 9.4s, EN.
28646_00189184 · in -14.2 dBFS · gain -5.8 dB · podcast-03317
(normally alert, slightly relaxed, steady, monologue) and enter no town of the Samaritans, but go rather to the (low mumble) lost sheep of the house of Israel.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.0/10; 7.6s, EN.
28646_00190116 · in -14.9 dBFS · gain -5.1 dB · podcast-03306
(interest, bitterness, anger · normally alert, slightly relaxed, fairly steady, monologue) And preach as you've been saying, you know, he gives them the authority and power to take all this away. So he originally came for the chosen people. And it's clear from so many of the parables, it's about the difference between the old covenant and the new. The old covenant misread God.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, bitterness, anger; style: monologue, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 2.5/10; 17.2s, EN.
28646_00190868 · in -15.4 dBFS · gain -4.6 dB · podcast-03788
(anger, thankfulness gratitude, malevolence malice · very low-energy, neutral tension, fairly steady, casual) The parable of the prodigal son, or the (low mumble) um the parable on the wedding feast. I mean, all of those. (low mumble) Um, so many of the parables keep setting the old covenant against the new and showing that (low mumble) um he came to bring the chosen people to fulfillment, that he is the Messiah that they've been waiting for.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as anger, thankfulness gratitude, malevolence malice; style: casual; average recording, quiet background; mildly explicit content; genuineness 3.9/6; vocal-burst blend 2.1/10; 23.2s, EN.
28646_00192584 · in -17.9 dBFS · gain -2.1 dB · podcast-03809
(anger, emotional numbness, contempt · very low-energy, slightly relaxed, fairly steady, casual) But just as they did with the prophets before Christ, they're doing with him. They've rejected them, they're taking this legalistic way of reading reading the world into what they do and they miss him. They don't see him. (low mumble) Um
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as anger, emotional numbness, contempt; style: casual, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.3/10; 13.3s, EN.
28646_00194904 · in -18.2 dBFS · gain -1.8 dB · podcast-03783
RESP — audible breath / respirationvoicenet__VN1__T0.80__C0.25__INTERNAL · #4

This is a VoiceNet dimension, not an emotion: audible breath / respiration (RESP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with audible breath / respiration (RESP) at the very top of the range — 0.91, higher than 91 % of clips in this corpus — and works its way down to low at 0.09, lower than 91 % of clips in this corpus. That is a total fall of 0.83.

It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.19, then -0.21, then -0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.51 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.42 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.51, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 44 s · es · podcast

hear it un-normalised (raw levels, max seam 4.5 dB)
k 5d_a -0.826d_b -0.826step_a 0.221step_b 0.221min_cos_consec 0.4226min_cos_anchor 0.5075dataset podcastlang esspeaker 317252track 317252total 44.4slevel spread 5.7 dBmax seam 4.5 dBcos from orange-id (speaker identity)
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert
(contempt, sourness, bitterness · fast, neutral tension, fairly steady, didactic) no tienen comprensión lectora. Ojo con eso, ¿eh? Que cuidado es que la comprensión lectora, (ahem) digamos que hay muchas personas que adolecen del nivel de comprensión lectora y no se enteran de lo que estás leyendo, ¿eh? Claro,
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, sourness, bitterness; style: didactic, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 7.0/10; 13.1s, ES.
317252_00103024 · in -26.0 dBFS · gain +6.0 dB · podcast-05086
(embarrassment, disappointment, relief · fast, slightly relaxed, fairly steady, casual) es que tú es muy. (ahem) Si no estás habituado a leer artículos, libros y tal, tampoco puedes pretender que entienda un cartel que pone prohibido fumar. Tienes que entender a la gente. Hay que dárselo más masticado todo.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, disappointment, relief; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 10.0/10; 10.1s, ES.
317252_00104352 · in -22.0 dBFS · gain +2.0 dB · podcast-05114
(relief, anger, impatience and irritability · normal-paced, neutral tension, moderately variable, conversational) Claro, claro, quizá con uno, quizá con un latiguillo que salga de algún sitio, ¡pum! ¡Espada, pan! Solucionado. (contented sigh) Vamos con el tuyo. Pues
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as relief, anger, impatience and irritability; style: conversational, dramatic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.6/10; 6.5s, ES.
317252_00105368 · in -24.8 dBFS · gain +4.8 dB · podcast-04035
(pride, embarrassment, infatuation · normal-paced, relaxed, fairly steady, casual) mira, mi. Tengo dos, ¿vale? Voy a decir uno muy rápido y más serio, que es (low mumble) (ahem) ayer, cuando acabó el partido, o el domingo cuando acabó el partido. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, embarrassment, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 10.2s, ES.
317252_00106024 · in -20.3 dBFS · gain +0.3 dB · podcast-05103
(astonishment surprise, confusion · normal-paced, slightly relaxed, moderately variable, conversational) El partido fue el domingo, ¿no? El sábado, ¿no?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, confusion; style: conversational, casual; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 1.4/10; 4.0s, ES.
317252_00107040 · in -22.9 dBFS · gain +2.9 dB · podcast-05086
S_NARR — style: narrationvoicenet__VN1__T0.80__C0.25__INTERNAL · #5

This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with style: narration (S_NARR) high — 0.85, higher than 85 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.01, lower than 99 % of clips in this corpus. That is a total fall of 0.84.

It takes 5 clips to get there. Clip to clip the moves are -0.24, then -0.24, then -0.24, then -0.12 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.30 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.43 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.30, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 69 s · en · podcast

hear it un-normalised (raw levels, max seam 1.5 dB)
k 5d_a -0.839d_b -0.839step_a 0.244step_b 0.244min_cos_consec 0.4309min_cos_anchor 0.2951dataset podcastlang enspeaker 422763track 422763total 68.8slevel spread 2.1 dBmax seam 1.5 dBcos from orange-id (speaker identity)
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · slightly bright, fairly smooth, moderately variable, some disfluency, wide pitch range
(affection, contentment · normal-paced, energised, slightly relaxed, conversational) right. So, first question: what are some key things you would like to see parents doing in order for their children to be healthy?
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, contentment; style: conversational, casual; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.1/10; 9.3s, EN.
422763_00016424 · in -30.2 dBFS · gain +10.2 dB · podcast-04413
(contentment, hope enthusiasm optimism, pleasure ecstasy · brisk, normally alert, neutral tension, casual) This is one of the few things really in my life that I'm very passionate about. (low mumble) Um, and I think there's some simple steps, and I basically have a list of them. I could go on for hours, but I think you'll
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as contentment, hope enthusiasm optimism, pleasure ecstasy; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 7.5/10; 14.0s, EN.
422763_00017464 · in -31.7 dBFS · gain +11.7 dB · podcast-05562
(fear, interest, distress · brisk, energised, neutral tension, casual) Wasn't this you need to eat this now or you'll see it again for breakfast. Think about what kind of emotion that stirs up in somebody. And everybody that I talk to, including people that are in their 90s now, will tell me, I don't eat oatmeal now because I was forced to eat it as a child. Now, if what they have experienced at five years old stays with them till they're 95 years old, it's probably not a good thing.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as fear, interest, distress; style: casual, cartoonish; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 8.9/10; 25.1s, EN.
422763_00042200 · in -31.1 dBFS · gain +11.1 dB · podcast-06272
(contemplation · brisk, energised, neutral tension, conversational) They might not, they might not choose to eat oatmeal anyway, but it was that emotion that it was attached to that feeling of you have to eat
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as contemplation; style: conversational, casual; below-average recording, quiet background; genuineness 2.5/6; vocal-burst blend 6.1/10; 8.2s, EN.
422763_00044712 · in -30.1 dBFS · gain +10.1 dB · podcast-04403
(doubt, fear, helplessness · brisk, energised, neutral tension, conversational) this now, otherwise you will, you know, I don't know what, die or or not be well nourished. I'm not sure what we're afraid of as parents. (surprised gasp) Um, but the other thing is realize that
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as doubt, fear, helplessness; style: conversational, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.2/10; 11.5s, EN.
422763_00045528 · in -29.6 dBFS · gain +9.6 dB · podcast-05557
ARSH — harshness of articulationvoicenet__VN1__T0.80__C0.25__INTERNAL · #6

This is a VoiceNet dimension, not an emotion: harshness of articulation (ARSH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with harshness of articulation (ARSH) low — 0.12, lower than 88 % of clips in this corpus — and ends with it at the very top of the range at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.83.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.20, then +0.24, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.07 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.07 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · en · emolia

k 5d_a 0.825d_b 0.825step_a 0.245step_b 0.245min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_2q-qSlJzL64track EN_2q-qSlJzL64total 33.2slevel spread 4.1 dBmax seam 2.6 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice
(measured, normally alert, slightly relaxed, monologue) And in that (low mumble) spirit, it's unlikely that we will add a PvP element.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 2.1/10; 4.7s, EN.
EN_2q-qSlJzL64_W000228 · in -20.9 dBFS · gain +0.9 dB · emolia-01905
(relief, sadness, fatigue exhaustion · normal-paced, normally alert, relaxed, conversational) But unfortunately, we have to end this now. It was a lot of fun today. And... As always! Yes!
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as relief, sadness, fatigue exhaustion; style: conversational, casual; below-average recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.8/10; 7.2s, EN.
EN_2q-qSlJzL64_W000229 · in -19.6 dBFS · gain -0.4 dB · emolia-01905
(contentment, hope enthusiasm optimism, pleasure ecstasy · normal-paced, normally alert, relaxed, conversational) Okay, yeaah, I wish you (ahem) a very nice evening and hope to see you soon and enjoy the Elvinar TV tomorrow and yeaah.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as contentment, hope enthusiasm optimism, pleasure ecstasy; style: conversational, casual; below-average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.3/10; 10.9s, EN.
EN_2q-qSlJzL64_W000232 · in -19.4 dBFS · gain -0.6 dB · emolia-01905
(elation, pleasure ecstasy, infatuation · normal-paced, normally alert, slightly relaxed, conversational) Yeah, look, the Alvinard TV show tomorrow, even though it's without her, but be sure she will be in the next episode again.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, pleasure ecstasy, infatuation; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.7/10; 6.5s, EN.
EN_2q-qSlJzL64_W000233 · in -19.5 dBFS · gain -0.5 dB · emolia-01905
(teasing, amusement, pleasure ecstasy · normal-paced, energised, slightly relaxed, casual) Let's do the keep on playing. (breathy giggle)
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, slightly relaxed, volatile; timbre is neutral-toned, bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, pleasure ecstasy; style: casual, conversational; good recording, no background noise; genuineness 4.2/6; vocal-burst blend 4.1/10; 3.3s, EN.
EN_2q-qSlJzL64_W000235 · in -16.9 dBFS · gain -3.1 dB · emolia-01905
COGL — cognitive loadvoicenet__VN1__T0.80__C0.25__INTERNAL · #7

This is a VoiceNet dimension, not an emotion: cognitive load (COGL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with cognitive load (COGL) at the very top of the range — 0.96, higher than 96 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.81.

It takes 5 clips to get there. Clip to clip the moves are -0.24, then -0.23, then -0.16, then -0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.23 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 37 s · en · podcast

k 5d_a -0.812d_b -0.812step_a 0.239step_b 0.239min_cos_consec 0.2269min_cos_anchor 0.1627dataset podcastlang enspeaker 604345track 604345total 36.9slevel spread 3.8 dBmax seam 2.8 dBcos from orange-id (speaker identity)
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a child somewhat feminine voice
(contemplation, awe, affection · slow, very low-energy, relaxed, monologue) but like it was a man giving birth and (ahem) um it was saying something about how it's linked to women going missing, primarily black women going missing
full caption & clip details
A child somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, vulnerable; reads as contemplation, awe, affection; style: monologue, casual; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.2/10; 15.1s, EN.
604345_00629224 · in -33.6 dBFS · gain +13.6 dB · podcast-03655
(fear, distress, helplessness · measured, normally alert, neutral tension, casual) Like and like when you think about a get out, like they were kidnapping black people so that they could take over their bodies so they could run and stay
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as fear, distress, helplessness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 5.4/10; 9.6s, EN.
604345_00631096 · in -32.7 dBFS · gain +12.7 dB · podcast-03666
(helplessness, astonishment surprise, fear · normal-paced, normally alert, fully relaxed, casual) all that they couldn't do like and think that that shit's not possible. Right,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as helplessness, astonishment surprise, fear; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 1.6/10; 3.5s, EN.
604345_00632440 · in -32.7 dBFS · gain +12.7 dB · podcast-03662
(impatience and irritability, disgust, contempt · brisk, energised, neutral tension, casual) craziest part. People be too uh (breathy giggle) you you got doctors who take heart transplants.
full caption & clip details
A child strongly feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, very rough, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly dominant, neutral openness; reads as impatience and irritability, disgust, contempt; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 3.5/6; vocal-burst blend 5.0/10; 4.5s, EN.
604345_00634632 · in -29.9 dBFS · gain +9.9 dB · podcast-03665
(impatience and irritability, sourness, bitterness · brisk, normally alert, neutral tension, casual) they gonna put it right back in and now you back up you that's like you remember everything.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as impatience and irritability, sourness, bitterness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 7.3/10; 3.6s, EN.
604345_00635880 · in -29.8 dBFS · gain +9.8 dB · podcast-03652
S_MONO — style: monologuevoicenet__VN1__T0.80__C0.25__INTERNAL · #8

This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with style: monologue (S_MONO) low — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the range at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.85.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.16, then +0.25, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 29 s · ko · emolia

k 5d_a 0.850d_b 0.850step_a 0.248step_b 0.248min_cos_consec min_cos_anchor dataset emolialang kospeaker KO_aWOsiMf_bxutrack KO_aWOsiMf_bxutotal 28.5slevel spread 3.9 dBmax seam 1.7 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, slightly relaxed, fairly steady, some disfluency
(helplessness, pain · fast, normally alert, average clarity, casual) 나중에 사이사이에 길이 만들어진다고 가정하면 이렇게 해야 될 것 같아요.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, pain; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.1/10; 3.6s, KO.
KO_aWOsiMf_bxu_W000103 · in -14.8 dBFS · gain -5.2 dB · emolia-03220
(measured, normally alert, slurred, casual) 일단 이거 받을게요. 운이라도 버려야 되니까. 이렇게 산대 받아주고.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.2/10; 4.4s, KO.
KO_aWOsiMf_bxu_W000104 · in -16.5 dBFS · gain -3.5 dB · emolia-03220
(confusion · measured, normally alert, slurred, casual) 하우스도 좀 더 확장을 할까요? (low mumble) 음, 2층으로 올려도 될 것 같긴 한데. 아, 그럴 필요 없구나. 자, 이거를.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.1/10; 7.4s, KO.
KO_aWOsiMf_bxu_W000105 · in -17.0 dBFS · gain -3.0 dB · emolia-03220
(doubt · measured, normally alert, slurred, monologue) 자, 연료가 부족해지니까 이쪽으로 이동해서 Fabricator Large를 만들고.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: monologue, ASMR; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 5.1/10; 5.2s, KO.
KO_aWOsiMf_bxu_W000106 · in -18.7 dBFS · gain -1.3 dB · emolia-03220
(contemplation · measured, subdued, somewhat unclear, monologue) 자, 자원이 모자르니까 좀 자원 많이 쥔 대로 좀, (low mumble) 음, 빨리 개간을 시켜야 될 것 같아요. 이런 데 있죠. 개간 시킵시다.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation; style: monologue, ASMR; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.6/10; 7.2s, KO.
KO_aWOsiMf_bxu_W000107 · in -17.3 dBFS · gain -2.7 dB · emolia-03220
ATCK — attack / onset sharpnessvoicenet__VN1__T0.80__C0.25__INTERNAL · #9

This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with attack / onset sharpness (ATCK) at the very top of the range — 0.91, higher than 91 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.80.

It takes 5 clips to get there. Clip to clip the moves are -0.11, then -0.23, then -0.25, then -0.21 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · en · emolia

k 5d_a -0.802d_b -0.802step_a 0.247step_b 0.247min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_mJIoF6iIPqMtrack EN_mJIoF6iIPqMtotal 32.7slevel spread 4.0 dBmax seam 3.9 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, fairly steady
(brisk, normally alert, neutral tension, casual) A lot of people were running close combat because you could use a fighting gem back then, but even then I didn't like using that because it would lower hitmontop's defenses.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; no dominant emotion; style: casual, storytelling; average recording, quiet background; mildly explicit content; genuineness 3.2/6; vocal-burst blend 6.5/10; 7.0s, EN.
EN_mJIoF6iIPqM_W000060 · in -13.5 dBFS · gain -6.5 dB · emolia-00656
(normal-paced, normally alert, slightly relaxed, casual) And that's the main reason to use Hitmontop over Haruyama or Scrappy. It's kind of a mixture of both.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.9/10; 4.9s, EN.
EN_mJIoF6iIPqM_W000061 · in -13.9 dBFS · gain -6.1 dB · emolia-00656
(normal-paced, normally alert, slightly relaxed, casual) (low mumble) Uhm, and they're all popular for different reasons, but hit them on top, get subscribed, these intimidate.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.6/10; 4.3s, EN.
EN_mJIoF6iIPqM_W000062 · in -13.6 dBFS · gain -6.4 dB · emolia-00656
(intoxication altered states of consciousness, embarrassment · measured, normally alert, slightly relaxed, casual) As bulky as either, (low mumble) uh, I'll just bring up my really fast here, (ahem) uh, Haruyama.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness, embarrassment; style: casual, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.0/10; 6.3s, EN.
EN_mJIoF6iIPqM_W000064 · in -17.5 dBFS · gain -2.5 dB · emolia-00656
(measured, subdued, slightly relaxed, monologue) And Scrafty, just so you can compare the stats there and see what I'm talking about. But we can see that Scrafty, of course,
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.8/10; 9.7s, EN.
EN_mJIoF6iIPqM_W000065 · in -17.5 dBFS · gain -2.5 dB · emolia-00656
STNC — stance / assertivenessvoicenet__VN1__T0.80__C0.25__INTERNAL · #10

This is a VoiceNet dimension, not an emotion: stance / assertiveness (STNC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with stance / assertiveness (STNC) at the very bottom of the range — 0.07, lower than 93 % of clips in this corpus — and ends with it high at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.82.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.24, then +0.24, then +0.12 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.58 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.64 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.58, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 67 s · en · podcast

k 5d_a 0.822d_b 0.822step_a 0.244step_b 0.244min_cos_consec 0.6403min_cos_anchor 0.5846dataset podcastlang enspeaker 802554track 802554total 67.3slevel spread 4.3 dBmax seam 4.3 dBcos from orange-id (speaker identity)
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, normally alert, somewhat unclear, light breath
(doubt, confusion, embarrassment · measured, relaxed, fairly steady, casual) (low mumble) uh yeah, every I mean that works. I mean that's not every person climbs a mountain different.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, confusion, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 1.9/10; 4.8s, EN.
802554_00222000 · in -23.9 dBFS · gain +3.9 dB · podcast-05037
(fatigue exhaustion, amusement, teasing · normal-paced, neutral tension, moderately variable, casual) life is a hard I mean everybody squeezes a lemon different, but the juice still comes out. (breathy giggle) Yeah, there you go. Unless there's shells in there and now you have to put that shit in the trash. Yeah,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as fatigue exhaustion, amusement, teasing; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 6.8/10; 13.3s, EN.
802554_00222856 · in -22.3 dBFS · gain +2.3 dB · podcast-00411
(confusion, fear, amusement · normal-paced, neutral tension, moderately variable, casual) but if you if you're like I still like I don't understand lady but like oh I'm I'm like I'm colorblind. And there was this kid, like I'm in high school. So like the kid in high school is just like (low mumble) uh he was like, You're not colorblind. Uh like, what do you mean? And like he was spreading like a false rumor around school, like that I was lying to everyone.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, fear, amusement; style: casual, conversational; average recording, quiet background; genuineness 6.0/6; vocal-burst blend 10.0/10; 17.4s, EN.
802554_00224200 · in -26.6 dBFS · gain +6.6 dB · podcast-05030
(doubt, intoxication altered states of consciousness, helplessness · normal-paced, neutral tension, moderately variable, casual) I mean, I can barely also tell. I mean, but it doesn't really matter. I I'm pretty sure that's still red green category. I don't I don't see I don't even give a fuck about being colorblind. I don't even put my I don't put that much research in looking at that shit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as doubt, intoxication altered states of consciousness, helplessness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 10.0/10; 11.7s, EN.
802554_00229375 · in -23.9 dBFS · gain +3.9 dB · podcast-05035
(amusement, elation, teasing · brisk, neutral tension, moderately variable, casual) I just know that the eye doctor said, Oh, you're colorblind deficient. You know that? I was like, (ahem) uh no. Thank you for telling me. And went to the eye doctor again. It's like, do you know? Like, yeah, the other doctor told me. Actually, in fact, you told me this. (breathy giggle) You forgot.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, elation, teasing; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 10.0/10; 19.6s, EN.
802554_00230592 · in -23.4 dBFS · gain +3.4 dB · podcast-05172
S_TECH — style: technicalvoicenet__VN1__T0.80__C0.25__INTERNAL · #11

This is a VoiceNet dimension, not an emotion: style: technical (S_TECH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with style: technical (S_TECH) at the very bottom of the range — 0.05, lower than 95 % of clips in this corpus — and ends with it at the very top of the range at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.87.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.24, then +0.20, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.42 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.42, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 62 s · es · podcast

k 5d_a 0.871d_b 0.871step_a 0.240step_b 0.240min_cos_consec 0.2427min_cos_anchor 0.4156dataset podcastlang esspeaker 138448track 138448total 61.9slevel spread 4.1 dBmax seam 4.1 dBcos from orange-id (speaker identity)
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · quiet background
(jealousy and envy, disgust, sourness · brisk, energised, neutral tension, casual) orgullo orduct a su edad, at the moment in the queen. Hoy en día lo que te doy iste ofrecen publicity and according to ubicación. El celular del capitán y el celular de la capitana está in el mismo lugar, and then the comercial en una pastilla azul, anda.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, disgust, sourness; style: casual, dramatic; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 10.0/10; 24.1s, ES.
138448_00123320 · in -33.3 dBFS · gain +13.3 dB · podcast-00876
(disappointment · fast, normally alert, slightly relaxed, casual) foto hace 10 años inicial. No, entonces esta persona le gusta viajar, no le gusta viajar, te puedo poner más comerciales de eso.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 4.1/10; 5.0s, ES.
138448_00127144 · in -37.4 dBFS · gain +17.4 dB · podcast-00858
(pride, intoxication altered states of consciousness · measured, normally alert, neutral tension, casual) Que funciona muy bien in Grecia y en Roma. Todos sabemos cómo terminó esta historia. Y bueno, ahí nos vamos. Algo más que agregar a esta
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as pride, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.2/10; 8.0s, ES.
138448_00135088 · in -33.4 dBFS · gain +13.4 dB · podcast-04540
(pride, affection · fast, normally alert, neutral tension, casual) Pues mira, más que eso, voté. Me parecía una cosa cool que un huevo le ganara a Kylie Jenner. Pero luego yo traté, y a ver, yo les quería preguntar, yo traté de hacer lo mismo y subí una foto a Instagram y me bloquearon la cuenta.
full caption & clip details
An elderly masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, affection; style: casual, playful; below-average recording, quiet background; genuineness 5.2/6; vocal-burst blend 9.7/10; 14.2s, ES.
138448_00137768 · in -34.1 dBFS · gain +14.2 dB · podcast-04560
(disgust, fear, malevolence malice · brisk, normally alert, neutral tension, casual) la cuenta se llama (purr) World-Records. Tiene ya 6.6 millones de seguidores. Solo sigue a 926 personas. Y solo tiene un post.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disgust, fear, malevolence malice; style: casual, authoritative; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 6.5/10; 10.0s, ES.
138448_00140272 · in -33.8 dBFS · gain +13.8 dB · podcast-04530
ATCK — attack / onset sharpnessvoicenet__VN1__T0.80__C0.25__INTERNAL · #12

This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with attack / onset sharpness (ATCK) low — 0.08, lower than 92 % of clips in this corpus — and ends with it high at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.82.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.23, then +0.17, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 97 s · en · podcast

k 5d_a 0.815d_b 0.815step_a 0.242step_b 0.242min_cos_consec 0.7306min_cos_anchor 0.7205dataset podcastlang enspeaker 276163track 276163total 96.8slevel spread 3.2 dBmax seam 2.1 dBcos from orange-id (speaker identity)
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult somewhat feminine voice · neutral-toned, fairly smooth, average recording, quiet background
(sourness, contempt, doubt · slow, normally alert, relaxed, conversational) well, let me move it like this. Unless you're in law enforcement, like you really.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sourness, contempt, doubt; style: conversational, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.2/10; 6.0s, EN.
276163_00194264 · in -24.0 dBFS · gain +4.0 dB · podcast-05397
(contemplation · normal-paced, normally alert, slightly relaxed, casual) There's different, like I said, there's just different professions, right? So there's different things I'm not gonna know about because I've not lived or experienced that profession or how that community interacts. But law enforcement.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.7/10; 14.6s, EN.
276163_00194868 · in -23.7 dBFS · gain +3.7 dB · podcast-05408
(contentment, shame, sourness · normal-paced, normally alert, neutral tension, casual) The way that I got in there, first off, I got in there at 22 years old. So you gotta think like it's a little 22-year-old brain going into this community of police officers that are embracing you, calling you all kind of, you know, sister, you're my sister, we're family. You know, you're in the group chat, you're invited to all the cookouts, all the weddings, all the baby showers, all the birthdays.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, shame, sourness; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 5.8/10; 23.2s, EN.
276163_00196328 · in -24.8 dBFS · gain +4.8 dB · podcast-05398
(sourness, amusement, contempt · brisk, normally alert, neutral tension, casual) I mean, that is your legit family. My mind you there might be some that you don't get along with that you don't like. (childlike giggle) Um, the systemic racism and all the other things, that's a whole other topic. But there was this camaraderie. You know, when I started, there was this camaraderie, and it was so tight, and we were so bonded. They gave me a nickname, you know. I had a nickname.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as sourness, amusement, contempt; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 7.5/10; 22.9s, EN.
276163_00198648 · in -22.7 dBFS · gain +2.7 dB · podcast-05403
(embarrassment, amusement, infatuation · brisk, energised, neutral tension, casual) I was called uh (low mumble) from my first squad, they would call me Taz, like the Tasmanian devil. Um, or they'll call me like the canine, and they would literally, my boy would grab me from like my vest and my gun belt, and he would pick me up and throw me in a car, and I would just look for the drugs. And it was like that was just like a little reckless little thing, you know. But it was just this this camaraderie, this way that we used to joke, this way that we used to be, these (ahem) um, you know, kind of morbid jokes that we used to make that
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, amusement, infatuation; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.0/10; 29.4s, EN.
276163_00200936 · in -21.6 dBFS · gain +1.6 dB · podcast-04535
VFLX — vocal flexibility / inflectionvoicenet__VN1__T0.80__C0.25__INTERNAL · #13

This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with vocal flexibility / inflection (VFLX) at the very top of the range — 0.95, higher than 95 % of clips in this corpus — and works its way down to low at 0.14, lower than 86 % of clips in this corpus. That is a total fall of 0.81.

It takes 5 clips to get there. Clip to clip the moves are -0.13, then -0.21, then -0.24, then -0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.12 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.12 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.12, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 82 s · en · podcast

k 5d_a -0.806d_b -0.806step_a 0.243step_b 0.243min_cos_consec 0.1157min_cos_anchor 0.1214dataset podcastlang enspeaker 47222track 47222total 82.5slevel spread 18.0 dBmax seam 18.0 dBcos from orange-id (speaker identity)
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, some disfluency, average clarity, light breath
(elation, contentment, pleasure ecstasy · brisk, normally alert, neutral tension, casual) pen situation size, and you can go lope circles in there and kind of do something there and go out. And (ahem) um, I'm a big fan of being outside. But like you said, that's kind of just like a privileged thing. Like where we live right now, it's really nice because we can just go ride down the road and we've, you know, from our house, there's plenty of places to go. (low mumble) Um, that technically isn't like owned by anyone. So
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as elation, contentment, pleasure ecstasy; style: casual, conversational; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.3/10; 19.9s, EN.
47222_00358560 · in -26.4 dBFS · gain +6.4 dB · podcast-04049
(shame, contentment, contemplation · normal-paced, very low-energy, relaxed, casual) Yeah, that's something that I huge kind of bucket list thing for me would be, you know, once Boudre's old enough to be able to like ride on his own is would be to just pack the trailer up and go up to like Wyoming, Montana, big sky country, and just go out and, you know, ride. (ahem) Um because that's something I've never gotten to do, and that's something I really want my son to be able to do. (surprised gasp) Um, so that's definitely on the to-do list in.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, contentment, contemplation; style: casual, ASMR; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 10.0/10; 25.4s, EN.
47222_00360576 · in -39.0 dBFS · gain +19.0 dB · podcast-04070
(thankfulness gratitude, contentment, hope enthusiasm optimism · normal-paced, normally alert, relaxed, casual) Well, you guys know we (low mumble) um you're always welcome to come with us. (breathy giggle) Come
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as thankfulness gratitude, contentment, hope enthusiasm optimism; style: casual, conversational; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 4.9/10; 4.4s, EN.
47222_00363168 · in -21.0 dBFS · gain +1.0 dB · podcast-04018
(embarrassment, contempt, amusement · brisk, energised, relaxed, casual) all up here and we'll take you guys on a little backpacking trip with them. Yeah, that's definitely definitely the goal in a few years.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, contempt, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 2.9/10; 7.1s, EN.
47222_00363600 · in -23.6 dBFS · gain +3.6 dB · podcast-04061
(doubt, jealousy and envy, amusement · brisk, normally alert, neutral tension, casual) (ahem) Um, push style horses versus free runners. But I don't know necessarily. It's just like the type of horse you run. Like, do you or do you ride? Do you would you rather a horse you have that has to be a little more push style that you have to get going? Or would you rather have a horse that you wanna like kind of have to pull back a little
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, jealousy and envy, amusement; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.0/10; 25.2s, EN.
47222_00364344 · in -24.2 dBFS · gain +4.2 dB · podcast-06478
VFLX — vocal flexibility / inflectionvoicenet__VN1__T0.80__C0.25__INTERNAL · #14

This is a VoiceNet dimension, not an emotion: vocal flexibility / inflection (VFLX) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with vocal flexibility / inflection (VFLX) low — 0.11, lower than 89 % of clips in this corpus — and ends with it at the very top of the range at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.82.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.24, then +0.24, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.31 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.31, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 56 s · es · podcast

k 5d_a 0.822d_b 0.822step_a 0.244step_b 0.244min_cos_consec 0.3122min_cos_anchor 0.3122dataset podcastlang esspeaker 194077track 194077total 56.3slevel spread 3.8 dBmax seam 2.7 dBcos from orange-id (speaker identity)
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · normally alert
(infatuation, affection, contentment · normal-paced, relaxed, fairly steady, casual) Wey, the only time you got a anti-gravity has been para accompanying my hermana. (low mumble) Yo me acuerdo que
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, very dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as infatuation, affection, contentment; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 9.4/10; 9.8s, ES.
194077_00121096 · in -29.2 dBFS · gain +9.2 dB · podcast-01542
(elation, intoxication altered states of consciousness, contentment · normal-paced, neutral tension, moderately variable, casual) hacer para ponerte un uniforme que parece, pues, pinche pinche toca de toda griega, pero pintada de naranja y azul, güey? Y unos pinches telos como de tres metros. O
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as elation, intoxication altered states of consciousness, contentment; style: casual, conversational; average recording, quiet background; genuineness 6.0/6; vocal-burst blend 10.0/10; 11.7s, ES.
194077_00139767 · in -31.3 dBFS · gain +11.3 dB · podcast-01807
(interest, jealousy and envy, teasing · brisk, neutral tension, moderately variable, casual) yo creo que más como la peluca, en sí, que más que el traje. Ándale, es que la peluca, además de que está bien pinche grande, y dependiendo de la forma del cojones, pues también depende del material, porque pues te puedes hacer un. No sé si se han visto la foto del fin de pinche, güey, de ese de Goku, o tercer fase, güey, que la tiene un racimo de plátano de metro ochenta.
full caption & clip details
A child masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, jealousy and envy, teasing; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 9.8/10; 21.4s, ES.
194077_00141224 · in -28.6 dBFS · gain +8.6 dB · podcast-01791
(fast, slightly relaxed, moderately variable, casual) que, güey, ya el simple hecho de vestirte como un personaje de (low mumble) caricatura japonesa, güey. Irte en una
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, dramatic; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 2.3/10; 7.1s, ES.
194077_00152040 · in -29.9 dBFS · gain +9.9 dB · podcast-01793
(teasing, triumph, amusement · fast, neutral tension, moderately variable, casual) Fuiste a la casa de tu compa de la secundaria, el que tenía el brazo derecho más grande que el izquierdo, güey.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as teasing, triumph, amusement; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.2/10; 5.7s, ES.
194077_00155760 · in -32.4 dBFS · gain +12.4 dB · podcast-01796
ROUG — roughnessvoicenet__VN1__T0.80__C0.25__INTERNAL · #15

This is a VoiceNet dimension, not an emotion: roughness (ROUG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with roughness (ROUG) at the very top of the range — 0.90, higher than 90 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.01, lower than 99 % of clips in this corpus. That is a total fall of 0.89.

It takes 5 clips to get there. Clip to clip the moves are -0.25, then -0.23, then -0.17, then -0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 26 s · snippets

k 5d_a -0.891d_b -0.891step_a 0.249step_b 0.249min_cos_consec min_cos_anchor dataset snippetslang ?speaker batch222_part0_batch222_patrack batch222_part0_batch222_patotal 25.9slevel spread 11.1 dBmax seam 11.1 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · normally alert
(elation · normal-paced, neutral tension, moderately variable, ranting) Oh, there are many places that you've got to in the market tomorrow. One of them, or my favorite one is called the sausage trip, right? It looks like three
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly dark, very rough, balanced body; average clarity, almost no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as elation; style: ranting, storytelling; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 3.0/10; 6.8s.
batch222_part0_batch222_part0_chunk_46_1_1175573 · in -29.6 dBFS · gain +9.6 dB · snippets-00634
(affection, hope enthusiasm optimism · measured, slightly relaxed, fairly steady, storytelling) And they know the best place for them to be safe is in the water and I want to believe that's where she has headed to.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, full; average clarity, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, neutral openness; reads as affection, hope enthusiasm optimism; style: storytelling, whispered; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 2.3/10; 6.7s.
batch222_part0_batch222_part0_chunk_46_1_1175614 · in -30.6 dBFS · gain +10.6 dB · snippets-00634
(measured, slightly relaxed, fairly steady, casual) what I want to do is to head out (ahem) uh towards the river.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 1.8/10; 4.1s.
batch222_part0_batch222_part0_chunk_46_1_1175800 · in -29.7 dBFS · gain +9.7 dB · snippets-00634
(amusement, embarrassment, doubt · normal-paced, slightly relaxed, moderately variable, casual) Once a (ahem) while, as is typical for lions or lionesses.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as amusement, embarrassment, doubt; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.4/10; 4.1s.
batch222_part0_batch222_part0_chunk_46_1_1175809 · in -28.9 dBFS · gain +8.9 dB · snippets-00634
(intoxication altered states of consciousness, fatigue exhaustion, amusement · measured, slightly relaxed, fairly steady, casual) in front of us, a couple of birds, red-billed queleas coming down to drink.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, amusement; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.4/10; 3.6s.
batch222_part0_batch222_part0_chunk_46_1_1175848 · in -39.9 dBFS · gain +19.9 dB · snippets-00634
WARM — warmthvoicenet__VN1__T0.80__C0.25__INTERNAL · #16

This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with warmth (WARM) at the very bottom of the range — 0.04, lower than 96 % of clips in this corpus — and ends with it high at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.84.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.25, then +0.20, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.51 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.51 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 41 s · en · emolia

k 5d_a 0.837d_b 0.837step_a 0.247step_b 0.247min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_73tfFiSZU8Itrack EN_73tfFiSZU8Itotal 40.6slevel spread 4.2 dBmax seam 2.6 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, normally alert
(amusement, teasing, intoxication altered states of consciousness · normal-paced, relaxed, moderately variable, casual) So you are promoting your experiments, right? (breathy giggle) Disclaimer, disclaimer.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, teasing, intoxication altered states of consciousness; style: casual, conversational; below-average recording, some background noise; genuineness 6.0/6; vocal-burst blend 0.1/10; 7.6s, EN.
EN_73tfFiSZU8I_W000388 · in -17.4 dBFS · gain -2.6 dB · emolia-01678
(doubt, interest · normal-paced, slightly relaxed, fairly steady, casual) Well, Swiggo will be very interesting (ahem) experiment. Do (low mumble) you believe this sensitivity curve is not built yet? (low mumble) It's not even fully designed yet?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, interest; style: casual, monologue; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.3/10; 8.4s, EN.
EN_73tfFiSZU8I_W000389 · in -15.9 dBFS · gain -4.1 dB · emolia-01678
(anger · measured, neutral tension, fairly steady, casual) Yeah, I mean, (surprised gasp) CTA is the opposite. CTA will not be built until (ahem) a few years from now. This is mainly SSTs in the south. Okay.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.1/10; 9.8s, EN.
EN_73tfFiSZU8I_W000393 · in -17.5 dBFS · gain -2.5 dB · emolia-01678
(measured, slightly relaxed, fairly steady, monologue) And Lasso is already taking data. They are already taking one year. So they will have an even better sensitivity for the whole sky. Okay. So there will be, I mean, for discoveries.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.2/10; 11.2s, EN.
EN_73tfFiSZU8I_W000394 · in -20.0 dBFS · gain +0.0 dB · emolia-01678
(normal-paced, slightly relaxed, fairly steady, casual) Well, also we'll be taking a lot of science here on CTA.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 0.9/10; 3.0s, EN.
EN_73tfFiSZU8I_W000395 · in -18.2 dBFS · gain -1.8 dB · emolia-01678
METL — metallic qualityvoicenet__VN1__T0.80__C0.25__INTERNAL · #17

This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with metallic quality (METL) low — 0.16, lower than 84 % of clips in this corpus — and ends with it at the very top of the range at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.80.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.17, then +0.20, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.22 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.18 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.22, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 45 s · nl · podcast

k 5d_a 0.803d_b 0.803step_a 0.237step_b 0.237min_cos_consec 0.1762min_cos_anchor 0.2238dataset podcastlang nlspeaker 189442track 189442total 45.3slevel spread 4.5 dBmax seam 4.5 dBcos from orange-id (speaker identity)
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, fairly steady, light breath
(normal-paced, normally alert, slightly relaxed, conversational) En volgens mij, (low mumble) ik had een soort van (low mumble) veganistische versie in dit geval. (low mumble) Namelijk met een soort van (low mumble) citrus (low mumble) vinicrette. En daar zat een onder andere greepvoet bij. En ook volgens mij een mandarijn.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 6.5/10; 16.2s, NL.
189442_00053944 · in -20.1 dBFS · gain +0.1 dB · podcast-05084
(longing, confusion · normal-paced, normally alert, slightly relaxed, narration) En een soort bloemkoolmoes. Die was ook heel erg lekker.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, confusion; style: narration, storytelling; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 6.6/10; 3.0s, NL.
189442_00055648 · in -22.5 dBFS · gain +2.5 dB · podcast-05084
(confusion, infatuation · measured, very low-energy, relaxed, conversational) Ja, (low mumble) die was echt (ahem) een goede combinatie, (low mumble) inderdaad. Nou, vervolgens werd er best wel veel ruimte (low mumble) tussen de gangen gegeven.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, infatuation; style: conversational, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.0/10; 10.4s, NL.
189442_00055968 · in -20.9 dBFS · gain +0.9 dB · podcast-05112
(contentment · normal-paced, normally alert, slightly relaxed, conversational) Klopt, want we mochten dus om zeven uur naar binnen en we moesten overbruggen, zeg maar, tot (ahem) oud jaar tot het aftellen. Dus (low mumble) ze hebben dan zes gangen in rustig tempo uitgeserveerd.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment; style: conversational, casual; good recording, no background noise; genuineness 3.7/6; vocal-burst blend 2.9/10; 11.8s, NL.
189442_00057040 · in -24.3 dBFS · gain +4.3 dB · podcast-05094
(intoxication altered states of consciousness · normal-paced, normally alert, slightly relaxed, casual) Overigens (low mumble) qua drank konden we eigenlijk zo'n beetje alles kiezen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.6/10; 3.2s, NL.
189442_00058279 · in -19.8 dBFS · gain -0.2 dB · podcast-05090
BKGN — background noise levelvoicenet__VN1__T0.80__C0.25__INTERNAL · #18

This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with background noise level (BKGN) at the very bottom of the range — 0.06, lower than 94 % of clips in this corpus — and ends with it high at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.80.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.19, then +0.24, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

k 5d_a 0.801d_b 0.801step_a 0.236step_b 0.236min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_PnRytmNi2Lytrack EN_PnRytmNi2Lytotal 48.2slevel spread 2.8 dBmax seam 1.8 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, balanced body
(interest, sadness · measured, very low-energy, slightly relaxed, didactic) Uh, (ahem) the species that we're concerned with are at the base of the food chain in oceans. The lower center picture of plankton, (low mumble) uh, very small, one micron, 10 micron, 100 micron.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as interest, sadness; style: didactic, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.4/10; 15.3s, EN.
EN_PnRytmNi2Ly_W000023 · in -15.9 dBFS · gain -4.1 dB · emolia-02082
(awe, astonishment surprise · slow, normally alert, slightly relaxed, casual) Tiny little things that you don't see all throughout the ocean. (low mumble) Uh, that's the base of the food chain.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, astonishment surprise; style: casual, didactic; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.8/10; 6.4s, EN.
EN_PnRytmNi2Ly_W000024 · in -15.7 dBFS · gain -4.3 dB · emolia-02082
(concentration · normal-paced, energised, neutral tension, didactic) Without those organisms, the only thing that might be left would be some jellyfish. And because these organisms are different in the sense that, from jellyfish, in the sense that they calcify, they take carbonate ion out of the water
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.6/10; 14.1s, EN.
EN_PnRytmNi2Ly_W000025 · in -14.7 dBFS · gain -5.3 dB · emolia-02082
(teasing · measured, normally alert, slightly relaxed, casual) Build calcium carbonate shells or skeletons like we have.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as teasing; style: casual, conversational; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.5/10; 5.2s, EN.
EN_PnRytmNi2Ly_W000026 · in -16.5 dBFS · gain -3.5 dB · emolia-02082
(pain, helplessness, distress · measured, normally alert, slightly relaxed, monologue) And when they die, they take that to the ocean floor. And that's carbon that's sequestered.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, helplessness, distress; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.5/10; 6.5s, EN.
EN_PnRytmNi2Ly_W000027 · in -17.5 dBFS · gain -2.5 dB · emolia-02082
ARSH — harshness of articulationvoicenet__VN1__T0.80__C0.25__INTERNAL · #19

This is a VoiceNet dimension, not an emotion: harshness of articulation (ARSH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with harshness of articulation (ARSH) at the very top of the range — 0.93, higher than 93 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.82.

It takes 5 clips to get there. Clip to clip the moves are -0.23, then -0.21, then -0.23, then -0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 38 s · en · podcast

k 5d_a -0.824d_b -0.824step_a 0.235step_b 0.235min_cos_consec 0.1567min_cos_anchor 0.0731dataset podcastlang enspeaker 863271track 863271total 37.6slevel spread 6.2 dBmax seam 6.0 dBcos from orange-id (speaker identity)
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, some disfluency
(interest, sourness, sexual lust · brisk, energised, neutral tension, conversational) demon with an axe, which can't that melt? Metal will melt. There's a reason steel Pokemon are weak to fire Pokemon. All right. I
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as interest, sourness, sexual lust; style: conversational, ranting; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.5/10; 7.3s, EN.
863271_00154760 · in -21.1 dBFS · gain +1.1 dB · podcast-02683
(sexual lust, pride, pain · normal-paced, normally alert, fully relaxed, casual) I'm a Pokemon nerd. I can tell you every type's weakness (low mumble) and
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, pride, pain; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.8/10; 3.6s, EN.
863271_00155656 · in -25.9 dBFS · gain +5.9 dB · podcast-02685
(impatience and irritability, confusion, malevolence malice · normal-paced, normally alert, neutral tension, casual) Steel loses to fire. All right. They also have a gun. This is to do with a machine gun, which in that case, I'm like, why doesn't everyone just have a machine gun?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as impatience and irritability, confusion, malevolence malice; style: casual, conversational; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.3/10; 6.7s, EN.
863271_00156152 · in -24.7 dBFS · gain +4.7 dB · podcast-02698
(disgust, sourness, amusement · normal-paced, normally alert, slightly relaxed, casual) then there's just the the holy girl who just sits there in the back and prays. She makes a like a triangle thing with a hand.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as disgust, sourness, amusement; style: casual, conversational; average recording, no background noise; genuineness 4.8/6; vocal-burst blend 4.0/10; 5.1s, EN.
863271_00156936 · in -27.3 dBFS · gain +7.3 dB · podcast-02690
(amusement, pleasure ecstasy, teasing · normal-paced, energised, relaxed, casual) Hey, she's a big fan of Jay-Z. (childlike giggle) You know? Hova. (ahem) I mean, it's a pretty fire album. Oh boy. I
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, relaxed, volatile; timbre is neutral-toned, slightly bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, pleasure ecstasy, teasing; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 4.4/6; vocal-burst blend 1.8/10; 14.4s, EN.
863271_00157480 · in -21.2 dBFS · gain +1.2 dB · podcast-00219
ATCK — attack / onset sharpnessvoicenet__VN1__T0.80__C0.25__INTERNAL · #20

This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.80.

The chain starts with attack / onset sharpness (ATCK) at the very bottom of the range — 0.04, lower than 96 % of clips in this corpus — and ends with it at the very top of the range at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.91.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.23, then +0.21, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.58 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 28 s · en · podcast

k 5d_a 0.910d_b 0.910step_a 0.238step_b 0.238min_cos_consec 0.5773min_cos_anchor 0.5936dataset podcastlang enspeaker 679068track 679068total 28.3slevel spread 5.1 dBmax seam 3.0 dBcos from orange-id (speaker identity)
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice
(doubt, embarrassment, amusement · slow, energised, neutral tension, casual) Banana Republic is like it it's refers to like a government that's not very like good. (breathy giggle) Like it it's chaos.
full caption & clip details
A young adult feminine voice; delivery is energised, slow, neutral tension, variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is negative, submissive, vulnerable; reads as doubt, embarrassment, amusement; style: casual, playful; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.5/10; 8.7s, EN.
679068_00406672 · in -34.9 dBFS · gain +14.9 dB · podcast-01561
(astonishment surprise, amusement, sourness · normal-paced, normally alert, neutral tension, casual) I thought this whole time I thought a Banana Republic was just like you know, the woman's store.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as astonishment surprise, amusement, sourness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 4.7/10; 3.8s, EN.
679068_00407848 · in -35.1 dBFS · gain +15.1 dB · podcast-01561
(teasing, intoxication altered states of consciousness, amusement · normal-paced, energised, relaxed, casual) sat on the wagon. (breathy giggle) Behind me a dragon.
full caption & clip details
An infant feminine voice; delivery is energised, normal-paced, relaxed, volatile; timbre is neutral-toned, dark, gravelly, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is negative, slightly submissive, very vulnerable; reads as teasing, intoxication altered states of consciousness, amusement; style: casual, dramatic; below-average recording, quiet background; mildly explicit content; genuineness 3.5/6; vocal-burst blend 2.1/10; 5.7s, EN.
679068_00410232 · in -36.6 dBFS · gain +16.6 dB · podcast-01561
(teasing, embarrassment, amusement · normal-paced, energised, neutral tension, casual) (ahem) Um first of all, he definitely rhymes orange with something in one of his songs. Second of all, (ahem) um what
full caption & clip details
A child somewhat masculine voice; delivery is energised, normal-paced, neutral tension, variable; timbre is slightly cool, dark, fairly smooth, slightly thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly submissive, vulnerable; reads as teasing, embarrassment, amusement; style: casual, conversational; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.1/10; 6.2s, EN.
679068_00414047 · in -34.5 dBFS · gain +14.5 dB · podcast-01574
(astonishment surprise, teasing, amusement · brisk, energised, very tense, casual) did you see that video yet? No. Bro. Oh wait. Which one?
full caption & clip details
A child strongly feminine voice; delivery is energised, brisk, very tense, volatile; timbre is neutral-toned, dark, gravelly, thin; slurred, heavy disfluency, very wide pitch range, heavy breath; affect is negative, very submissive, very vulnerable; reads as astonishment surprise, teasing, amusement; style: casual, dramatic; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 6.4/10; 3.4s, EN.
679068_00414704 · in -31.5 dBFS · gain +11.5 dB · podcast-01568