Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. one corpus in isolation Sampled from 144,000 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
TEMP — tempo ↓c-snippets-VN1 · #1
This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with tempo (TEMP) high — 0.75, higher than 75 % of clips in this corpus — and works its way down to low at 0.22, lower than 78 % of clips in this corpus. That is a total fall of 0.53.
It takes 4 clips to get there. Clip to clip the moves are -0.13, then -0.16, then -0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 20 s · snippets
hear it un-normalised (raw levels, max seam 1.8 dB)
k 4d_a -0.531d_b -0.531step_a 0.244step_b 0.244min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch241_part4_batch241_patrack batch241_part4_batch241_patotal 19.7slevel spread 2.2 dBmax seam 1.8 dB
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · balanced body, no background noise, normally alert, slightly relaxed, fairly steady, light breath
(emotional numbness · normal-paced, no disfluency, clear, dramatic)This kind of thing started to feel distinctly low tech when a new vision appeared in the 1990s.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: dramatic, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.3/10; 6.3s.
batch241_part4_batch241_part4_chunk_643_1_457838 · in -25.2 dBFS · gain +5.2 dB · snippets-00741
(normal-paced, no disfluency, clear, dramatic)Today, it's taking on much fiercer competition than shops.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, authoritative; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.3/10; 4.0s.
batch241_part4_batch241_part4_chunk_643_1_457940 · in -25.5 dBFS · gain +5.5 dB · snippets-00741
(infatuation· normal-paced, some disfluency, average clarity, casual)(ahem) Um, you can cut this, have whatever size you like.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as infatuation; style: casual, conversational; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 1.7/10; 3.3s.
batch241_part4_batch241_part4_chunk_643_1_458125 · in -27.4 dBFS · gain +7.3 dB · snippets-00741
(measured, some disfluency, somewhat unclear, casual)If you ask the people building them, you'll learn that they were actually more expensive than just buying a cheap desk.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 3.7/10; 5.6s.
batch241_part4_batch241_part4_chunk_643_1_458175 · in -27.4 dBFS · gain +7.4 dB · snippets-00741
This is a VoiceNet dimension, not an emotion: aesthetic pleasantness (ESTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with aesthetic pleasantness (ESTH) below average — 0.40, lower than 60 % of clips in this corpus — and ends with it high at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.36.
It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 14 s · snippets
hear it un-normalised (raw levels, max seam 2.1 dB)
k 3d_a 0.356d_b 0.356step_a 0.232step_b 0.232min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch79_part3_batch79_parttrack batch79_part3_batch79_parttotal 14.4slevel spread 4.2 dBmax seam 2.1 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, no background noise, normal-paced, normally alert, some disfluency, average clarity
(embarrassment · slightly relaxed, fairly steady, light breath, conversational)My list of investors is (low mumble) uh, pretty A-list.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment; style: conversational, casual; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 1.5/10; 5.3s.
batch79_part3_batch79_part3_chunk_1703_1_1520102 · in -26.2 dBFS · gain +6.2 dB · snippets-01296
(astonishment surprise, amusement, pleasure ecstasy·fully relaxed, moderately variable, normal breath, casual)I think we had the best luck. (ahem) And like it it sounds crazy but
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, amusement, pleasure ecstasy; style: casual, conversational; average recording, no background noise; explicit content; genuineness 5.1/6; vocal-burst blend 4.6/10; 3.3s.
batch79_part3_batch79_part3_chunk_1703_1_1520128 · in -28.3 dBFS · gain +8.3 dB · snippets-01296
(embarrassment, infatuation, longing·slightly relaxed, fairly steady, light breath, casual)it was time for Billy to face the music, and I definitely left out some important details along the way.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as embarrassment, infatuation, longing; style: casual, monologue; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 3.8/10; 5.5s.
batch79_part3_batch79_part3_chunk_1703_1_1520179 · in -30.4 dBFS · gain +10.4 dB · snippets-01296
RANG — pitch range used ↓c-snippets-VN1 · #3
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with pitch range used (RANG) around average — 0.43, lower than 57 % of clips in this corpus — and works its way down to low at 0.19, lower than 81 % of clips in this corpus. That is a total fall of 0.25.
It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 8 s · snippets
hear it un-normalised (raw levels, max seam 7.9 dB)
k 2d_a -0.246d_b -0.246step_a 0.246step_b 0.246min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch270_part3_batch270_patrack batch270_part3_batch270_patotal 8.5slevel spread 7.9 dBmax seam 7.9 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(sourness, helplessness · some disfluency, average clarity, conversational, casual)to use his own death so that fellow Republicans could meet en masse at his funeral.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sourness, helplessness; style: conversational, casual; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 2.4/10; 4.3s.
batch270_part3_batch270_part3_chunk_897_1_722675 · in -22.7 dBFS · gain +2.7 dB · snippets-00891
(emotional numbness·no disfluency, clear, monologue, formal)prepared himself for the entrance exams to prestigious Ecole Polytechnique.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, formal; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.6/10; 4.0s.
batch270_part3_batch270_part3_chunk_897_1_722691 · in -30.6 dBFS · gain +10.6 dB · snippets-00891
S_ASMR — style: asmr ↑c-snippets-VN1 · #4
This is a VoiceNet dimension, not an emotion: style: asmr (S_ASMR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: asmr (S_ASMR) low — 0.20, lower than 80 % of clips in this corpus — and ends with it above average at 0.71, higher than 71 % of clips in this corpus. That is a total rise of 0.51.
It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.25, then +0.18 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · snippets
hear it un-normalised (raw levels, max seam 10.0 dB)
k 4d_a 0.512d_b 0.512step_a 0.246step_b 0.246min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch166_part2_batch166_patrack batch166_part2_batch166_patotal 33.5slevel spread 13.1 dBmax seam 10.0 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, average recording, quiet background
(sexual lust, sourness, interest · normal-paced, normally alert, neutral tension, casual)I mean like, is anyone, is there any scientific facts on like soap is better for you than just rinsing? I think we know that soap is antibacterial. Not all of it, like not like the
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, sourness, interest; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 7.2/10; 10.9s.
batch166_part2_batch166_part2_chunk_2492_1_2055509 · in -29.3 dBFS · gain +9.3 dB · snippets-00350
(disgust, intoxication altered states of consciousness, pleasure ecstasy· normal-paced, energised, relaxed, casual)people have decided that the brand (chuckle) is really gross. (low mumble) um when they go Bert pool showers. I pool I pool showered my whole life. Is it is it a salt water pool?
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as disgust, intoxication altered states of consciousness, pleasure ecstasy; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 0.8/10; 12.0s.
batch166_part2_batch166_part2_chunk_2492_1_2055547 · in -26.2 dBFS · gain +6.2 dB · snippets-00350
(sexual lust, infatuation, intoxication altered states of consciousness ·measured, normally alert, relaxed, casual)weed, you know? So, it's (low mumble) um, it just gets me ready to go to sleep. So I love it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, infatuation, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.9/10; 4.4s.
batch166_part2_batch166_part2_chunk_2492_1_2055695 · in -36.2 dBFS · gain +16.2 dB · snippets-00350
(helplessness, impatience and irritability, fear· measured, normally alert, relaxed, casual)It's gotten, (contented sigh) it's not doing very well on Instagram. No? No, because not everyone's swiping, they don't know what to do.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, steady; timbre is neutral-toned, dark, slightly rough, thin; slurred, some disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as helplessness, impatience and irritability, fear; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.8/10; 5.6s.
batch166_part2_batch166_part2_chunk_2492_1_2055716 · in -39.2 dBFS · gain +19.2 dB · snippets-00350
METL — metallic quality ↓c-snippets-VN1 · #5
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with metallic quality (METL) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to around average at 0.56, higher than 56 % of clips in this corpus. That is a total fall of 0.37.
It takes 4 clips to get there. Clip to clip the moves are -0.19, then -0.07, then -0.12 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 27 s · snippets
hear it un-normalised (raw levels, max seam 1.0 dB)
k 4d_a -0.374d_b -0.374step_a 0.188step_b 0.188min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch270_part3_batch270_patrack batch270_part3_batch270_patotal 27.4slevel spread 1.2 dBmax seam 1.0 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, quiet background, normally alert, moderate pitch range, light breath
(measured, slightly relaxed, fairly steady, casual)uh (ahem) Chola Temple which uh (ahem) (ahem) recently celebrated
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 1.9/10; 3.9s.
batch270_part3_batch270_part3_chunk_902_1_799663 · in -17.1 dBFS · gain -2.9 dB · snippets-00891
(normal-paced, slightly relaxed, fairly steady, didactic)find some patterns in in India's architecture, patterns of geometry and certain concepts which we'll come to.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.3/10; 7.4s.
batch270_part3_batch270_part3_chunk_902_1_799887 · in -16.8 dBFS · gain -3.2 dB · snippets-00891
(contemplation, concentration· normal-paced, neutral tension, moderately variable, casual)is symbolizes the outer world. Up to this point, you're still in the outer world because anything can take place here. As I said, discourses, cultural activities.
full caption & clip details
An elderly masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, concentration; style: casual, playful; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.3/10; 10.8s.
batch270_part3_batch270_part3_chunk_902_1_799950 · in -17.0 dBFS · gain -3.0 dB · snippets-00891
(emotional numbness· normal-paced, slightly relaxed, fairly steady, casual)(ahem) it all started from this desire to construct complex alters like this one.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: casual, playful; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.5/10; 4.9s.
batch270_part3_batch270_part3_chunk_902_1_800061 · in -18.0 dBFS · gain -2.0 dB · snippets-00891
R_ORAL — resonance: oral ↓c-snippets-VN1 · #6
This is a VoiceNet dimension, not an emotion: resonance: oral (R_ORAL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: oral (R_ORAL) around average — 0.56, higher than 56 % of clips in this corpus — and works its way down to low at 0.22, lower than 78 % of clips in this corpus. That is a total fall of 0.34.
It takes 4 clips to get there. Clip to clip the moves are -0.18, then -0.13, then -0.03 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 21 s · snippets
k 4d_a -0.342d_b -0.342step_a 0.177step_b 0.177min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch92_part4_batch92_parttrack batch92_part4_batch92_parttotal 21.2slevel spread 5.8 dBmax seam 4.6 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(no disfluency, clear, formal, narration)The scanner projects low intensity X-ray beams at high speed.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 3.9s.
batch92_part4_batch92_part4_chunk_1829_1_1972400 · in -27.7 dBFS · gain +7.7 dB · snippets-01370
(no disfluency, clear, formal, narration)The narcotics unit shifts their focus to the next international departure.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.7/10; 4.3s.
batch92_part4_batch92_part4_chunk_1829_1_1972441 · in -29.0 dBFS · gain +9.0 dB · snippets-01370
(fear·almost no disfluency, clear, formal, narration)Once there, a radiologist will get him hospitalized where the suspect will be under observation until he naturally passes each one of the objects out of his body.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.2s.
batch92_part4_batch92_part4_chunk_1829_1_1972547 · in -28.9 dBFS · gain +8.9 dB · snippets-01370
(distress, sadness, helplessness·some disfluency, average clarity, conversational, casual)Es el mismo metal que reacciona con botas alemanas. Es un poco
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, sadness, helplessness; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.2/10; 3.4s.
batch92_part4_batch92_part4_chunk_1829_1_1972564 · in -33.5 dBFS · gain +13.5 dB · snippets-01370
R_MASK — resonance: mask ↑c-snippets-VN1 · #7
This is a VoiceNet dimension, not an emotion: resonance: mask (R_MASK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: mask (R_MASK) around average — 0.44, lower than 56 % of clips in this corpus — and ends with it high at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.43.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.22, then -0.15, then +0.17 — not a clean run: step 3 moves back the other way by 0.15 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 26 s · snippets
k 5d_a 0.426d_b 0.426step_a 0.219step_b 0.219min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch262_part1_batch262_patrack batch262_part1_batch262_patotal 26.1slevel spread 5.6 dBmax seam 3.0 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, moderate pitch range
(measured, fairly steady, little disfluency, formal)In a way it was a kind of a race to match the Russians.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.0/10; 3.2s.
batch262_part1_batch262_part1_chunk_831_1_1105953 · in -21.8 dBFS · gain +1.8 dB · snippets-00845
(measured, fairly steady, almost no disfluency, monologue)launched experiments as high as 135 miles, then tilted downward and fired its third and fourth stages toward Earth.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.1/10; 8.2s.
batch262_part1_batch262_part1_chunk_831_1_1106039 · in -24.4 dBFS · gain +4.4 dB · snippets-00845
(normal-paced, fairly steady, no disfluency, formal)And trouble arises when the radar tracking indicates that the rocket is off course.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.1/10; 4.6s.
batch262_part1_batch262_part1_chunk_831_1_1106054 · in -27.4 dBFS · gain +7.4 dB · snippets-00845
(pride·measured, steady, almost no disfluency, formal)These scout people created a launch vehicle that set a standard for productivity and reliability.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.7/10; 5.9s.
batch262_part1_batch262_part1_chunk_831_1_1106067 · in -25.9 dBFS · gain +5.9 dB · snippets-00845
(fatigue exhaustion, helplessness·normal-paced, steady, no disfluency, formal)Spin up motors and separation sections needed for flight.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, helplessness; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.4/10; 3.6s.
batch262_part1_batch262_part1_chunk_831_1_1106136 · in -25.3 dBFS · gain +5.3 dB · snippets-00845
RANG — pitch range used ↑c-snippets-VN1 · #8
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with pitch range used (RANG) low — 0.17, lower than 83 % of clips in this corpus — and ends with it below average at 0.40, lower than 60 % of clips in this corpus. That is a total rise of 0.24.
It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 8 s · snippets
k 2d_a 0.236d_b 0.236step_a 0.236step_b 0.236min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch163_part4_batch163_patrack batch163_part4_batch163_patotal 7.5slevel spread 8.3 dBmax seam 8.3 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(affection, embarrassment · slow, frequent disfluency, slurred, whispered)But he's a very sad character actually. And (low mumble) um,
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; reads as affection, embarrassment; style: whispered, casual; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.5/10; 3.2s.
batch163_part4_batch163_part4_chunk_245_1_136268 · in -32.1 dBFS · gain +12.1 dB · snippets-00337
(sadness, emotional numbness, longing·normal-paced, no disfluency, clear, narration)She was working for herself in the same Amsterdam canalside window.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sadness, emotional numbness, longing; style: narration, storytelling; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.7/10; 4.2s.
batch163_part4_batch163_part4_chunk_245_1_136432 · in -23.8 dBFS · gain +3.8 dB · snippets-00337
S_CASU — style: casual ↓c-snippets-VN1 · #9
This is a VoiceNet dimension, not an emotion: style: casual (S_CASU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: casual (S_CASU) high — 0.78, higher than 78 % of clips in this corpus — and works its way down to low at 0.24, lower than 76 % of clips in this corpus. That is a total fall of 0.54.
It takes 4 clips to get there. Clip to clip the moves are -0.16, then -0.19, then -0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 14 s · snippets
k 4d_a -0.536d_b -0.536step_a 0.192step_b 0.192min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch254_part0_batch254_patrack batch254_part0_batch254_patotal 13.9slevel spread 7.2 dBmax seam 7.2 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, slightly relaxed, light breath
(pain · normal-paced, normally alert, fairly steady, casual)Mă simt ca față de anul trecut, mă simt mai puternică.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain; style: casual, conversational; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 0.0/10; 3.2s.
batch254_part0_batch254_part0_chunk_752_1_1355819 · in -28.2 dBFS · gain +8.2 dB · snippets-00803
(helplessness, longing, sadness·measured, subdued, fairly steady, whispered)eu nu mai am nevoie de persoanele din anturajul ăla, nu le mai simt lipsa.
batch254_part0_batch254_part0_chunk_752_1_1355878 · in -32.8 dBFS · gain +12.8 dB · snippets-00803
(sadness, distress, helplessness · measured, normally alert, fairly steady, storytelling)Elena was left for dead by her traffickers.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as sadness, distress, helplessness; style: storytelling, narration; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 4.7/10; 3.0s.
batch254_part0_batch254_part0_chunk_752_1_1355887 · in -25.6 dBFS · gain +5.6 dB · snippets-00803
(pain· measured, subdued, steady, monologue)they are on the streets since they are 13.
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, ASMR; very good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.7/10; 3.2s.
batch254_part0_batch254_part0_chunk_752_1_1355909 · in -27.8 dBFS · gain +7.8 dB · snippets-00803
R_MIXD — resonance: mixed ↑c-snippets-VN1 · #10
This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: mixed (R_MIXD) below average — 0.31, lower than 69 % of clips in this corpus — and ends with it high at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.53.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.20, then +0.13, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 29 s · snippets
k 5d_a 0.529d_b 0.529step_a 0.209step_b 0.209min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch76_part3_batch76_parttrack batch76_part3_batch76_parttotal 29.2slevel spread 6.6 dBmax seam 4.8 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · fairly smooth, normally alert, slightly relaxed, moderate pitch range, light breath
(thankfulness gratitude, affection, awe · measured, fairly steady, frequent disfluency, casual)to ask advice and to seek guidance on a matter I feel only you.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, affection, awe; style: casual, formal; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.1/10; 3.8s.
batch76_part3_batch76_part3_chunk_1685_1_1929723 · in -18.7 dBFS · gain -1.3 dB · snippets-01281
(brisk, steady, no disfluency, narration)much simpler methods such as linear classifiers gradually overtook neural networks in machine learning popularity.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, steady; timbre is slightly warm, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: narration, whispered; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.5/10; 6.6s.
batch76_part3_batch76_part3_chunk_1685_1_1929897 · in -23.5 dBFS · gain +3.5 dB · snippets-01281
(measured, steady, no disfluency, whispered)This model paved the way for neural network research to split into two approaches.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; no dominant emotion; style: whispered, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.3/10; 4.6s.
batch76_part3_batch76_part3_chunk_1685_1_1929911 · in -24.1 dBFS · gain +4.1 dB · snippets-01281
(emotional numbness· measured, steady, no disfluency, whispered)the SP 2012 segmentation of neuronal structures in M stacks challenge, the image net competition and others.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: whispered, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.2/10; 7.1s.
batch76_part3_batch76_part3_chunk_1685_1_1929926 · in -25.3 dBFS · gain +5.3 dB · snippets-01281
(measured, fairly steady, no disfluency, whispered)However, using neural networks transformed some domains, such as the prediction of protein structures.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, narration; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.6/10; 6.4s.
batch76_part3_batch76_part3_chunk_1685_1_1929942 · in -24.1 dBFS · gain +4.1 dB · snippets-01281
METL — metallic quality ↑c-snippets-VN1 · #11
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with metallic quality (METL) low — 0.21, lower than 79 % of clips in this corpus — and ends with it above average at 0.64, higher than 64 % of clips in this corpus. That is a total rise of 0.43.
It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.02, then +0.22, then -0.06 — not a clean run: step 4 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 43 s · snippets
k 5d_a 0.428d_b 0.428step_a 0.248step_b 0.248min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch139_part2_batch139_patrack batch139_part2_batch139_patotal 42.9slevel spread 4.0 dBmax seam 4.0 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · normally alert
(measured, slightly relaxed, steady, newsreading)The post office started to contract with private stagecoach companies to carry mail, and these companies worked with cities and towns to build roads. In this way, the post helped carve out the early transportation infrastructure of the country, connecting disparate communities.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 14.8s.
batch139_part2_batch139_part2_chunk_2246_1_2740313 · in -22.7 dBFS · gain +2.7 dB · snippets-00203
(elation, interest, astonishment surprise·normal-paced, neutral tension, fairly steady, casual)Big data. Oh, big data. So, researchers have been trying to solve this problem for years before New Jersey starts messing around with bail reform.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation, interest, astonishment surprise; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.9/10; 9.1s.
batch139_part2_batch139_part2_chunk_2246_1_2740320 · in -24.4 dBFS · gain +4.5 dB · snippets-00203
(emotional numbness, sexual lust, bitterness·measured, slightly relaxed, steady, storytelling)They act rationally for purposes they agree on, such as assassinating me.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, very full; average clarity, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as emotional numbness, sexual lust, bitterness; style: storytelling, whispered; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.9/10; 5.8s.
batch139_part2_batch139_part2_chunk_2246_1_2740334 · in -23.2 dBFS · gain +3.2 dB · snippets-00203
(longing·normal-paced, slightly relaxed, steady, whispered)In the early 60s, Engelbart got funding to start his own lab where he could experiment with all of his ideas.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as longing; style: whispered, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.7/10; 6.5s.
batch139_part2_batch139_part2_chunk_2246_1_2740386 · in -20.8 dBFS · gain +0.8 dB · snippets-00203
(contempt, sourness, teasing· normal-paced, slightly relaxed, steady, narration)But surely the prime minister does not make a man a duke solely in order to blast him with a charge of bribery.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, sourness, teasing; style: narration, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.1/10; 6.1s.
batch139_part2_batch139_part2_chunk_2246_1_2740500 · in -24.9 dBFS · gain +4.9 dB · snippets-00203
S_CART — style: cartoonish ↓c-snippets-VN1 · #12
This is a VoiceNet dimension, not an emotion: style: cartoonish (S_CART) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: cartoonish (S_CART) above average — 0.66, higher than 66 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.05, lower than 95 % of clips in this corpus. That is a total fall of 0.61.
It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.16, then -0.15, then -0.11 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · snippets
k 5d_a -0.614d_b -0.614step_a 0.197step_b 0.197min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch13_part1_batch13_parttrack batch13_part1_batch13_parttotal 31.7slevel spread 4.8 dBmax seam 4.3 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normally alert, moderate pitch range
(emotional numbness, contempt, impatience and irritability · normal-paced, slightly relaxed, fairly steady, whispered)talk about the intersections, then you're not dealing with it.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, contempt, impatience and irritability; style: whispered, storytelling; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.8/10; 4.1s.
batch13_part1_batch13_part1_chunk_1110_1_1104390 · in -30.5 dBFS · gain +10.5 dB · snippets-00207
(thankfulness gratitude· normal-paced, slightly relaxed, fairly steady, conversational)Great, thank you for that. And then how do you define disability?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude; style: conversational, casual; good recording, no background noise; genuineness 3.6/6; vocal-burst blend 4.3/10; 3.1s.
batch13_part1_batch13_part1_chunk_1110_1_1104516 · in -26.2 dBFS · gain +6.2 dB · snippets-00207
(impatience and irritability, contempt·brisk, slightly relaxed, moderately variable, casual)It's also true what you're saying that like many employers do not accept that responsibility or do not think about it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, contempt; style: casual, conversational; good recording, no background noise; genuineness 3.9/6; vocal-burst blend 5.1/10; 6.6s.
batch13_part1_batch13_part1_chunk_1110_1_1104530 · in -28.1 dBFS · gain +8.1 dB · snippets-00207
(infatuation·normal-paced, slightly relaxed, fairly steady, casual)and it's a federal law and person first versus identity first central language is pretty self-explanatory it's the idea that you
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as infatuation; style: casual, conversational; good recording, no background noise; genuineness 3.9/6; vocal-burst blend 3.9/10; 7.8s.
batch13_part1_batch13_part1_chunk_1110_1_1104588 · in -25.7 dBFS · gain +5.7 dB · snippets-00207
(distress, fear, sadness· normal-paced, neutral tension, moderately variable, casual)It doesn't just happen one morning that you wake up and you're like, I need a mental health break, I'm burnt out. There was a whole period of time that led up to that. And so
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as distress, fear, sadness; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 10.0/10; 9.5s.
batch13_part1_batch13_part1_chunk_1110_1_1104870 · in -25.9 dBFS · gain +5.8 dB · snippets-00207
R_THRT — resonance: throat ↑c-snippets-VN1 · #13
This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: throat (R_THRT) around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the range at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.37.
It takes 4 clips to get there. Clip to clip the moves are +0.05, then +0.24, then +0.09 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 20 s · snippets
k 4d_a 0.375d_b 0.375step_a 0.236step_b 0.236min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch41_part1_batch41_parttrack batch41_part1_batch41_parttotal 20.2slevel spread 3.2 dBmax seam 3.2 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · no background noise
(measured, normally alert, slightly relaxed, whispered)to, (low mumble) um, practice, literally, to practice the performance that's up and coming.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.8/10; 6.2s.
batch41_part1_batch41_part1_chunk_1364_1_1207732 · in -20.5 dBFS · gain +0.5 dB · snippets-01104
(pain·slow, normally alert, fully relaxed, casual)(ahem) me. (low mumble) Um, the race is (low mumble) uh smooth.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, fully relaxed, steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; style: casual, monologue; average recording, no background noise; reads as pain; genuineness 3.4/6; vocal-burst blend 1.5/10; 3.3s.
batch41_part1_batch41_part1_chunk_1364_1_1207747 · in -20.2 dBFS · gain +0.2 dB · snippets-01104
(measured, normally alert, slightly relaxed, casual)in perhaps a transformed mode throughout our lives.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 2.9/10; 3.2s.
batch41_part1_batch41_part1_chunk_1364_1_1207800 · in -18.3 dBFS · gain -1.7 dB · snippets-01104
(doubt·slow, very low-energy, relaxed, narration)What's the architecture? I guess it's sort of a modern or is it an ancient? It's an old thing. But that was only two of the four pairings.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt; style: narration, whispered; below-average recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.0/10; 7.0s.
batch41_part1_batch41_part1_chunk_1364_1_1207813 · in -21.5 dBFS · gain +1.5 dB · snippets-01104
S_NEWS — style: newsreading ↓c-snippets-VN1 · #14
This is a VoiceNet dimension, not an emotion: style: newsreading (S_NEWS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: newsreading (S_NEWS) above average — 0.70, higher than 70 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.59.
It takes 5 clips to get there. Clip to clip the moves are -0.22, then -0.12, then -0.01, then -0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 35 s · snippets
k 5d_a -0.590d_b -0.590step_a 0.245step_b 0.245min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch45_part1_batch45_parttrack batch45_part1_batch45_parttotal 34.6slevel spread 4.3 dBmax seam 4.3 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, moderate pitch range, light breath
(awe · fairly steady, some disfluency, clear, whispered)It's just that we are reaching the plateau where we can see the patterns and come up with the codes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe; style: whispered, casual; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.6/10; 5.6s.
batch45_part1_batch45_part1_chunk_1402_1_1767450 · in -22.6 dBFS · gain +2.6 dB · snippets-01121
(contemplation, astonishment surprise, emotional numbness·moderately variable, some disfluency, average clarity, casual)(ahem) uh of an interview after he was done speaking and it was something that if you weren't paying attention to it it it it happened so quickly in the interview and the person interviewing him totally skipped over they didn't they didn't really capture what he was saying
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, astonishment surprise, emotional numbness; style: casual, didactic; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.2/10; 14.8s.
batch45_part1_batch45_part1_chunk_1402_1_1767557 · in -21.5 dBFS · gain +1.5 dB · snippets-01121
(contemplation, doubt·fairly steady, some disfluency, average clarity, casual)What what do you make of that similarity and and if if you do see anything in that similarity?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, doubt; style: casual, conversational; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.5/10; 5.9s.
batch45_part1_batch45_part1_chunk_1402_1_1767796 · in -24.5 dBFS · gain +4.5 dB · snippets-01121
(fairly steady, some disfluency, average clarity, casual)as fast as those jobs are moving away. What's going to happen is
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.4/10; 3.8s.
batch45_part1_batch45_part1_chunk_1402_1_1767858 · in -25.0 dBFS · gain +5.0 dB · snippets-01121
(doubt, confusion, contemplation·moderately variable, frequent disfluency, somewhat unclear, conversational)(low mumble) Um, I think so. Yeah, I think so. I mean it it gives
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as doubt, confusion, contemplation; style: conversational, casual; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.5/10; 3.9s.
batch45_part1_batch45_part1_chunk_1402_1_1767893 · in -20.7 dBFS · gain +0.7 dB · snippets-01121
RANG — pitch range used ↑c-snippets-VN1 · #15
This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with pitch range used (RANG) around average — 0.55, higher than 55 % of clips in this corpus — and ends with it high at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.24.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.00 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 11 s · snippets
k 3d_a 0.244d_b 0.244step_a 0.240step_b 0.240min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch191_part0_batch191_patrack batch191_part0_batch191_patotal 11.1slevel spread 2.0 dBmax seam 2.0 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, normally alert, moderate pitch range, light breath
(brisk, slightly relaxed, fairly steady, casual)Failing to produce discovery in a timely manner.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.0/10; 3.1s.
batch191_part0_batch191_part0_chunk_2712_1_2655521 · in -19.8 dBFS · gain -0.2 dB · snippets-00476
(pleasure ecstasy, awe, elation·normal-paced, relaxed, moderately variable, casual)It was very powerful. It was very revealing.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, slightly vulnerable; reads as pleasure ecstasy, awe, elation; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 4.6/10; 3.4s.
batch191_part0_batch191_part0_chunk_2712_1_2655555 · in -21.7 dBFS · gain +1.7 dB · snippets-00476
(normal-paced, neutral tension, fairly steady, casual)And if Barry hid Suzanne's body during that four-hour window.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.4/10; 4.3s.
batch191_part0_batch191_part0_chunk_2712_1_2655565 · in -19.7 dBFS · gain -0.3 dB · snippets-00476
FULL — fullness of tone ↓c-snippets-VN1 · #16
This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with fullness of tone (FULL) around average — 0.49, right about the corpus median — and works its way down to at the very bottom of the range at 0.02, lower than 98 % of clips in this corpus. That is a total fall of 0.47.
It takes 5 clips to get there. Clip to clip the moves are -0.15, then +0.11, then -0.20, then -0.23 — not a clean run: step 2 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 26 s · snippets
k 5d_a -0.474d_b -0.474step_a 0.233step_b 0.233min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch111_part2_batch111_patrack batch111_part2_batch111_patotal 26.0slevel spread 6.3 dBmax seam 4.7 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice
(relief, thankfulness gratitude, emotional numbness · measured, normally alert, slightly relaxed, storytelling)I have instructed my assistant to have paid into your...
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, dark, very rough, very full; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, dominant, fairly guarded; reads as relief, thankfulness gratitude, emotional numbness; style: storytelling, whispered; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 2.1/10; 3.2s.
batch111_part2_batch111_part2_chunk_1998_1_1950416 · in -35.7 dBFS · gain +15.7 dB · snippets-00060
(emotional numbness, infatuation· measured, normally alert, slightly relaxed, casual)To reach Calba, you must first make contact with a man called Fekesh.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, infatuation; style: casual, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 1.6/10; 4.0s.
batch111_part2_batch111_part2_chunk_1998_1_1950552 · in -34.1 dBFS · gain +14.1 dB · snippets-00060
(longing, sadness, disappointment·fast, energised, slightly tense, ranting)this nights have passed since we came to Cino, but although our industriousness has been a solid ten, our net results have amounted to less than zip.
full caption & clip details
An adult masculine voice; delivery is energised, fast, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, guarded; reads as longing, sadness, disappointment; style: ranting, cartoonish; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.7/10; 7.2s.
batch111_part2_batch111_part2_chunk_1998_1_1950581 · in -29.4 dBFS · gain +9.4 dB · snippets-00060
(triumph, hope enthusiasm optimism, elation· fast, frantic, very tense, cartoonish)Time for this. I've got to come up with a new contest style stat. (childlike giggle) Hold the brakes, Sparky. What about putting our training advice?
full caption & clip details
A child somewhat masculine voice; delivery is frantic, fast, very tense, volatile; timbre is slightly cool, bright, very rough, thin; clear, no disfluency, very wide pitch range, normal breath; affect is elated, very dominant, guarded; reads as triumph, hope enthusiasm optimism, elation; style: cartoonish, ranting; below-average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.8/10; 7.6s.
batch111_part2_batch111_part2_chunk_1998_1_1950754 · in -29.4 dBFS · gain +9.4 dB · snippets-00060
(teasing, impatience and irritability, malevolence malice·brisk, energised, slightly tense, storytelling)But don't let your guard down for a second when you battle against Paul.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, fairly guarded; reads as teasing, impatience and irritability, malevolence malice; style: storytelling, casual; good recording, no background noise; mildly explicit content; genuineness 1.1/6; vocal-burst blend 4.3/10; 3.5s.
batch111_part2_batch111_part2_chunk_1998_1_1950885 · in -31.2 dBFS · gain +11.2 dB · snippets-00060
R_HEAD — resonance: head ↓c-snippets-VN1 · #17
This is a VoiceNet dimension, not an emotion: resonance: head (R_HEAD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: head (R_HEAD) high — 0.85, higher than 85 % of clips in this corpus — and works its way down to below average at 0.34, lower than 66 % of clips in this corpus. That is a total fall of 0.50.
It takes 5 clips to get there. Clip to clip the moves are -0.18, then -0.15, then -0.04, then -0.13 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 19 s · snippets
k 5d_a -0.503d_b -0.503step_a 0.177step_b 0.177min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch48_part1_batch48_parttrack batch48_part1_batch48_parttotal 18.8slevel spread 3.5 dBmax seam 3.5 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, monologue)and that we go very slowly and with high rotation speed.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.8/10; 3.6s.
batch48_part1_batch48_part1_chunk_1425_1_1551296 · in -26.1 dBFS · gain +6.1 dB · snippets-01134
(formal, monologue)Gdansk, Poland is home to one of the largest sailmakers in Europe.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.3/10; 3.8s.
batch48_part1_batch48_part1_chunk_1425_1_1551399 · in -22.6 dBFS · gain +2.6 dB · snippets-01134
(formal, monologue)508 passengers have reached their destination.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.9/10; 3.2s.
batch48_part1_batch48_part1_chunk_1425_1_1551419 · in -23.7 dBFS · gain +3.7 dB · snippets-01134
(narration, formal)Historian, Sean Bohannan, is an expert in the history of the bomber.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.4/10; 4.4s.
batch48_part1_batch48_part1_chunk_1425_1_1551447 · in -23.4 dBFS · gain +3.4 dB · snippets-01134
(monologue, narration)Nevertheless, it can transport many more containers.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 4.6/10; 3.3s.
batch48_part1_batch48_part1_chunk_1425_1_1551459 · in -23.5 dBFS · gain +3.5 dB · snippets-01134
R_NASL — resonance: nasal ↓c-snippets-VN1 · #18
This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: nasal (R_NASL) below average — 0.35, lower than 65 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.08, lower than 92 % of clips in this corpus. That is a total fall of 0.27.
It takes 4 clips to get there. Clip to clip the moves are -0.24, then -0.01, then -0.02 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 22 s · snippets
k 4d_a -0.271d_b -0.271step_a 0.245step_b 0.245min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch41_part4_batch41_parttrack batch41_part4_batch41_parttotal 21.6slevel spread 3.2 dBmax seam 3.2 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(helplessness, sadness, distress · steady, no disfluency, narration, monologue)Restless legs syndrome is a disorder in which the sufferer feels uncomfortable unless they keep their legs, and sometimes arms or torso.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, sadness, distress; style: narration, monologue; good recording, no background noise; mildly explicit content; genuineness 0.3/6; vocal-burst blend 1.5/10; 6.8s.
batch41_part4_batch41_part4_chunk_1366_1_1243196 · in -22.8 dBFS · gain +2.8 dB · snippets-01106
(fear·fairly steady, almost no disfluency, casual, monologue)It is important to seek medical attention if you suspect you have bruxism to prevent long-term damage to your teeth and jaw.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: casual, monologue; good recording, no background noise; explicit content; genuineness 1.4/6; vocal-burst blend 1.6/10; 5.9s.
batch41_part4_batch41_part4_chunk_1366_1_1243211 · in -26.0 dBFS · gain +6.0 dB · snippets-01106
(pain, emotional numbness, helplessness· fairly steady, no disfluency, monologue, formal)So the symptoms are alleviated with amphetamine based drugs.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, emotional numbness, helplessness; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 4.8/10; 3.1s.
batch41_part4_batch41_part4_chunk_1366_1_1243251 · in -23.9 dBFS · gain +3.9 dB · snippets-01106
(fairly steady, no disfluency, narration, monologue)We will explore some of the most common and bizarre sleep disorders that affect millions of people worldwide.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; explicit content; genuineness 0.0/6; vocal-burst blend 3.3/10; 5.3s.
batch41_part4_batch41_part4_chunk_1366_1_1243305 · in -22.8 dBFS · gain +2.8 dB · snippets-01106
FULL — fullness of tone ↑c-snippets-VN1 · #19
This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with fullness of tone (FULL) around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the range at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.43.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 17 s · snippets
k 3d_a 0.434d_b 0.434step_a 0.242step_b 0.242min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch13_part1_batch13_parttrack batch13_part1_batch13_parttotal 16.6slevel spread 0.7 dBmax seam 0.7 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, normally alert, clear, light breath
(emotional numbness · brisk, neutral tension, fairly steady, casual)It is stretched yet more when it's rolled up, adding a few kilometers.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness; style: casual, storytelling; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.0/10; 4.0s.
batch13_part1_batch13_part1_chunk_1114_1_1155305 · in -17.1 dBFS · gain -2.9 dB · snippets-00207
(measured, slightly relaxed, moderately variable, storytelling)It ends here in Norden, in this building. Hier bei uns in Norden hier in diesem Gebäude.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, very smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: storytelling, dramatic; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.5/10; 4.6s.
batch13_part1_batch13_part1_chunk_1114_1_1155378 · in -17.7 dBFS · gain -2.3 dB · snippets-00207
(relief, triumph·normal-paced, slightly relaxed, fairly steady, formal)After many failed attempts, the first transatlantic cable was finally successfully laid between Ireland and Newfoundland.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, triumph; style: formal, newsreading; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.4/10; 7.8s.
batch13_part1_batch13_part1_chunk_1114_1_1155391 · in -17.0 dBFS · gain -3.0 dB · snippets-00207
R_MASK — resonance: mask ↓c-snippets-VN1 · #20
This is a VoiceNet dimension, not an emotion: resonance: mask (R_MASK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: mask (R_MASK) high — 0.80, higher than 80 % of clips in this corpus — and works its way down to around average at 0.47, lower than 53 % of clips in this corpus. That is a total fall of 0.32.
It takes 3 clips to get there. Clip to clip the moves are -0.14, then -0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 68 s · snippets
k 3d_a -0.324d_b -0.324step_a 0.181step_b 0.181min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch59_part3_batch59_parttrack batch59_part3_batch59_parttotal 68.2slevel spread 2.7 dBmax seam 2.7 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, slightly relaxed
(concentration, doubt, contemplation · subdued, fairly steady, some disfluency, monologue)some of the exposure that we've had as a culture to some of the (low mumble) um media images of people having been trafficked have (ahem) uh kind of set us up to to think of it only one way. So again, some of the media images might (low mumble) uh involve someone being held captive literally, you know, in bondage or kept away, you know, in a dark room or shackled. And that's not always the case. The unfortunate impact of that media exposure is if we don't see those same scenarios being depicted, sometimes we don't have the same type of empathic response as we should. We begin to say, well, if the person had an opportunity to, you know, get away or if they weren't really being held physically captive, then they somehow must have wanted to be there. And
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, doubt, contemplation; style: monologue, casual; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.5/10; 39.8s.
batch59_part3_batch59_part3_chunk_1526_1_1668642 · in -25.3 dBFS · gain +5.3 dB · snippets-01188
(affection, longing, contentment·very low-energy, moderately variable, some disfluency, casual)before we left, I did have a relationship with my sister. She would love the chance of babysitting my son. So, (low mumble) um, there will be times when I would just, you know, call my mom and say, hey, you know, can I borrow my sister, you know, for the weekend? And so there were times that he was safe home, and then there were times where (ahem) she couldn't babysit, and then I would, (ahem) um, take him to the night care provided for us.
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as affection, longing, contentment; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 6.3/10; 24.8s.
batch59_part3_batch59_part3_chunk_1526_1_1668737 · in -24.0 dBFS · gain +4.0 dB · snippets-01188
(sexual lust, pain, shame·normally alert, fairly steady, no disfluency, casual)was lured into sex trafficking when I was 18.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, pain, shame; style: casual, monologue; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 5.1/10; 3.4s.
batch59_part3_batch59_part3_chunk_1526_1_1668766 · in -26.8 dBFS · gain +6.8 dB · snippets-01188