rare-VN1-dims

VoiceNet: 2 chains from each of the 10 scarcest of the 57 dimensions (supply 42,740-63,981).

Rule. VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C
Source. trajectories_v5.parquet  |  Family. the scarcest cells -- the ones the owner said matter most
Sampled from 4,011,767 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
GEND — perceived genderrare-VN1-dims · #1

This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with perceived gender (GEND) below average — 0.42, lower than 58 % of clips in this corpus — and ends with it above average at 0.63, higher than 63 % of clips in this corpus. That is a total rise of 0.22.

It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · french · mls

hear it un-normalised (raw levels, max seam 0.2 dB)
k 2d_a 0.217d_b 0.217step_a 0.217step_b 0.217min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|judith_08_fillion_64total 32.0slevel spread 0.2 dBmax seam 0.2 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, measured, subdued, slightly relaxed
(longing, pride, malevolence malice · frequent disfluency, whispered, monologue) éprouvez aussi si ce que j'ai résolu de faire vient de lui et priez-le afin qu'il affermisse mon dessein vous vous tiendrez cette nvit à la porte de la ville et je sortirai avec ma servante
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, pride, malevolence malice; style: whispered, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.4/10; 16.8s, FRENCH.
10065_10612_000209 · in -27.2 dBFS · gain +7.2 dB · mls-00045
(sadness, helplessness · no disfluency, whispered, monologue) qu'on ne fasse autre chose que de prier pour moi le seigneur notre dieu eia prince de juda lui dit allez en paix et que le seigneur soit avec vous pour se venger de nos ennemis
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, helplessness; style: whispered, monologue; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.1s, FRENCH.
10065_10612_000002 · in -27.0 dBFS · gain +7.0 dB · mls-00045
GEND — perceived genderrare-VN1-dims · #2

This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with perceived gender (GEND) above average — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the range at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · portuguese · mls

hear it un-normalised (raw levels, max seam 0.5 dB)
k 2d_a 0.227d_b 0.227step_a 0.227step_b 0.227min_cos_consec min_cos_anchor dataset mlslang portuguesespeaker 10107track 10107|demonios_03_azevedo_total 24.6slevel spread 0.5 dBmax seam 0.5 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · slightly dark, very low-energy, relaxed, fairly narrow pitch, normal breath
(infatuation, sadness, thankfulness gratitude · slow, steady, frequent disfluency, whispered) em que elas como um bando de náiades alegres vinham aos saltos tontas de alegria quebrar na praia as suas risadas de prata
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, sadness, thankfulness gratitude; style: whispered, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.8/10; 12.1s, PORTUGUESE.
10107_12266_000261 · in -26.6 dBFS · gain +6.6 dB · mls-00114
(disgust, thankfulness gratitude, sadness · measured, fairly steady, almost no disfluency, narration) pobre mar pobre atleta nada mais lhe restava agora sobre o plúmbeo dorso fosforescente do que tristes esqueletos dos últimos navios
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, slightly dark, rough, very full; somewhat unclear, almost no disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as disgust, thankfulness gratitude, sadness; style: narration, storytelling; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 3.4/10; 12.3s, PORTUGUESE.
10107_12266_000274 · in -27.1 dBFS · gain +7.1 dB · mls-00114
REGS — register (speaking pitch region)rare-VN1-dims · #3

This is a VoiceNet dimension, not an emotion: register (speaking pitch region) (REGS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with register (speaking pitch region) (REGS) above average — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the range at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · german · mls

hear it un-normalised (raw levels, max seam 2.3 dB)
k 2d_a 0.227d_b 0.227step_a 0.227step_b 0.227min_cos_consec min_cos_anchor dataset mlslang germanspeaker 10148track 10148|annakarenina_055_toltotal 30.4slevel spread 2.3 dBmax seam 2.3 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · neutral-toned, neutral-bright, balanced body, good recording, normally alert, slightly relaxed, fairly steady, clear
(teasing, fear, impatience and irritability · measured, almost no disfluency, narration, storytelling) ja aber ein mann kann nicht nähren wendete peszow ein hingegen kann eine frau nun ein engländer hat doch einmal auf dem schiffe seinen säugling selbst genährt
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as teasing, fear, impatience and irritability; style: narration, storytelling; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.1/10; 12.7s, GERMAN.
10148_10349_006084 · in -24.5 dBFS · gain +4.5 dB · mls-00001
(fear, malevolence malice, sourness · normal-paced, some disfluency, narration, monologue) erwiderte der alte fürst der sich durch die anwesenheit seiner töchter nicht davon abhalten ließ sich im gespräche ein bißchen gehenzulassen es wird wohl so herauskommen daß es ebensoviel weibliche beamte gibt wie solche engländer sagte sergei
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, malevolence malice, sourness; style: narration, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.3/10; 17.6s, GERMAN.
10148_10349_007144 · in -26.8 dBFS · gain +6.8 dB · mls-00002
REGS — register (speaking pitch region)rare-VN1-dims · #4

This is a VoiceNet dimension, not an emotion: register (speaking pitch region) (REGS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with register (speaking pitch region) (REGS) high — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the range at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · german · mls

hear it un-normalised (raw levels, max seam 0.6 dB)
k 2d_a 0.234d_b 0.234step_a 0.234step_b 0.234min_cos_consec min_cos_anchor dataset mlslang germanspeaker 10148track 10148|annakarenina_111_toltotal 30.2slevel spread 0.6 dBmax seam 0.6 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, subdued, slightly relaxed, clear
(emotional numbness, infatuation, fear · measured, steady, almost no disfluency, didactic) wir in der stadt bekommen jetzt nichts anderes zu hören als vom serbischen kriege nun wie stellt sich denn mein freund dazu wahrscheinlich nicht so wie andere leute
full caption & clip details
An adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, infatuation, fear; style: didactic, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 11.9s, GERMAN.
10148_10349_010637 · in -27.6 dBFS · gain +7.6 dB · mls-00003
(bitterness, jealousy and envy, infatuation · normal-paced, fairly steady, no disfluency, narration) o doch ziemlich ebenso wie alle antwortete kitty etwas verlegen und blickte sich dabei nach sergei iwanowitsch um ich will hinschicken und ihn rufen lassen bei uns ist auch papa zu besuch er ist erst vor kurzem aus dem ausland zurückgekommen
full caption & clip details
A middle-aged feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as bitterness, jealousy and envy, infatuation; style: narration, storytelling; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 18.2s, GERMAN.
10148_10349_010792 · in -27.0 dBFS · gain +7.0 dB · mls-00003
S_FORM — style: formalrare-VN1-dims · #5

This is a VoiceNet dimension, not an emotion: style: formal (S_FORM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: formal (S_FORM) around average — 0.43, lower than 57 % of clips in this corpus — and ends with it above average at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 40 s · portuguese · mls

hear it un-normalised (raw levels, max seam 0.9 dB)
k 2d_a 0.227d_b 0.227step_a 0.227step_b 0.227min_cos_consec min_cos_anchor dataset mlslang portuguesespeaker 10107track 10107|antologiabrasileira1total 40.1slevel spread 0.9 dBmax seam 0.9 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady
(anger, awe, interest · very low-energy, some disfluency, somewhat unclear, storytelling) o desenvolvimento de forças locaes demasiado dominadoras em vez de grandes barões eu pudera dizer que o ambiente só produziu baronetes todos os elementos conservadores reuniram-se entretanto na familia do visconde do uruguay com
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, awe, interest; style: storytelling, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 3.7/10; 20.0s, PORTUGUESE.
10107_6390_000422 · in -24.2 dBFS · gain +4.2 dB · mls-00121
(jealousy and envy · normally alert, frequent disfluency, slurred, monologue) no centro dirigida nos nossos dias pelo conselheiro paulino josé soares de souza a um tempo politico e lavrador os elementos liberaes muito dispersos estiveram por annos ao dispor de francisco octaviano poeta eterno
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy; style: monologue, didactic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.5/10; 20.0s, PORTUGUESE.
10107_6390_000434 · in -23.3 dBFS · gain +3.3 dB · mls-00121
S_FORM — style: formalrare-VN1-dims · #6

This is a VoiceNet dimension, not an emotion: style: formal (S_FORM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: formal (S_FORM) around average — 0.52, higher than 52 % of clips in this corpus — and ends with it above average at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.22.

It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · portuguese · mls

k 2d_a 0.225d_b 0.225step_a 0.225step_b 0.225min_cos_consec min_cos_anchor dataset mlslang portuguesespeaker 10107track 10107|iracema_02_dealencartotal 30.4slevel spread 0.3 dBmax seam 0.3 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · warm, slightly dark, slightly rough, balanced body, average recording, measured, subdued, slightly relaxed
(shame, affection, infatuation · fairly steady, frequent disfluency, narration, monologue) não porque irapuã vai ser punido pela mão de iracema seu primeiro passo é o passo da morte a virgem retraiu dum salto o avanço que tomara e vibrou o arco
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as shame, affection, infatuation; style: narration, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.3/10; 13.7s, PORTUGUESE.
10107_11436_000776 · in -25.6 dBFS · gain +5.6 dB · mls-00114
(shame, malevolence malice, contentment · steady, no disfluency, whispered, narration) o chefe cerrou ainda o punho do formidável tacape mas pela vez primeira sentiu que pesava ao braço robusto o golpe que devia ferir iracema ainda não alçado já lhe trespassava a ele próprio o coração
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is warm, slightly dark, slightly rough, balanced body; slurred, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, malevolence malice, contentment; style: whispered, narration; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.1/10; 16.6s, PORTUGUESE.
10107_11436_000855 · in -25.9 dBFS · gain +5.9 dB · mls-00114
S_CASU — style: casualrare-VN1-dims · #7

This is a VoiceNet dimension, not an emotion: style: casual (S_CASU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: casual (S_CASU) below average — 0.34, lower than 66 % of clips in this corpus — and works its way down to low at 0.09, lower than 91 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · german · mls

k 2d_a -0.246d_b -0.246step_a 0.246step_b 0.246min_cos_consec min_cos_anchor dataset mlslang germanspeaker 10148track 10148|annakarenina_036_toltotal 26.4slevel spread 2.5 dBmax seam 2.5 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(disgust, fear, bitterness · normal-paced, moderately variable, almost no disfluency, narration) ohne hilfe dahinstirbt rohe unwissende bauernweiber die als geburtshelferinnen dienen martern die kinder zu tode und das volk verharrt in tiefster unwissenheit und ist der willkür jedes schreibers preisgegeben
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as disgust, fear, bitterness; style: narration, storytelling; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.4/10; 15.8s, GERMAN.
10148_10349_008268 · in -23.8 dBFS · gain +3.8 dB · mls-00002
(disappointment, sadness, helplessness · measured, fairly steady, no disfluency, narration) dir aber ist das mittel dem abzuhelfen in die hand gelegt und doch schaffst du keine abhilfe weil das alles deiner meinung nach nicht wichtig ist
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, sadness, helplessness; style: narration, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.8/10; 10.5s, GERMAN.
10148_10349_008028 · in -26.3 dBFS · gain +6.3 dB · mls-00002
S_CASU — style: casualrare-VN1-dims · #8

This is a VoiceNet dimension, not an emotion: style: casual (S_CASU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: casual (S_CASU) below average — 0.41, lower than 59 % of clips in this corpus — and works its way down to low at 0.16, lower than 84 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · german · mls

k 2d_a -0.249d_b -0.249step_a 0.249step_b 0.249min_cos_consec min_cos_anchor dataset mlslang germanspeaker 10148track 10148|annakarenina_063_toltotal 30.1slevel spread 2.5 dBmax seam 2.5 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · neutral-toned, neutral-bright, balanced body, good recording, slightly relaxed, fairly steady, clear, light breath
(affection, contentment, anger · normal-paced, normally alert, little disfluency, narration) nun dann wollen wir es so machen bringe du ihn in unserm wagen hin und sergei iwanowitsch ist wohl so gut zum hochzeitsmarschall zu fahren und uns nachher den wagen zu schicken schön sehr gern ich fahre sofort mit ihm hin ist dein gepäck schon weggeschafft fragte stepan
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, contentment, anger; style: narration, storytelling; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.8/10; 19.0s, GERMAN.
10148_10349_010796 · in -27.6 dBFS · gain +7.6 dB · mls-00003
(malevolence malice · measured, very low-energy, frequent disfluency, storytelling) jawohl jawohl antwortete ljewin und befahl kusma ihm beim ankleiden behilflich zu sein
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as malevolence malice; style: storytelling, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.0/10; 10.9s, GERMAN.
10148_10349_010905 · in -30.1 dBFS · gain +10.1 dB · mls-00003
S_NARR — style: narrationrare-VN1-dims · #9

This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: narration (S_NARR) above average — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the range at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.24.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 28 s · french · mls

k 2d_a 0.238d_b 0.238step_a 0.238step_b 0.238min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_01_letotal 28.3slevel spread 1.1 dBmax seam 1.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a child masculine voice · slightly dark, balanced body, average recording, slow, slightly relaxed, steady, frequent disfluency, fairly narrow pitch
(awe, emotional numbness, infatuation · very low-energy, somewhat unclear, minimal breath, whispered) et aussitôt qu'il fut sorti de l'eau il vit les cieux s'ouvrir et l'esprit comme une colombe descendre et demeurer sur lui
full caption & clip details
A child masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as awe, emotional numbness, infatuation; style: whispered, monologue; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 11.2s, FRENCH.
10065_10039_000222 · in -27.3 dBFS · gain +7.3 dB · mls-00043
(contentment, awe, infatuation · subdued, slurred, light breath, whispered) et une voix se fit entendre du ciel vous êtes mon fils bien aiméj c'est en vous evangile chap que j'ai mis toute mon affection et aussitôt après l'esprit le poussa dans le désert où il demeura quarante jours et quarante nuits
full caption & clip details
A middle-aged somewhat feminine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, awe, infatuation; style: whispered, narration; average recording, quiet background; genuineness 0.4/6; vocal-burst blend 0.6/10; 16.9s, FRENCH.
10065_10039_000411 · in -26.3 dBFS · gain +6.3 dB · mls-00043
S_NARR — style: narrationrare-VN1-dims · #10

This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: narration (S_NARR) at the very top of the range — 1.00, virtually no clip in this corpus scores higher — and works its way down to high at 0.76, higher than 76 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 22 s · portuguese · mls

k 2d_a -0.237d_b -0.237step_a 0.237step_b 0.237min_cos_consec min_cos_anchor dataset mlslang portuguesespeaker 10107track 10107|antologiabrasileira1total 22.4slevel spread 1.4 dBmax seam 1.4 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · normally alert, fairly steady, somewhat unclear
(fear, disgust, sourness · slow, slightly relaxed, almost no disfluency, storytelling) ruja a tormenta embora cerre se a noite sobre este triste mundo que parece querer voltar para o paganismo
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, almost no disfluency, very wide pitch range, audible breath; affect is positive, neutral stance, slightly guarded; reads as fear, disgust, sourness; style: storytelling, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.9/10; 10.0s, PORTUGUESE.
10107_6390_000520 · in -26.6 dBFS · gain +6.6 dB · mls-00121
(awe, contentment, thankfulness gratitude · measured, neutral tension, frequent disfluency, storytelling) os pharóes estão accesos a costa toda illumi nada a doutrina catbolica se afisrma em toda a sua força em toda a sua belleza
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is warm, neutral-bright, slightly rough, booming; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as awe, contentment, thankfulness gratitude; style: storytelling, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 3.1/10; 12.2s, PORTUGUESE.
10107_6390_000574 · in -25.3 dBFS · gain +5.3 dB · mls-00121
S_AUTH — style: authoritativerare-VN1-dims · #11

This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: authoritative (S_AUTH) below average — 0.35, lower than 65 % of clips in this corpus — and ends with it above average at 0.58, higher than 58 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · portuguese · mls

k 2d_a 0.230d_b 0.230step_a 0.230step_b 0.230min_cos_consec min_cos_anchor dataset mlslang portuguesespeaker 10107track 10107|antologiabrasileira1total 25.5slevel spread 0.3 dBmax seam 0.3 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · warm, rough, average recording, no background noise, steady, fairly narrow pitch
(awe, sadness, affection · slow, very low-energy, relaxed, narration) era nosso paiz na pedra isolada do valle na arvore gigante da montanha no píncaro agreste da serrania na terra no céu e nas aguas por toda a parte
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is warm, dark, rough, very full; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, fairly guarded; reads as awe, sadness, affection; style: narration, whispered; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.8/10; 15.1s, PORTUGUESE.
10107_6390_000634 · in -22.5 dBFS · gain +2.5 dB · mls-00121
(awe, contentment, affection · measured, normally alert, slightly relaxed, narration) deus estampou o verbo eterno da liberdade creadora na face da natureza antes de graval a fia consciência do homem
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is warm, slightly dark, rough, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as awe, contentment, affection; style: narration, storytelling; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.1/10; 10.2s, PORTUGUESE.
10107_6390_000729 · in -22.2 dBFS · gain +2.2 dB · mls-00121
S_AUTH — style: authoritativerare-VN1-dims · #12

This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: authoritative (S_AUTH) above average — 0.59, higher than 59 % of clips in this corpus — and works its way down to below average at 0.38, lower than 62 % of clips in this corpus. That is a total fall of 0.21.

It takes 2 clips to get there. Clip to clip the moves are -0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 40 s · portuguese · mls

k 2d_a -0.207d_b -0.207step_a 0.207step_b 0.207min_cos_consec min_cos_anchor dataset mlslang portuguesespeaker 10107track 10107|demonios_04_azevedo_total 40.1slevel spread 1.1 dBmax seam 1.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly warm, slightly dark, rough, good recording, no background noise, measured, subdued, slightly relaxed
(bitterness, disappointment, sadness · clear, narration, monologue) nos meus seus braços esgalhados e procurando unir sua boca à minha boca e assim nos quedamos para sempre aí plantados e seguros sem nunca mais nos soltarmos um do outro nem mais podermos mover com os nossos duros membros
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, rough, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as bitterness, disappointment, sadness; style: narration, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 5.5/10; 20.0s, PORTUGUESE.
10107_12266_000125 · in -27.3 dBFS · gain +7.3 dB · mls-00114
(bitterness, helplessness, pain · somewhat unclear, narration, monologue) e pouco a pouco nossos cabelos e nossos pêlos se nos foram desprendendo e caindo lentamente pelo corpo abaixo e cada poro que eles deixavam era um novo respiradouro que se abria para beber a noite tenebrosa
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, rough, very full; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as bitterness, helplessness, pain; style: narration, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 5.4/10; 20.0s, PORTUGUESE.
10107_12266_000295 · in -28.4 dBFS · gain +8.4 dB · mls-00114
S_NEWS — style: newsreadingrare-VN1-dims · #13

This is a VoiceNet dimension, not an emotion: style: newsreading (S_NEWS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: newsreading (S_NEWS) at the very top of the range — 0.91, higher than 91 % of clips in this corpus — and works its way down to above average at 0.66, higher than 66 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · french · mls

k 2d_a -0.245d_b -0.245step_a 0.245step_b 0.245min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_16_letotal 23.5slevel spread 1.6 dBmax seam 1.6 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, subdued, slightly relaxed, steady, somewhat unclear, fairly narrow pitch
(pain, sexual lust, sadness · slow, no disfluency, light breath, whispered) ils prendront les serpents avec la main et s'ils boivent quelque breuvage mortel il ne leur fera point de mal ils imposeront les mains sur les malades et les malades seront guéris
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, sexual lust, sadness; style: whispered, monologue; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 13.3s, FRENCH.
10065_10039_000226 · in -24.4 dBFS · gain +4.4 dB · mls-00043
(awe, pain, contemplation · measured, frequent disfluency, normal breath, monologue) le seigneur jésus après leur avoir ainsi parlé fut élevé dans le ciel où il est assis à la droite de dieu
full caption & clip details
A child masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pain, contemplation; style: monologue; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.0s, FRENCH.
10065_10039_000121 · in -26.1 dBFS · gain +6.0 dB · mls-00043
S_NEWS — style: newsreadingrare-VN1-dims · #14

This is a VoiceNet dimension, not an emotion: style: newsreading (S_NEWS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: newsreading (S_NEWS) above average — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the range at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · german · mls

k 2d_a 0.250d_b 0.250step_a 0.250step_b 0.250min_cos_consec min_cos_anchor dataset mlslang germanspeaker 10148track 10148|annakarenina_011_toltotal 28.6slevel spread 0.1 dBmax seam 0.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(fatigue exhaustion, contempt, disappointment · measured, some disfluency, whispered, monologue) bei bobrischtschews geht es immer lustig zu auch bei nikitins aber bei meschkows ist es immer langweilig ist ihnen das nicht auch aufgefallen
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, contempt, disappointment; style: whispered, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.9/10; 10.1s, GERMAN.
10148_10349_009430 · in -25.6 dBFS · gain +5.6 dB · mls-00002
(affection, emotional numbness, longing · normal-paced, no disfluency, narration, monologue) nein mein herz für mich gibt es solche bälle nicht mehr auf denen es lustig ist erwiderte anna und kitty erblickte in ihren augen jene besondere welt die ihr verschlossen blieb für mich gibt es nur solche auf denen es weniger öde und langweilig ist als auf anderen
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, emotional numbness, longing; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 18.3s, GERMAN.
10148_10349_009929 · in -25.7 dBFS · gain +5.7 dB · mls-00002
DFLU — disfluencyrare-VN1-dims · #15

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with disfluency (DFLU) below average — 0.26, lower than 74 % of clips in this corpus — and ends with it around average at 0.51, right about the corpus median. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · french · mls

k 2d_a 0.248d_b 0.248step_a 0.248step_b 0.248min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_09_letotal 30.3slevel spread 1.4 dBmax seam 1.4 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a child masculine voice · slightly dark, slightly rough, balanced body, quiet background, slow, slightly relaxed, steady, frequent disfluency
(awe, malevolence malice, longing · subdued, somewhat unclear, light breath, monologue) ses vêtements devinrent tout brillants de lumière et blancs comme la neige en sorte qu'il n'y a point de foulon sur la terre qui puisse en faire d'aussi blancs
full caption & clip details
A child masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, malevolence malice, longing; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.7s, FRENCH.
10065_10039_000000 · in -27.8 dBFS · gain +7.8 dB · mls-00043
(malevolence malice, emotional numbness, affection · very low-energy, slurred, audible breath, whispered) et ils virent paraître elie et moïse qui s'entretenaient avec jésus alors pierre dit à jésus maître nous sommes bien ici j faisons y trois tentes une pour vous une pour moïse et une pour
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, emotional numbness, affection; style: whispered, ASMR; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.5/10; 18.5s, FRENCH.
10065_10039_000442 · in -26.4 dBFS · gain +6.4 dB · mls-00043
DFLU — disfluencyrare-VN1-dims · #16

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with disfluency (DFLU) low — 0.12, lower than 88 % of clips in this corpus — and ends with it below average at 0.37, lower than 63 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · french · mls

k 2d_a 0.246d_b 0.246step_a 0.246step_b 0.246min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_10_letotal 34.5slevel spread 2.3 dBmax seam 2.3 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly dark, balanced body, average recording, quiet background, frequent disfluency, narrow pitch range
(awe, relief, malevolence malice · measured, normally alert, slightly relaxed, monologue) un aveugle nommé bartimée fils de timée qui était assis sur le chemin pour demander l'aumône ayant appris que c'était jésus de nazareth se mit à chap lo ii selon s marc crier jésus fils de david ayez pitié de moi
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, relief, malevolence malice; style: monologue, whispered; average recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.8/10; 15.8s, FRENCH.
10065_10039_000230 · in -24.1 dBFS · gain +4.1 dB · mls-00043
(malevolence malice, anger, jealousy and envy · slow, very low-energy, relaxed, storytelling) et plusieurs le reprenaient et lui disaient qu'il se tûtj mais il criait encore beaucoup plus haut fils de david ayez pitié de moi alors jésus s'étant arrêté commanda qu'on l'appelât
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is warm, slightly dark, rough, balanced body; slurred, frequent disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, fairly guarded; reads as malevolence malice, anger, jealousy and envy; style: storytelling, narration; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 3.5/10; 18.6s, FRENCH.
10065_10039_000046 · in -26.4 dBFS · gain +6.5 dB · mls-00043
DARC — darkness of timbrerare-VN1-dims · #17

This is a VoiceNet dimension, not an emotion: darkness of timbre (DARC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with darkness of timbre (DARC) low — 0.15, lower than 85 % of clips in this corpus — and ends with it below average at 0.38, lower than 62 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · french · mls

k 2d_a 0.233d_b 0.233step_a 0.233step_b 0.233min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_05_letotal 24.3slevel spread 0.4 dBmax seam 0.4 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, measured, slightly relaxed, steady, fairly narrow pitch
(awe, emotional numbness, thankfulness gratitude · subdued, no disfluency, somewhat unclear, monologue) cet homme s'en étant allé commença à publier dans la décapole les grandes grâces que jésus lui avait faites et tout le monde en était dans l'admiration
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, emotional numbness, thankfulness gratitude; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.3s, FRENCH.
10065_10039_000362 · in -26.8 dBFS · gain +6.8 dB · mls-00043
(emotional numbness, awe, contemplation · normally alert, little disfluency, clear, monologue) étant encore reeassé dans la barque à l'autre ord lorsqu'il était auprès de la mer une grande multitude de peuple s'assembla autour de lui
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, awe, contemplation; style: monologue, whispered; average recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 11.9s, FRENCH.
10065_10039_000234 · in -26.4 dBFS · gain +6.3 dB · mls-00043
DARC — darkness of timbrerare-VN1-dims · #18

This is a VoiceNet dimension, not an emotion: darkness of timbre (DARC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with darkness of timbre (DARC) below average — 0.39, lower than 61 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 35 s · french · mls

k 2d_a -0.242d_b -0.242step_a 0.242step_b 0.242min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_06_letotal 34.8slevel spread 1.3 dBmax seam 1.3 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · balanced body, average recording, quiet background, measured, subdued, slightly relaxed, steady, frequent disfluency
(emotional numbness, infatuation, longing · somewhat unclear, light breath, whispered, monologue) car la fille d'hérotliade y étant entre'e et ayant dansé devant hérode elle lui plut tellement et à ceux qui étaien avec lui qu il lui dit demandez-moi ce que vous voudrez et je vous le donnerai
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, infatuation, longing; style: whispered, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.4s, FRENCH.
10065_10039_000309 · in -25.7 dBFS · gain +5.7 dB · mls-00043
(anger, bitterness, shame · slurred, audible breath, whispered, narration) et il ajouta avec serment oui je vous donnerai tout ce que vous me demanderez quand ce serait la moitié de mon royaume elle étant sortie dit à sa mère que demanderai je sa mère lui répondit la tête de
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is warm, dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as anger, bitterness, shame; style: whispered, narration; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.2/10; 19.2s, FRENCH.
10065_10039_000293 · in -27.0 dBFS · gain +7.0 dB · mls-00043
AROU — arousal / activationrare-VN1-dims · #19

This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with arousal / activation (AROU) at the very bottom of the range — 0.03, lower than 97 % of clips in this corpus — and ends with it below average at 0.28, lower than 72 % of clips in this corpus. That is a total rise of 0.24.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · french · mls

k 2d_a 0.243d_b 0.243step_a 0.243step_b 0.243min_cos_consec min_cos_anchor dataset mlslang frenchspeaker 10065track 10065|evangilestmarc_14_letotal 31.7slevel spread 0.1 dBmax seam 0.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly dark, balanced body, average recording, steady, frequent disfluency, fairly narrow pitch, light breath
(anger, fear, malevolence malice · slow, very low-energy, relaxed, whispered) mais pierre insistait encore davantage quand il me faudrait mourir avec vous je ne vous renoncerai point et tous les autres en dirent autant
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, fairly guarded; reads as anger, fear, malevolence malice; style: whispered, ASMR; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.3/10; 12.0s, FRENCH.
10065_10039_000032 · in -26.8 dBFS · gain +6.8 dB · mls-00043
(malevolence malice, jealousy and envy, disappointment · measured, subdued, slightly relaxed, whispered) ils allèrent ensuite au lieu appelé gethsémani où il dit à ses disciples asseyezvous ici jusqu'à ce que j'aie fait ma prière et ayant pris avec lui pierre jacques et jean il commença à être saisi de frayeur et pénétré d'une extrême
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, jealousy and envy, disappointment; style: whispered, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.3/10; 19.5s, FRENCH.
10065_10039_000283 · in -26.7 dBFS · gain +6.7 dB · mls-00043
AROU — arousal / activationrare-VN1-dims · #20

This is a VoiceNet dimension, not an emotion: arousal / activation (AROU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with arousal / activation (AROU) low — 0.15, lower than 85 % of clips in this corpus — and ends with it below average at 0.38, lower than 62 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · german · mls

k 2d_a 0.230d_b 0.230step_a 0.230step_b 0.230min_cos_consec min_cos_anchor dataset mlslang germanspeaker 10148track 10148|annakarenina_005_toltotal 28.9slevel spread 0.8 dBmax seam 0.8 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, slightly relaxed, fairly steady
(jealousy and envy, bitterness, fear · measured, very low-energy, little disfluency, whispered) langweilen sie sich denn im winter auf dem lande gar nicht fragte sie nein ich langweile mich nicht ich habe viel zu tun erwiderte er und fühlte dabei daß sie ihn durch ihren ruhigen ton in schranken hielt
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy, bitterness, fear; style: whispered, narration; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.3/10; 14.8s, GERMAN.
10148_10349_002015 · in -26.0 dBFS · gain +6.0 dB · mls-00000
(sadness, helplessness, confusion · normal-paced, normally alert, no disfluency, narration) und daß er jetzt ebensowenig imstande sein werde diese schranken zu durchbrechen wie er es zu anfang des winters gekonnt hatte sind sie zu längerem aufenthalte nach moskau gekommen fragte ihn kitty
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sadness, helplessness, confusion; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 13.9s, GERMAN.
10148_10349_001404 · in -26.8 dBFS · gain +6.8 dB · mls-00000