k-VN1-k2

VN1 at chain length k=2, all corpora, at the mining floor.

Rule. VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C
Source. trajectories_v5.parquet  |  Family. rule x chain length
Sampled from 672,000 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
S_TECH — style: technicalk-VN1-k2 · #1

This is a VoiceNet dimension, not an emotion: style: technical (S_TECH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: technical (S_TECH) below average — 0.39, lower than 61 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 39 s · en · emolia

hear it un-normalised (raw levels, max seam 0.9 dB)
k 2d_a -0.244d_b -0.244step_a 0.244step_b 0.244min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_nm3YFc8IBy4track EN_nm3YFc8IBy4total 38.8slevel spread 0.9 dBmax seam 0.9 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(astonishment surprise · no disfluency, light breath, newsreading, formal) Background checks on the 1st Ordnance Squadron revealed that it had several escaped convicts in its ranks. Uwana surmised that enlisting in the army under false names was an easy way of escaping detection during wartime.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.2s, EN.
EN_nm3YFc8IBy4_W000025 · in -14.0 dBFS · gain -6.0 dB · emolia-02343
(pride, triumph · almost no disfluency, minimal breath, newsreading, formal) Since skilled technicians were hard to find, Tibbets elected to keep them, threatening to send them back to prison for any dereliction of duty or security breaches.Uana oversaw the movement of the 509th from its training base in Wendover Army Air Field, Utah to Tinian Island in the Western Pacific, traveling by air with the Project Alberta Advance Party of 34 in a Douglas C-54 Skymaster.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as pride, triumph; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 26.4s, EN.
EN_nm3YFc8IBy4_W000026 · in -14.8 dBFS · gain -5.2 dB · emolia-02343
ROUG — roughnessk-VN1-k2 · #2

This is a VoiceNet dimension, not an emotion: roughness (ROUG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with roughness (ROUG) above average — 0.74, higher than 74 % of clips in this corpus — and works its way down to around average at 0.50, right about the corpus median. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.64 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.64 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.64, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 47 s · ru · podcast

hear it un-normalised (raw levels, max seam 3.7 dB)
k 2d_a -0.242d_b -0.242step_a 0.242step_b 0.242min_cos_consec 0.6366min_cos_anchor 0.6366dataset podcastlang ruspeaker 320045track 320045total 47.2slevel spread 3.7 dBmax seam 3.7 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(pride, disappointment, elation · monologue, casual) Может быть, ему что-то не нравится в этой карте, но вот это вот мое единственное пожелание, потому что это вот то, что поможет сделать На'ви менее прогнозируемой командой, что ли. И поможет нивелировать те же проблемы с Ньюком, потому что появится возможность, например, в том же матчапе с Маус, в теории банить карту, на которой они чувствуют, что они слабее. Ну, в общем, я вот за это.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, disappointment, elation; style: monologue, casual; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 6.2/10; 23.1s, RU.
320045_00054712 · in -26.7 dBFS · gain +6.7 dB · podcast-03722
(interest, contentment, elation · monologue, casual) Слушай, вот ты же вообще чуть ли не самый главный фанат (ahem) Vertiga. Не только в русском комьюнити, на самом деле, во всем (ahem) комьюнити. Во всем комьюнити. То есть я видел очень много роликов, где игроков спрашивают про самую худшую карту. Там 90% называют на самом деле Vertigo. Почему тебе так нравится эта карта? Почему ты думаешь, что другие игроки так негативно отзываются, а ты прям полностью утопишь за вертигов уже очень долгое время?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, contentment, elation; style: monologue, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 10.0/10; 23.9s, RU.
320045_00057064 · in -30.4 dBFS · gain +10.4 dB · podcast-04817
RANG — pitch range usedk-VN1-k2 · #3

This is a VoiceNet dimension, not an emotion: pitch range used (RANG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with pitch range used (RANG) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it low at 0.25, lower than 75 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · en · emolia

hear it un-normalised (raw levels, max seam 0.5 dB)
k 2d_a 0.246d_b 0.246step_a 0.246step_b 0.246min_cos_consec min_cos_anchor dataset emolialang enspeaker EN__5Y0LTHdj64track EN__5Y0LTHdj64total 19.5slevel spread 0.5 dBmax seam 0.5 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · slightly rough, slightly thin, quiet background, slow, relaxed, frequent disfluency, normal breath
(doubt · lethargic, steady, slurred, whispered) (low mumble) Uhm, Amy, can we change your name for your channel? Amy, can we change the name for your channel? (low mumble) Uhm... This is a good question.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, neutral openness; reads as doubt; style: whispered, monologue; below-average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.1/10; 13.8s, EN.
EN__5Y0LTHdj64_W000189 · in -20.9 dBFS · gain +0.9 dB · emolia-02481
(very low-energy, fairly steady, somewhat unclear, casual) Alright, so under advanced settings, it looks like I can change this here.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.2/10; 5.6s, EN.
EN__5Y0LTHdj64_W000190 · in -20.4 dBFS · gain +0.5 dB · emolia-02481
WARM — warmthk-VN1-k2 · #4

This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with warmth (WARM) above average — 0.60, higher than 60 % of clips in this corpus — and ends with it high at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 46 s · en · podcast

hear it un-normalised (raw levels, max seam 3.1 dB)
k 2d_a 0.246d_b 0.246step_a 0.246step_b 0.246min_cos_consec 0.9273min_cos_anchor 0.9273dataset podcastlang enspeaker 618062track 618062total 46.0slevel spread 3.1 dBmax seam 3.1 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, neutral tension, moderately variable, some disfluency, average clarity
(hope enthusiasm optimism, triumph, pride · brisk, energised, casual, conversational) There are people like this, folks. (chuckle) There really are people like this. What do we have to do? Well, we need to study together. Well, how soon can we do this? Well, camp meeting's coming up. Well, I want to get started. So we figured out a way to do that just before camp meeting, and we had our first Bible study with her. She came to camp meeting the first Sabbath, and (low mumble) uh and she got up at four o'clock in the morning.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, triumph, pride; style: casual, conversational; average recording, some background noise; genuineness 3.3/6; vocal-burst blend 6.4/10; 25.9s, EN.
618062_00274656 · in -20.7 dBFS · gain +0.7 dB · podcast-02269
(awe, fear, triumph · normal-paced, normally alert, casual, monologue) Two days ago to get her son up here. And this is some of one of the ways that the imp (low mumble) the uh this was impacted. Now, this could have been a deacon, but it turned out to be a pastor because I needed somebody who could do it quickly. And (low mumble) um, and I was hoping that if he couldn't do it, he had a deacon or deaconess who would be able to do this work.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as awe, fear, triumph; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.6/10; 19.9s, EN.
618062_00277244 · in -23.8 dBFS · gain +3.8 dB · podcast-02226
TENS — tensionk-VN1-k2 · #5

This is a VoiceNet dimension, not an emotion: tension (TENS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with tension (TENS) high — 0.82, higher than 82 % of clips in this corpus — and works its way down to above average at 0.58, higher than 58 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · en · emolia

hear it un-normalised (raw levels, max seam 2.1 dB)
k 2d_a -0.242d_b -0.242step_a 0.242step_b 0.242min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_ZjPMdGH03VItrack EN_ZjPMdGH03VItotal 11.7slevel spread 2.1 dBmax seam 2.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, moderate pitch range
(awe, pleasure ecstasy, disgust · normal-paced, steady, almost no disfluency, formal) Sniffing and touching the cookies. Each one with a nice illustration of Cookie Monster.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, pleasure ecstasy, disgust; style: formal, monologue; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.8/10; 5.2s, EN.
EN_ZjPMdGH03VI_W000049 · in -21.4 dBFS · gain +1.4 dB · emolia-01707
(emotional numbness · measured, fairly steady, some disfluency, monologue) For the listening, he has that classic, hands up to the ear, listening to the cookie. And then he has one where he is...
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, casual; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.4/10; 6.3s, EN.
EN_ZjPMdGH03VI_W000050 · in -19.2 dBFS · gain -0.8 dB · emolia-01707
R_MIXD — resonance: mixedk-VN1-k2 · #6

This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: mixed (R_MIXD) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to above average at 0.74, higher than 74 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · en · emolia

k 2d_a -0.238d_b -0.238step_a 0.238step_b 0.238min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_tBWtn6M2UAYtrack EN_tBWtn6M2UAYtotal 24.3slevel spread 0.6 dBmax seam 0.6 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, steady, clear, moderate pitch range
(malevolence malice, concentration, interest · normal-paced, little disfluency, light breath, formal) There's a crack in everything and that's how the light gets in. That's a Lenin quote. When faced with this shared challenge that confronts the global population.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, concentration, interest; style: formal, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.3/10; 8.5s, EN.
EN_tBWtn6M2UAY_W000046 · in -18.9 dBFS · gain -1.1 dB · emolia-01274
(interest, concentration, sexual lust · measured, some disfluency, minimal breath, whispered) I think an opportunity emerged to form shared goals, bringing all hands on deck to crowdsource for idea that unlocks our collective and connective intelligence to march toward a free, open, democratic and resilient Indo-Pacific.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, some disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, concentration, sexual lust; style: whispered, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 2.0/10; 15.7s, EN.
EN_tBWtn6M2UAY_W000047 · in -18.3 dBFS · gain -1.7 dB · emolia-01274
CHNK — chunking / phrasing densityk-VN1-k2 · #7

This is a VoiceNet dimension, not an emotion: chunking / phrasing density (CHNK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with chunking / phrasing density (CHNK) high — 0.78, higher than 78 % of clips in this corpus — and works its way down to around average at 0.55, higher than 55 % of clips in this corpus. That is a total fall of 0.23.

It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 16 s · zh · emolia

k 2d_a -0.232d_b -0.232step_a 0.232step_b 0.232min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00014_S03050track ZH_B00014_S03050total 15.5slevel spread 3.0 dBmax seam 3.0 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly rough, good recording, no background noise, almost no disfluency, clear
(malevolence malice, anger, distress · brisk, energised, neutral tension, narration) Jimmi's bar, i checked four other bars. Before i found you, we went to the scene of a homicide. The victim's name was carlos or teas, he blood in my memory.
full caption & clip details
A middle-aged masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, full; clear, almost no disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, neutral openness; reads as malevolence malice, anger, distress; style: narration, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.6/10; 10.5s, ZH.
ZH_B00014_S03050_W000004 · in -18.3 dBFS · gain -1.7 dB · emolia-03412
(sadness, distress, fear · slow, very low-energy, slightly relaxed, narration) His name is coal, and they just turned six at the time of the accident.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, very full; clear, almost no disfluency, fairly narrow pitch, light breath; affect is negative, neutral stance, neutral openness; reads as sadness, distress, fear; style: narration, whispered; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.5/10; 4.8s, ZH.
ZH_B00014_S03050_W000005 · in -21.3 dBFS · gain +1.3 dB · emolia-03412
CLRT — clarity / intelligibilityk-VN1-k2 · #8

This is a VoiceNet dimension, not an emotion: clarity / intelligibility (CLRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with clarity / intelligibility (CLRT) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to above average at 0.74, higher than 74 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.94 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.94 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 17 s · en · emolia

k 2d_a -0.235d_b -0.235step_a 0.235step_b 0.235min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00029_S03803track EN_B00029_S03803total 17.5slevel spread 0.8 dBmax seam 0.8 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a child masculine voice · neutral-toned, neutral-bright, good recording, no background noise, slightly relaxed, wide pitch range
(teasing, malevolence malice, fatigue exhaustion · normal-paced, energised, moderately variable, narration) You should be off to bed. After all, this may be the last peaceful night's sleep you'll have for several weeks, he said with a wink.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as teasing, malevolence malice, fatigue exhaustion; style: narration, playful; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 8.7s, EN.
EN_B00029_S03803_W000042 · in -18.6 dBFS · gain -1.4 dB · emolia-00807
(relief, thankfulness gratitude, contentment · measured, very low-energy, fairly steady, conversational) Okay, sir. Thank you for the advice. Maybe I'll see you on the ship, Patrick said as he turned to go home. Stormy seas.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, no disfluency, wide pitch range, minimal breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, thankfulness gratitude, contentment; style: conversational, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.8/10; 8.6s, EN.
EN_B00029_S03803_W000043 · in -19.4 dBFS · gain -0.6 dB · emolia-00807
S_CASU — style: casualk-VN1-k2 · #9

This is a VoiceNet dimension, not an emotion: style: casual (S_CASU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: casual (S_CASU) above average — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the range at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · en · podcast

k 2d_a 0.249d_b 0.249step_a 0.249step_b 0.249min_cos_consec 0.8667min_cos_anchor 0.8667dataset podcastlang enspeaker 111638track 111638total 34.5slevel spread 1.5 dBmax seam 1.5 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, slightly rough, below-average recording, neutral tension, fairly steady, somewhat unclear, moderate pitch range, normal breath
(jealousy and envy, helplessness, intoxication altered states of consciousness · measured, very low-energy, frequent disfluency, casual) Like at the end of the day, I knew that, okay, boom. We had the Bulls. We had the Wizards. We had (low mumble) the Hornets. They all got better. So my thing is who was going to drop? And I didn't think the Hawks were going to be one of the teams. And it's still early. But I didn't think the Hawks was going to be one of the teams that fell. I thought Knicks, you know, (low mumble) uh, I knew Sixers were going to get worse because, of course, you lose
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as jealousy and envy, helplessness, intoxication altered states of consciousness; style: casual, monologue; below-average recording, some background noise; genuineness 4.6/6; vocal-burst blend 8.6/10; 27.2s, EN.
111638_00332212 · in -17.4 dBFS · gain -2.6 dB · podcast-05797
(relief, triumph · brisk, normally alert, some disfluency, casual) Jaron Jackson been hooping, stepped this game up. Dylan Brooks came back off his injury. (ahem) Uh the kid Bain been out there hooping. Yeah.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, very dark, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, neutral openness; reads as relief, triumph; style: casual; below-average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 9.3/10; 7.1s, EN.
111638_00345152 · in -15.9 dBFS · gain -4.1 dB · podcast-05498
S_AUTH — style: authoritativek-VN1-k2 · #10

This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: authoritative (S_AUTH) below average — 0.29, lower than 71 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.05, lower than 95 % of clips in this corpus. That is a total fall of 0.23.

It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · en · podcast

k 2d_a -0.233d_b -0.233step_a 0.233step_b 0.233min_cos_consec 0.8990min_cos_anchor 0.8990dataset podcastlang enspeaker 471056track 471056total 32.5slevel spread 1.0 dBmax seam 1.0 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, slightly rough, balanced body, below-average recording, quiet background, normal-paced, fairly steady
(contemplation, sourness, impatience and irritability · normally alert, neutral tension, some disfluency, casual) It is all about knowing what you want. And you go you understand knowing what you want by going through relationships though. Like not I'm not saying necessarily going through mad relationships, but you understand what you want by going through certain things. Cause you're not gonna learn how to be in a relationship just by being in one relationship. Like
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contemplation, sourness, impatience and irritability; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 5.8/6; vocal-burst blend 10.0/10; 16.9s, EN.
471056_00887224 · in -25.2 dBFS · gain +5.2 dB · podcast-00436
(disgust, amusement, bitterness · very low-energy, relaxed, frequent disfluency, casual) She was like she did this lyric prank. And it's basically like where you basically tell somebody like, Oh, feel me. You do like a lyric joke around and it was like the lyric joke around was like, Oh, I'm in love with somebody (ahem) else. It was like lyrics from a song.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disgust, amusement, bitterness; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 5.7/6; vocal-burst blend 10.0/10; 15.4s, EN.
471056_00892384 · in -24.2 dBFS · gain +4.2 dB · podcast-03483
R_MASK — resonance: maskk-VN1-k2 · #11

This is a VoiceNet dimension, not an emotion: resonance: mask (R_MASK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: mask (R_MASK) above average — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the range at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · pt · eurospeech

k 2d_a 0.246d_b 0.246step_a 0.246step_b 0.246min_cos_consec min_cos_anchor dataset eurospeechlang ptspeaker portugal_13_2_103track portugal_13_2_103total 32.9slevel spread 1.6 dBmax seam 1.6 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly cool, neutral-bright, rough, thin, below-average recording, highly aroused, tense, moderately variable
(pride, anger, triumph · measured, frequent disfluency, authoritative, cartoonish) Mas é preciso que o PSD se entenda consigo próprio, porque acaba de dar entrada na Mesa um projeto de lei subscrito por largo consenso da Câmara
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, measured, tense, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, frequent disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as pride, anger, triumph; style: authoritative, cartoonish; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.6/10; 16.7s, PT.
portugal_13_2_103_638295_655024 · in -18.8 dBFS · gain -1.2 dB · eurospeech-02543
(anger, contempt, malevolence malice · brisk, almost no disfluency, dramatic, authoritative) que tem como fundamento — e não posso de deixar de ler — o seguinte: «Compete ao Parlamento criar as condições para que os esclarecimentos devidos possam ser obtidos de forma empenhada, isenta e credível.
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, almost no disfluency, wide pitch range, normal breath; affect is positive, dominant, fairly guarded; reads as anger, contempt, malevolence malice; style: dramatic, authoritative; below-average recording, some background noise; genuineness 1.5/6; vocal-burst blend 2.6/10; 16.0s, PT.
portugal_13_2_103_655024_671072 · in -17.2 dBFS · gain -2.8 dB · eurospeech-02543
S_RANT — style: rantingk-VN1-k2 · #12

This is a VoiceNet dimension, not an emotion: style: ranting (S_RANT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: ranting (S_RANT) below average — 0.32, lower than 68 % of clips in this corpus — and ends with it around average at 0.56, higher than 56 % of clips in this corpus. That is a total rise of 0.24.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 7 s · zh · emolia

k 2d_a 0.244d_b 0.244step_a 0.244step_b 0.244min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00024_S08030track ZH_B00024_S08030total 7.3slevel spread 2.9 dBmax seam 2.9 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, measured, slightly relaxed, light breath
(emotional numbness · subdued, steady, frequent disfluency, formal) 金太祖就是蜿蜒阿骨打。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, whispered; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 1.2/10; 3.0s, ZH.
ZH_B00024_S08030_W000005 · in -22.2 dBFS · gain +2.2 dB · emolia-03512
(normally alert, fairly steady, some disfluency, formal) 那么他向南就是面对的谁呢?就是契丹了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 3.3/10; 4.2s, ZH.
ZH_B00024_S08030_W000006 · in -19.3 dBFS · gain -0.7 dB · emolia-03512
R_CHST — resonance: chestk-VN1-k2 · #13

This is a VoiceNet dimension, not an emotion: resonance: chest (R_CHST) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: chest (R_CHST) around average — 0.46, lower than 54 % of clips in this corpus — and works its way down to low at 0.21, lower than 79 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 38 s · en · podcast

k 2d_a -0.247d_b -0.247step_a 0.247step_b 0.247min_cos_consec 0.9381min_cos_anchor 0.9381dataset podcastlang enspeaker 209404track 209404total 38.4slevel spread 4.4 dBmax seam 4.4 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, measured, relaxed, fairly steady, frequent disfluency, somewhat unclear, normal breath
(pride, shame, contemplation · subdued, moderate pitch range, monologue, whispered) real validation of this path I had taken that it took two years or whatever it was, that I had made my own stepping stone (low mumble)
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as pride, shame, contemplation; style: monologue, whispered; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 2.4/10; 9.7s, EN.
209404_00202184 · in -31.9 dBFS · gain +11.8 dB · podcast-05037
(relief, interest, thankfulness gratitude · very low-energy, fairly narrow pitch, monologue, whispered) and found the end goal of I shouldn't say the end goal, because there's always a future goal, but the the goal of getting into the industry and being around what I wanted to be around and actually signing the contract. Like it it as I think of all the moments of my life of that really resonate with me. I I will always remember the table, the time where I kind of signed the contract, (low mumble) uh my boss had sent it through
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, interest, thankfulness gratitude; style: monologue, whispered; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 4.5/10; 28.6s, EN.
209404_00203148 · in -27.4 dBFS · gain +7.4 dB · podcast-05012
S_NARR — style: narrationk-VN1-k2 · #14

This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: narration (S_NARR) below average — 0.41, lower than 59 % of clips in this corpus — and ends with it above average at 0.62, higher than 62 % of clips in this corpus. That is a total rise of 0.22.

It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.37 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.37 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 27 s · de · emolia

k 2d_a 0.217d_b 0.217step_a 0.217step_b 0.217min_cos_consec min_cos_anchor dataset emolialang despeaker DE_Q92Gb-PapJgtrack DE_Q92Gb-PapJgtotal 27.2slevel spread 0.4 dBmax seam 0.4 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(slightly relaxed, casual, conversational) Okay. Eka, wie war das bei deinen Studierenden?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.5/10; 3.3s, DE.
DE_Q92Gb-PapJg_W000048 · in -22.5 dBFS · gain +2.5 dB · emolia-00157
(interest, jealousy and envy, disappointment · neutral tension, casual) Ja, ähnlich auch. Also, wenn ich mal auf die (ahem) (ahem) Lehramts-Kohorte zurückgreife, da habe ich (ahem) logischerweise das Modul (ahem) KI und Sprache genommen aus dem (ahem) (low mumble) Kurs Schule macht KI. Und das konnte man auch wunderbar nur als einzelnes Modul als Snippet rausnehmen, ohne dass ich den ganzen Kurs machen muss. War durchweg positiv. Die Lehramtsstudien waren aber auch sehr kritisch. Also, ich werde das jetzt mal als kritisch
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as interest, jealousy and envy, disappointment; style: casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.6/10; 23.8s, DE.
DE_Q92Gb-PapJg_W000049 · in -22.9 dBFS · gain +2.9 dB · emolia-00157
EXPL — expressivenessk-VN1-k2 · #15

This is a VoiceNet dimension, not an emotion: expressiveness (EXPL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with expressiveness (EXPL) high — 0.88, higher than 88 % of clips in this corpus — and works its way down to above average at 0.64, higher than 64 % of clips in this corpus. That is a total fall of 0.23.

It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 7 s · zh · emolia

k 2d_a -0.233d_b -0.233step_a 0.233step_b 0.233min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00060_S08495track ZH_B00060_S08495total 7.4slevel spread 5.1 dBmax seam 5.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, normally alert, slightly relaxed, some disfluency, light breath
(normal-paced, fairly steady, slurred, casual) 这边这边也是顺势一波掉了,所以吧。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.1/10; 3.7s, ZH.
ZH_B00060_S08495_W000005 · in -30.5 dBFS · gain +10.5 dB · emolia-03876
(astonishment surprise · fast, moderately variable, average clarity, casual) 对,所以感觉下一把的话,就看他们的阵容是怎么去选择了。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as astonishment surprise; style: casual, conversational; average recording, some background noise; genuineness 5.8/6; vocal-burst blend 4.1/10; 3.5s, ZH.
ZH_B00060_S08495_W000006 · in -25.4 dBFS · gain +5.4 dB · emolia-03876
CHNK — chunking / phrasing densityk-VN1-k2 · #16

This is a VoiceNet dimension, not an emotion: chunking / phrasing density (CHNK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with chunking / phrasing density (CHNK) above average — 0.67, higher than 67 % of clips in this corpus — and works its way down to around average at 0.42, lower than 58 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.37 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.37 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.37, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 22 s · en · podcast

k 2d_a -0.246d_b -0.246step_a 0.246step_b 0.246min_cos_consec 0.3680min_cos_anchor 0.3680dataset podcastlang enspeaker 970642track 970642total 21.8slevel spread 0.9 dBmax seam 0.9 dBcos from orange-id (speaker identity)
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, moderate pitch range
(intoxication altered states of consciousness, affection, contemplation · subdued, relaxed, fairly steady, casual) But you know, I I've been around you and I've seen you doing it. I think when we go with you know, a couple of times when we came up when you came up here and helped me working and stuff like that. I I noticed both of us kind of disassociating in those periods.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, affection, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 10.0/10; 14.7s, EN.
970642_00062800 · in -15.0 dBFS · gain -5.0 dB · podcast-05624
(doubt, confusion, interest · normally alert, slightly relaxed, moderately variable, casual) So what is it? I mean, I'm really curious. What what is that look like? Disassociate? You mean just kind of like drift off or
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, confusion, interest; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.6/10; 7.0s, EN.
970642_00064288 · in -14.1 dBFS · gain -5.9 dB · podcast-05619
SMTH — smoothnessk-VN1-k2 · #17

This is a VoiceNet dimension, not an emotion: smoothness (SMTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with smoothness (SMTH) at the very bottom of the range — 0.07, lower than 93 % of clips in this corpus — and ends with it below average at 0.32, lower than 68 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · hr · eurospeech

k 2d_a 0.250d_b 0.250step_a 0.250step_b 0.250min_cos_consec min_cos_anchor dataset eurospeechlang hrspeaker croatia_20121018174934-172track croatia_20121018174934-172total 30.4slevel spread 6.8 dBmax seam 6.8 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · balanced body, average recording, quiet background, normally alert, somewhat unclear, light breath
(triumph, thankfulness gratitude, pride · normal-paced, slightly relaxed, steady, monologue) Želi li predstavnica predlagatelja dodatno obrazložiti prijedlog? Da, poštovana zamjenica ministra Sandra Artuković KUnšt. Izvolite.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, thankfulness gratitude, pride; style: monologue, narration; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 11.4s, HR.
croatia_20121018174934-1722_44976_56337 · in -31.1 dBFS · gain +11.1 dB · eurospeech-01358
(shame, thankfulness gratitude, pride · measured, neutral tension, moderately variable, casual) Hvala (low mumble) vam lijepo.
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, thankfulness gratitude, pride; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.4/10; 18.9s, HR.
croatia_20121018174934-1722_56337_75216 · in -24.3 dBFS · gain +4.3 dB · eurospeech-01358
FOCS — vocal focusk-VN1-k2 · #18

This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with vocal focus (FOCS) around average — 0.47, lower than 53 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.22.

It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.95 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.95 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 14 s · en · emolia

k 2d_a 0.220d_b 0.220step_a 0.220step_b 0.220min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_9TDULxjsE0Itrack EN_9TDULxjsE0Itotal 14.4slevel spread 0.8 dBmax seam 0.8 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration · steady, formal, authoritative) Not all materials need to be decomposed fully. Coal, a fossil fuel formed over vast tracts of time in swamp ecosystems, is one example
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.6s, EN.
EN_9TDULxjsE0I_W000168 · in -15.3 dBFS · gain -4.7 dB · emolia-01386
(fairly steady, formal, authoritative) Contemporary evolutionary theory sees death as an important part of the process of natural selection
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.4/10; 5.7s, EN.
EN_9TDULxjsE0I_W000170 · in -14.5 dBFS · gain -5.5 dB · emolia-01386
S_MONO — style: monologuek-VN1-k2 · #19

This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: monologue (S_MONO) below average — 0.31, lower than 69 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.06, lower than 94 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 35 s · en · podcast

k 2d_a -0.243d_b -0.243step_a 0.243step_b 0.243min_cos_consec 0.8879min_cos_anchor 0.8879dataset podcastlang enspeaker 615558track 615558total 34.7slevel spread 1.1 dBmax seam 1.1 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · slightly cool, slightly bright, thin, brisk, variable, some disfluency, very clear, wide pitch range
(contemplation, concentration, interest · energised, neutral tension, dramatic, ranting) When we look at people who are saved in the Bible and we compare their lives with our lives, we usually think to ourselves, yeah, that's not me, that's probably the pastor or one of the elders. In fact, look what the Bible says about Noah. And I think this is very key. Look at Genesis chapter 6 one more time. Let's start with verse 9.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, normal breath; affect is positive, dominant, fairly guarded; reads as contemplation, concentration, interest; style: dramatic, ranting; average recording, some background noise; genuineness 1.9/6; vocal-burst blend 1.6/10; 21.5s, EN.
615558_00229260 · in -23.6 dBFS · gain +3.6 dB · podcast-06340
(contempt, pride, anger · highly aroused, tense, dramatic, cartoonish) This is the genealogy of who? (ahem) Noah. Noah was a what? Just man. Now, how many people are gonna say to themselves, right now, well, I'm a just person.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, variable; timbre is slightly cool, slightly bright, slightly rough, thin; very clear, some disfluency, wide pitch range, normal breath; affect is elated, very dominant, fairly guarded; reads as contempt, pride, anger; style: dramatic, cartoonish; below-average recording, quiet background; mildly explicit content; genuineness 2.3/6; vocal-burst blend 1.0/10; 13.0s, EN.
615558_00231412 · in -22.6 dBFS · gain +2.6 dB · podcast-04334
R_NASL — resonance: nasalk-VN1-k2 · #20

This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: nasal (R_NASL) below average — 0.38, lower than 62 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 16 s · snippets

k 2d_a -0.237d_b -0.237step_a 0.237step_b 0.237min_cos_consec min_cos_anchor dataset snippetslang ?speaker batch265_part3_batch265_patrack batch265_part3_batch265_patotal 15.7slevel spread 12.2 dBmax seam 12.2 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(casual, monologue) The relative wage would depends on on on the relative employment, right?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.6/10; 3.6s.
batch265_part3_batch265_part3_chunk_859_1_862376 · in -15.3 dBFS · gain -4.7 dB · snippets-00861
(pride, relief · conversational, casual) Uh, (low mumble) you can also, you can also do something statistically where you say, okay, you have retail sector here and you find most people have a high school and then there's some people working there who have a college degree and you say, hey, they stick out.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, relief; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.6/10; 12.0s.
batch265_part3_batch265_part3_chunk_859_1_862430 · in -27.6 dBFS · gain +7.6 dB · snippets-00861