c-snippets-PXR

Corpus snippets in isolation, rule PXR.

Rule. PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable
Source. trajectories_v5.parquet  |  Family. one corpus in isolation
Sampled from 4,248 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Infatuation ↓  /  Emotional Numbnessc-snippets-PXR · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.27.

At the same time Infatuation goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.00 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · snippets

hear it un-normalised (raw levels, max seam 5.9 dB)
k 3d_a -0.424d_b 0.272step_a 0.216step_b 0.149min_cos_consec -0.0038min_cos_anchor 0.0706dataset snippetslang ?speaker batch214_part4_batch214_patrack batch214_part4_batch214_patotal 28.9slevel spread 8.3 dBmax seam 5.9 dBcos from orange-id (speaker identity)
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult somewhat feminine voice · good recording, normally alert, slightly relaxed
(infatuation, affection, longing · normal-paced, fairly steady, some disfluency, casual) And I said, (low mumble) uh, my name and who I was related to and who he was related to, because he the family was a first cousin.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, affection, longing; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.4/10; 9.3s.
batch214_part4_batch214_part4_chunk_3_1_73922 · in -14.1 dBFS · gain -5.9 dB · snippets-00594
(slow, steady, frequent disfluency, monologue) globally that generates about a sixth of Ford's global automotive revenue. This strike includes
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is warm, slightly dark, slightly rough, full; very clear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; no dominant emotion; style: monologue, storytelling; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 9.2s.
batch214_part4_batch214_part4_chunk_3_1_74107 · in -20.0 dBFS · gain -0.0 dB · snippets-00594
(normal-paced, steady, no disfluency, narration) Drawing from the teachings of Ninjitsu, Chloe embodies the principle of hiding in plain sight, an unsuspecting flower company executive by day, and a relentless vigilante by night.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.1s.
batch214_part4_batch214_part4_chunk_3_1_74240 · in -22.4 dBFS · gain +2.5 dB · snippets-00594
Hope Enthusiasm Optimism ↓  /  Infatuationc-snippets-PXR · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.40.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.24, then +0.02 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.35 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.35, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · snippets

hear it un-normalised (raw levels, max seam 0.5 dB)
k 4d_a -0.300d_b 0.397step_a 0.187step_b 0.239min_cos_consec 0.1271min_cos_anchor 0.3512dataset snippetslang ?speaker batch2_part3_batch2_part3_track batch2_part3_batch2_part3_total 30.0slevel spread 0.8 dBmax seam 0.5 dBcos from orange-id (speaker identity)
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, fairly steady, average clarity
(hope enthusiasm optimism · fully relaxed, some disfluency, casual, conversational) and he's going to be one of those guys. I think Nate's legacy in freestyle motocross is bringing a professionalism.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 6.2/10; 6.4s.
batch2_part3_batch2_part3_chunk_1010_1_1017301 · in -33.0 dBFS · gain +13.0 dB · snippets-01045
(sadness, contentment, contemplation · neutral tension, frequent disfluency, casual, whispered) And I know that's what keeps him really centered and really narrowed with what what the choices are that he makes. You know, I think Nate and I hit it off so well because we have the same faith, we're both Christian and you know, we I think God plays a huge, Jesus plays a huge emphasis in our lives.
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sadness, contentment, contemplation; style: casual, whispered; average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 3.3/10; 15.5s.
batch2_part3_batch2_part3_chunk_1010_1_1017336 · in -33.4 dBFS · gain +13.4 dB · snippets-01045
(infatuation, contentment, embarrassment · slightly relaxed, frequent disfluency, casual, conversational) really (ahem) is. I'd describe Nate as (low mumble) um, just a really nice, humble guy.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as infatuation, contentment, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 3.3/10; 4.4s.
batch2_part3_batch2_part3_chunk_1010_1_1017386 · in -33.1 dBFS · gain +13.1 dB · snippets-01045
(infatuation, pride, triumph · slightly relaxed, little disfluency, casual, conversational) That was my second contest I've ever ridden professionally.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, pride, triumph; style: casual, conversational; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 3.2/10; 3.2s.
batch2_part3_batch2_part3_chunk_1010_1_1017418 · in -32.6 dBFS · gain +12.7 dB · snippets-01045
Contemplation ↓  /  Fatigue Exhaustionc-snippets-PXR · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.30.

At the same time Contemplation goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.20 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · snippets

hear it un-normalised (raw levels, max seam 0.2 dB)
k 3d_a -0.271d_b 0.296step_a 0.147step_b 0.199min_cos_consec 0.7546min_cos_anchor 0.7546dataset snippetslang ?speaker batch23_part0_batch23_parttrack batch23_part0_batch23_parttotal 23.0slevel spread 0.2 dBmax seam 0.2 dBcos from orange-id (speaker identity)
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady, moderate pitch range
(measured, some disfluency, somewhat unclear, casual) graphic novel ideas that we've transitioned over into (ahem) visual novels. So, uh, (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.8/10; 6.1s.
batch23_part0_batch23_part0_chunk_1205_1_1123248 · in -22.3 dBFS · gain +2.3 dB · snippets-00727
(distress, helplessness, fear · normal-paced, some disfluency, average clarity, casual) You're going to have failures a lot. If you're the type of person that collapses and has a has a breakdown every time there's a failure,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, helplessness, fear; style: casual, monologue; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.5/10; 8.4s.
batch23_part0_batch23_part0_chunk_1205_1_1123295 · in -22.5 dBFS · gain +2.5 dB · snippets-00727
(measured, frequent disfluency, average clarity, casual) getting close to half. And (ahem) it morphs the music and uh (low mumble) film industry, (ahem) uh the gaming industry more morphs.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.8/10; 8.2s.
batch23_part0_batch23_part0_chunk_1205_1_1123312 · in -22.4 dBFS · gain +2.4 dB · snippets-00727
Malevolence Malice ↓  /  Fatigue Exhaustionc-snippets-PXR · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion around average — 0.42, lower than 58 % of clips in this corpus — and ends with it clearly present at 0.71, higher than 71 % of clips in this corpus. That is a total rise of 0.29.

At the same time Malevolence Malice goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.05, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · snippets

hear it un-normalised (raw levels, max seam 2.9 dB)
k 3d_a -0.261d_b 0.290step_a 0.142step_b 0.239min_cos_consec 0.9092min_cos_anchor 0.8997dataset snippetslang ?speaker batch278_part4_batch278_patrack batch278_part4_batch278_patotal 22.8slevel spread 2.9 dBmax seam 2.9 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(brisk, almost no disfluency, clear, formal) The herbicide Orange was destroyed by incineration aboard a German incinerator ship, MV Vulcanus, in 1977.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.8/10; 6.7s.
batch278_part4_batch278_part4_chunk_975_1_1189051 · in -18.6 dBFS · gain -1.4 dB · snippets-00931
(normal-paced, almost no disfluency, clear, narration) Marine Jim Divine described the changes to Sand Island in 1944. Sand Island at the time consisted of two islands connected by a single lane coral roadway about 600 feet long.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 9.2s.
batch278_part4_batch278_part4_chunk_975_1_1189074 · in -21.5 dBFS · gain +1.5 dB · snippets-00931
(brisk, little disfluency, average clarity, playful) One of the island's anti-aircraft batteries have been destined for Wake Island before the island fell to the Japanese on December 23rd, 1941.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, authoritative; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.7/10; 6.6s.
batch278_part4_batch278_part4_chunk_975_1_1189242 · in -19.8 dBFS · gain -0.2 dB · snippets-00931
Contempt ↓  /  Thankfulness Gratitudec-snippets-PXR · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.31.

At the same time Contempt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · snippets

hear it un-normalised (raw levels, max seam 3.6 dB)
k 3d_a -0.376d_b 0.314step_a 0.248step_b 0.220min_cos_consec 0.1483min_cos_anchor 0.1483dataset snippetslang ?speaker batch135_part1_batch135_patrack batch135_part1_batch135_patotal 17.7slevel spread 3.6 dBmax seam 3.6 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, clear, moderate pitch range
(contempt, disgust, bitterness · measured, steady, almost no disfluency, formal) Like everyone else, indigenous children need certain books, not any books. Here's why.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust, bitterness; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.0/10; 6.0s.
batch135_part1_batch135_part1_chunk_2212_1_2084823 · in -18.4 dBFS · gain -1.6 dB · snippets-00185
(confusion, doubt · measured, fairly steady, some disfluency, didactic) But he also has some questions for Ju Galang. What if Cao Cao does not come my way?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, doubt; style: didactic, formal; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.8/10; 6.0s.
batch135_part1_batch135_part1_chunk_2212_1_2084843 · in -22.1 dBFS · gain +2.0 dB · snippets-00185
(thankfulness gratitude · brisk, fairly steady, no disfluency, casual) This episode was produced by Brett Bachman and Rebecca Ramirez, and edited by Andrea Kissick.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude; style: casual, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.0/10; 5.3s.
batch135_part1_batch135_part1_chunk_2212_1_2084919 · in -21.4 dBFS · gain +1.4 dB · snippets-00185
Doubt ↓  /  Fatigue Exhaustionc-snippets-PXR · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.31.

At the same time Doubt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.13 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.13, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 10 s · snippets

k 3d_a -0.354d_b 0.310step_a 0.214step_b 0.193min_cos_consec 0.0832min_cos_anchor 0.1317dataset snippetslang ?speaker batch73_part3_batch73_parttrack batch73_part3_batch73_parttotal 10.2slevel spread 9.3 dBmax seam 9.3 dBcos from orange-id (speaker identity)
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, quiet background, normally alert, fully relaxed, fairly steady, moderate pitch range
(doubt · measured, frequent disfluency, slurred, casual) (ahem) um and (low mumble) uh he talked about, you know,
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: casual, conversational; poor recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.2/10; 3.1s.
batch73_part3_batch73_part3_chunk_1658_1_1546962 · in -27.2 dBFS · gain +7.2 dB · snippets-01267
(disgust, pain, helplessness · normal-paced, frequent disfluency, slurred, casual) And like killing themselves or anything. Girls? No.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disgust, pain, helplessness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.7/10; 3.2s.
batch73_part3_batch73_part3_chunk_1658_1_1547030 · in -28.0 dBFS · gain +8.0 dB · snippets-01267
(fatigue exhaustion, emotional numbness · normal-paced, some disfluency, somewhat unclear, casual) yera ma kipe (ahem) (low mumble) agying wana nabatigitsi
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as fatigue exhaustion, emotional numbness; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 2.6/10; 3.6s.
batch73_part3_batch73_part3_chunk_1658_1_1547165 · in -18.7 dBFS · gain -1.3 dB · snippets-01267
Infatuation ↓  /  Sournessc-snippets-PXR · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sourness below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.62.

At the same time Infatuation goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.16, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.27 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 15 s · snippets

k 4d_a -0.501d_b 0.617step_a 0.246step_b 0.228min_cos_consec 0.2739min_cos_anchor 0.0647dataset snippetslang ?speaker batch27_part3_batch27_parttrack batch27_part3_batch27_parttotal 15.1slevel spread 9.2 dBmax seam 4.8 dBcos from orange-id (speaker identity)
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(infatuation · normal-paced, some disfluency, average clarity, conversational) And (low mumble) uh, but beyond that, you never heard about White Canvas trip again.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: conversational, casual; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 2.6/10; 3.7s.
batch27_part3_batch27_part3_chunk_1236_1_885437 · in -18.2 dBFS · gain -1.8 dB · snippets-00940
(measured, frequent disfluency, somewhat unclear, casual) inappropriate, (low mumble) uh, and probably illegal, (low mumble) uh.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; style: casual, conversational; average recording, no background noise; no dominant emotion; genuineness 4.2/6; vocal-burst blend 1.7/10; 3.2s.
batch27_part3_batch27_part3_chunk_1236_1_885592 · in -23.0 dBFS · gain +3.0 dB · snippets-00940
(normal-paced, some disfluency, average clarity, casual) posted on Twitter saying, does anyone have any information on this person's name?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 2.5/10; 3.8s.
batch27_part3_batch27_part3_chunk_1236_1_885829 · in -27.4 dBFS · gain +7.4 dB · snippets-00940
(normal-paced, some disfluency, average clarity, casual) George Soros doesn't even hold a candlestick to Peter Thiel in terms of (low mumble) uh.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; explicit content; genuineness 3.4/6; vocal-burst blend 2.1/10; 3.9s.
batch27_part3_batch27_part3_chunk_1236_1_885953 · in -24.6 dBFS · gain +4.5 dB · snippets-00940
Intoxication Altered States of Consciousness ↓  /  Concentrationc-snippets-PXR · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.31.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.75 (higher than 75 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.01, then +0.12, then +0.20 — not a clean run: step 1 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 31 s · snippets

k 4d_a -0.256d_b 0.312step_a 0.246step_b 0.201min_cos_consec 0.7748min_cos_anchor 0.6978dataset snippetslang ?speaker batch79_part0_batch79_parttrack batch79_part0_batch79_parttotal 30.7slevel spread 4.3 dBmax seam 3.9 dBcos from orange-id (speaker identity)
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, light breath
(steady, almost no disfluency, clear, monologue) That is the root cause and because of that he was skeptical to accept Sita.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.5/10; 5.0s.
batch79_part0_batch79_part0_chunk_1704_1_1535438 · in -19.1 dBFS · gain -0.9 dB · snippets-01293
(pain · fairly steady, some disfluency, average clarity, monologue) And that is what people have always been doing with regards to Uttara Ramayan. And the reason for this is their strong emotional connect.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pain; style: monologue, casual; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.2/10; 7.5s.
batch79_part0_batch79_part0_chunk_1704_1_1535454 · in -16.9 dBFS · gain -3.1 dB · snippets-01293
(steady, little disfluency, clear, monologue) So that is Skanda Puranam for you, standing as a strong argument that Uttara Kanda or Uttara Ramayanam is authentic and is for real.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.9/10; 8.5s.
batch79_part0_batch79_part0_chunk_1704_1_1535701 · in -17.3 dBFS · gain -2.7 dB · snippets-01293
(concentration, emotional numbness, contemplation · fairly steady, some disfluency, somewhat unclear, monologue) That is the reason people often jump to the conclusions that Uttara Ramayanam is very contentious or controversial and that is the reason it is a prakship, it is not part of Ramayan.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, emotional numbness, contemplation; style: monologue; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 1.4/10; 9.2s.
batch79_part0_batch79_part0_chunk_1704_1_1535805 · in -21.2 dBFS · gain +1.2 dB · snippets-01293
Concentration ↓  /  Emotional Numbnessc-snippets-PXR · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.40.

At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.07 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · snippets

k 3d_a -0.382d_b 0.399step_a 0.228step_b 0.233min_cos_consec 0.0673min_cos_anchor 0.0673dataset snippetslang ?speaker batch76_part0_batch76_parttrack batch76_part0_batch76_parttotal 22.5slevel spread 5.0 dBmax seam 3.3 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, moderate pitch range, light breath
(concentration · measured, steady, almost no disfluency, narration) If you take the measurement of the base perimeter of the great pyramid, go along and measure each of its four sides and add that together.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: narration, formal; good recording, no background noise; explicit content; genuineness 1.3/6; vocal-burst blend 0.9/10; 7.4s.
batch76_part0_batch76_part0_chunk_167_1_363035 · in -35.6 dBFS · gain +15.6 dB · snippets-01278
(awe · brisk, fairly steady, almost no disfluency, newsreading) lithic building with gigantic sculpted stones is considered by almost all the experts to have begun around 2,500 BC. From the Great Pyramid of Giza to Stonehenge in England.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe; style: newsreading, narration; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.0/10; 11.3s.
batch76_part0_batch76_part0_chunk_167_1_363113 · in -32.3 dBFS · gain +12.3 dB · snippets-01278
(slow, fairly steady, some disfluency, casual) And so they existed on the peripheries of France, of
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 3.5s.
batch76_part0_batch76_part0_chunk_167_1_363140 · in -30.6 dBFS · gain +10.6 dB · snippets-01278
Concentration ↓  /  Emotional Numbnessc-snippets-PXR · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.34.

At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.05 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · snippets

k 3d_a -0.264d_b 0.338step_a 0.210step_b 0.215min_cos_consec -0.0497min_cos_anchor -0.1447dataset snippetslang ?speaker batch108_part3_batch108_patrack batch108_part3_batch108_patotal 27.1slevel spread 5.3 dBmax seam 5.3 dBcos from orange-id (speaker identity)
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth
(concentration · measured, normally alert, slightly relaxed, didactic) For this I'm going to use this data which has X1 to X 19 variable and the Y variable.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.8/10; 8.6s.
batch108_part3_batch108_part3_chunk_1970_1_2189316 · in -23.6 dBFS · gain +3.5 dB · snippets-00041
(relief, pride · normal-paced, normally alert, neutral tension, conversational) for the world. And so you can probably tell that (ahem) that myself and many other aid workers kind of view the net benefit as pretty limited in relation to what's going on.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, pride; style: conversational, casual; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 6.3/10; 8.9s.
batch108_part3_batch108_part3_chunk_1970_1_2189343 · in -28.9 dBFS · gain +8.9 dB · snippets-00041
(slow, very low-energy, slightly relaxed, monologue) Some, more than, increased by, plus, and together, all mean addition.
full caption & clip details
A child feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is cool, slightly bright, fairly smooth, very thin; crisply articulate, almost no disfluency, very wide pitch range, minimal breath; affect is positive, neutral stance, neutral openness; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.8/10; 9.3s.
batch108_part3_batch108_part3_chunk_1970_1_2189482 · in -25.6 dBFS · gain +5.6 dB · snippets-00041
Affection ↓  /  Contemplationc-snippets-PXR · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.33.

At the same time Affection goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.25 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.25 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.25, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 24 s · snippets

k 3d_a -0.314d_b 0.330step_a 0.175step_b 0.181min_cos_consec 0.2495min_cos_anchor 0.2495dataset snippetslang ?speaker batch48_part2_batch48_parttrack batch48_part2_batch48_parttotal 24.0slevel spread 3.9 dBmax seam 3.9 dBcos from orange-id (speaker identity)
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, normal-paced, normally alert
(affection, interest · slightly relaxed, fairly steady, some disfluency, casual) a fleet of (ahem) um thinking, a lead if you will of thinking women. She is participatory in politics. She was president of the Republicans Colored Women's League. She's a part of the national.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, interest; style: casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.5/10; 13.4s.
batch48_part2_batch48_part2_chunk_1428_1_1598057 · in -21.9 dBFS · gain +1.9 dB · snippets-01135
(slightly relaxed, fairly steady, some disfluency, casual) Can you talk a little bit about that and talk about the importance?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.2/10; 3.8s.
batch48_part2_batch48_part2_chunk_1428_1_1598270 · in -18.0 dBFS · gain -2.0 dB · snippets-01135
(neutral tension, moderately variable, frequent disfluency, casual) of how she reads the period of enslavement. So what she's doing is changing
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.6/10; 6.5s.
batch48_part2_batch48_part2_chunk_1428_1_1598323 · in -18.9 dBFS · gain -1.1 dB · snippets-01135
Sexual Lust ↓  /  Intoxication Altered States of Consciousnessc-snippets-PXR · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness around average — 0.53, higher than 53 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.26.

At the same time Sexual Lust goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 13 s · snippets

k 3d_a -0.258d_b 0.257step_a 0.159step_b 0.221min_cos_consec 0.7360min_cos_anchor 0.7137dataset snippetslang ?speaker batch272_part3_batch272_patrack batch272_part3_batch272_patotal 13.3slevel spread 0.9 dBmax seam 0.9 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, narration) To do so, first professional divers were hired to blast bombs underwater.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 4.9s.
batch272_part3_batch272_part3_chunk_918_1_1036433 · in -23.4 dBFS · gain +3.4 dB · snippets-00901
(narration, monologue) For the south side, the hard strata was 50 feet below the seabed level.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.4/10; 4.4s.
batch272_part3_batch272_part3_chunk_918_1_1036690 · in -24.2 dBFS · gain +4.2 dB · snippets-00901
(formal, monologue) Now, the length of the unsupported bridge deck is reduced.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.9/10; 3.6s.
batch272_part3_batch272_part3_chunk_918_1_1036753 · in -23.2 dBFS · gain +3.2 dB · snippets-00901
Sexual Lust ↓  /  Pleasure Ecstasyc-snippets-PXR · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Pleasure Ecstasy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pleasure Ecstasy clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.

At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.18, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.41 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.41 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.41, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 53 s · snippets

k 4d_a -0.341d_b 0.360step_a 0.187step_b 0.185min_cos_consec 0.4056min_cos_anchor 0.4056dataset snippetslang ?speaker batch33_part0_batch33_parttrack batch33_part0_batch33_parttotal 53.4slevel spread 6.7 dBmax seam 5.9 dBcos from orange-id (speaker identity)
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, average clarity, light breath
(sexual lust, sourness, amusement · brisk, energised, neutral tension, casual) Like and it she there needs to be some like the thing underneath has to be it it was the same red vinyl latex so you couldn't tell that something changed.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sexual lust, sourness, amusement; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 8.0/10; 8.4s.
batch33_part0_batch33_part0_chunk_1297_1_1090726 · in -21.8 dBFS · gain +1.8 dB · snippets-01062
(infatuation, sexual lust, amusement · normal-paced, normally alert, slightly relaxed, casual) uh, (low mumble) that she got to sing when you got to mama on on television.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as infatuation, sexual lust, amusement; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.8/10; 4.8s.
batch33_part0_batch33_part0_chunk_1297_1_1090911 · in -24.2 dBFS · gain +4.2 dB · snippets-01062
(confusion, astonishment surprise, amusement · brisk, energised, neutral tension, casual) I was, I don't know why Kevin Bacon was the one that really took me out. I was like, this is the last person I expected to see on, I know why he's there. I was just
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, astonishment surprise, amusement; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 9.2/10; 8.2s.
batch33_part0_batch33_part0_chunk_1297_1_1090935 · in -22.6 dBFS · gain +2.6 dB · snippets-01062
(pleasure ecstasy, pride, hope enthusiasm optimism · brisk, energised, slightly relaxed, casual) Now Alma has a diverse network of therapists to fit your unique needs. In the easy-to-use directory, you can filter for gender, sexual orientation, race, etcetera. Every therapist at Alma has a detailed profile so you can get a better picture of their working style, expertise, and even why they pursued a career in mental health. And it's sometimes important to make sure that your, in fact it is always important to make sure that your therapist is compatible with you and learning about their experience helps a lot. I think it is so important to take care of your mental health and I love how easy Alma makes it.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, pride, hope enthusiasm optimism; style: casual, conversational; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 7.1/10; 31.5s.
batch33_part0_batch33_part0_chunk_1297_1_1091115 · in -28.5 dBFS · gain +8.5 dB · snippets-01062
Concentration ↓  /  Contemplationc-snippets-PXR · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.30.

At the same time Concentration goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.05 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 21 s · snippets

k 3d_a -0.377d_b 0.299step_a 0.203step_b 0.246min_cos_consec 0.7188min_cos_anchor 0.7013dataset snippetslang ?speaker batch71_part3_batch71_parttrack batch71_part3_batch71_parttotal 20.8slevel spread 1.9 dBmax seam 1.9 dBcos from orange-id (speaker identity)
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, average clarity
(concentration · slightly relaxed, fairly steady, some disfluency, authoritative) So we randomized people between two conditions. They either get the (ahem) meditation while they're undergoing ultraviolet light or they just get the ultraviolet light by itself.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: authoritative, conversational; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.5/10; 8.9s.
batch71_part3_batch71_part3_chunk_1636_1_1902236 · in -19.4 dBFS · gain -0.6 dB · snippets-01257
(contemplation · neutral tension, moderately variable, frequent disfluency, conversational) (ahem) Uh, if you stare at that word for too long, it doesn't mean anything, as you know. But, (ahem) uh, I want to make a distinc- (ahem) a distinction between
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.3/10; 8.0s.
batch71_part3_batch71_part3_chunk_1636_1_1902334 · in -21.3 dBFS · gain +1.3 dB · snippets-01257
(contemplation, doubt · slightly relaxed, fairly steady, frequent disfluency, casual) so when the doing comes out of being I think you know then (surprised gasp)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, doubt; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 4.3/10; 3.6s.
batch71_part3_batch71_part3_chunk_1636_1_1902368 · in -19.7 dBFS · gain -0.3 dB · snippets-01257
Intoxication Altered States of Consciousness ↓  /  Infatuationc-snippets-PXR · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.49.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.24, then +0.23 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.21 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.09 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.21, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 18 s · snippets

k 4d_a -0.425d_b 0.494step_a 0.231step_b 0.240min_cos_consec -0.0856min_cos_anchor -0.2058dataset snippetslang ?speaker batch247_part4_batch247_patrack batch247_part4_batch247_patotal 17.9slevel spread 2.8 dBmax seam 2.8 dBcos from orange-id (speaker identity)
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, measured, normally alert, fairly steady
(relaxed, frequent disfluency, slurred, casual) मैं जो बोता हूं वह तीन से चार एकड़ गेहूं में बोता हूं।
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 2.9/10; 4.9s.
batch247_part4_batch247_part4_chunk_693_1_496782 · in -30.7 dBFS · gain +10.7 dB · snippets-00771
(slightly relaxed, no disfluency, clear, formal) around 7 million tons were exported to other countries.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.6/10; 3.5s.
batch247_part4_batch247_part4_chunk_693_1_496793 · in -29.7 dBFS · gain +9.7 dB · snippets-00771
(emotional numbness · slightly relaxed, little disfluency, clear, formal) After 107 million tons of wheat produced in 2021,
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.5/10; 4.5s.
batch247_part4_batch247_part4_chunk_693_1_496817 · in -28.0 dBFS · gain +8.0 dB · snippets-00771
(slightly relaxed, no disfluency, clear, narration) India ranked 101 out of 116 countries.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: narration, storytelling; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.1/10; 4.4s.
batch247_part4_batch247_part4_chunk_693_1_497240 · in -30.8 dBFS · gain +10.8 dB · snippets-00771
Emotional Numbness ↓  /  Fearc-snippets-PXR · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.40.

At the same time Emotional Numbness goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.01, then +0.19 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.07 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 25 s · snippets

k 4d_a -0.396d_b 0.400step_a 0.245step_b 0.197min_cos_consec -0.0665min_cos_anchor -0.1356dataset snippetslang ?speaker batch77_part1_batch77_parttrack batch77_part1_batch77_parttotal 25.0slevel spread 2.1 dBmax seam 2.1 dBcos from orange-id (speaker identity)
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · measured
(subdued, relaxed, steady, monologue) is (low mumble) uh, meet the sea. So, due to sea level rise, (low mumble) um, there's a large number of sun energy infusion which is
full caption & clip details
A child masculine voice; delivery is subdued, measured, relaxed, steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.6/10; 8.4s.
batch77_part1_batch77_part1_chunk_1690_1_1998559 · in -24.4 dBFS · gain +4.4 dB · snippets-01284
(doubt, longing, distress · normally alert, slightly relaxed, fairly steady, casual) a little bit about Bangladesh government right now, but how is the relationship going with them?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, longing, distress; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.5/10; 5.1s.
batch77_part1_batch77_part1_chunk_1690_1_1998641 · in -23.9 dBFS · gain +3.9 dB · snippets-01284
(subdued, slightly relaxed, fairly steady, casual) policies that has been amended in terms of including climate change
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; below-average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.6/10; 5.1s.
batch77_part1_batch77_part1_chunk_1690_1_1998721 · in -25.9 dBFS · gain +5.9 dB · snippets-01284
(fear · normally alert, slightly relaxed, fairly steady, whispered) just like systems or things that can change that would help you in your battle against climate change.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fear; style: whispered, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.4/10; 6.0s.
batch77_part1_batch77_part1_chunk_1690_1_1998736 · in -25.4 dBFS · gain +5.5 dB · snippets-01284
Astonishment Surprise ↓  /  Infatuationc-snippets-PXR · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.28.

At the same time Astonishment Surprise goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.09, then +0.16, then +0.21 — not a clean run: step 1 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 23 s · snippets

k 4d_a -0.410d_b 0.280step_a 0.165step_b 0.210min_cos_consec 0.6970min_cos_anchor 0.6970dataset snippetslang ?speaker batch183_part2_batch183_patrack batch183_part2_batch183_patotal 23.4slevel spread 1.1 dBmax seam 1.1 dBcos from orange-id (speaker identity)
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, no disfluency, formal, monologue) is apparently derived from the name of the goddess Mumba.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.0/10; 3.1s.
batch183_part2_batch183_part2_chunk_263_1_402087 · in -25.9 dBFS · gain +6.0 dB · snippets-00434
(normal-paced, almost no disfluency, newsreading, formal) These cultural characteristics are defined by various influences, including geographical, spiritual, and agricultural considerations.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 8.4s.
batch183_part2_batch183_part2_chunk_263_1_402172 · in -26.4 dBFS · gain +6.4 dB · snippets-00434
(normal-paced, almost no disfluency, monologue, storytelling) His holiness, the Dalai Lama, is the highest political as well as spiritual authority of Tibet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, storytelling; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.7/10; 5.7s.
batch183_part2_batch183_part2_chunk_263_1_402247 · in -25.3 dBFS · gain +5.3 dB · snippets-00434
(infatuation · brisk, almost no disfluency, conversational, narration) Do you find it difficult to make up your mind? A private fashion show with professional models is here to help you.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as infatuation; style: conversational, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.2/10; 5.7s.
batch183_part2_batch183_part2_chunk_263_1_402302 · in -25.9 dBFS · gain +5.9 dB · snippets-00434
Fatigue Exhaustion ↓  /  Intoxication Altered States of Consciousnessc-snippets-PXR · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.40.

At the same time Fatigue Exhaustion goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 13 s · snippets

k 3d_a -0.328d_b 0.403step_a 0.210step_b 0.242min_cos_consec 0.0171min_cos_anchor 0.0411dataset snippetslang ?speaker batch212_part0_batch212_patrack batch212_part0_batch212_patotal 13.1slevel spread 14.0 dBmax seam 14.0 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, light breath
(measured, fairly steady, some disfluency, whispered) We definitely started to see a lot of traction coming out of the Philippines, especially after
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.4/10; 5.8s.
batch212_part0_batch212_part0_chunk_379_1_37395 · in -28.8 dBFS · gain +8.8 dB · snippets-00583
(contemplation · measured, steady, no disfluency, monologue) fundamental change in the nature of work.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, whispered; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.8/10; 3.0s.
batch212_part0_batch212_part0_chunk_379_1_37418 · in -29.8 dBFS · gain +9.8 dB · snippets-00583
(intoxication altered states of consciousness · fast, fairly steady, frequent disfluency, casual) Kahit po sana kumita po ng kahit 3000 po a month.
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.1/10; 4.0s.
batch212_part0_batch212_part0_chunk_379_1_37532 · in -15.8 dBFS · gain -4.2 dB · snippets-00583
Impatience and Irritability ↓  /  Emotional Numbnessc-snippets-PXR · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.30.

At the same time Impatience and Irritability goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.21, then -0.02 — not a clean run: step 3 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 16 s · snippets

k 4d_a -0.327d_b 0.300step_a 0.212step_b 0.212min_cos_consec 0.0597min_cos_anchor 0.0610dataset snippetslang ?speaker batch112_part3_batch112_patrack batch112_part3_batch112_patotal 16.1slevel spread 3.4 dBmax seam 2.2 dBcos from orange-id (speaker identity)
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · slightly relaxed, fairly steady
(impatience and irritability, intoxication altered states of consciousness, distress · fast, normally alert, some disfluency, casual) 訳達者、その芝居出しとったところが手前から約
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, dark, very rough, thin; slurred, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, guarded; reads as impatience and irritability, intoxication altered states of consciousness, distress; style: casual, playful; poor recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.6/10; 3.4s.
batch112_part3_batch112_part3_chunk_2007_1_2065547 · in -32.2 dBFS · gain +12.2 dB · snippets-00066
(pain · fast, normally alert, some disfluency, casual) (low mumble) 杨教授呢,就是因为杨教授是担任这个章程的翻译。
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.5/10; 3.9s.
batch112_part3_batch112_part3_chunk_2007_1_2065634 · in -33.4 dBFS · gain +13.4 dB · snippets-00066
(sadness, pain, helplessness · measured, subdued, no disfluency, whispered) very cognizant of losing the person she once was.
full caption & clip details
An elderly feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly warm, dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sadness, pain, helplessness; style: whispered, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.2/10; 3.7s.
batch112_part3_batch112_part3_chunk_2007_1_2065646 · in -35.6 dBFS · gain +15.6 dB · snippets-00066
(emotional numbness · normal-paced, normally alert, little disfluency, casual) One person, actually just one idea, can start a war or end one.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: casual, storytelling; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.5/10; 4.6s.
batch112_part3_batch112_part3_chunk_2007_1_2065721 · in -35.1 dBFS · gain +15.1 dB · snippets-00066
Pride ↓  /  Astonishment Surprisec-snippets-PXR · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.35.

At the same time Pride goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.10 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · snippets

k 3d_a -0.355d_b 0.348step_a 0.191step_b 0.222min_cos_consec 0.1011min_cos_anchor 0.1011dataset snippetslang ?speaker batch45_part2_batch45_parttrack batch45_part2_batch45_parttotal 14.4slevel spread 0.3 dBmax seam 0.2 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, normal-paced, light breath
(normally alert, slightly relaxed, fairly steady, formal) his hard man reputation and formidable record.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.9/10; 3.0s.
batch45_part2_batch45_part2_chunk_1404_1_1787228 · in -22.3 dBFS · gain +2.3 dB · snippets-01122
(helplessness, disappointment, confusion · energised, neutral tension, moderately variable, casual) Nobody wanted to give it to me, so I just went by my head and did what I thought was the best thing for me to do.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as helplessness, disappointment, confusion; style: casual, storytelling; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.5/10; 6.2s.
batch45_part2_batch45_part2_chunk_1404_1_1787270 · in -22.1 dBFS · gain +2.1 dB · snippets-01122
(normally alert, slightly relaxed, fairly steady, narration) from the first bell, the crowd expected Liston to march over and knock Clay out.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.0/10; 4.9s.
batch45_part2_batch45_part2_chunk_1404_1_1787296 · in -22.0 dBFS · gain +2.0 dB · snippets-01122