k-PXR-k3

PXR at chain length k=3, all corpora, at the mining floor.

Rule. PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable
Source. trajectories_v5.parquet  |  Family. rule x chain length
Sampled from 169,361 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Disappointment ↓  /  Doubtk-PXR-k3 · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Doubt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Doubt around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.44.

At the same time Disappointment goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · en · emolia

hear it un-normalised (raw levels, max seam 1.7 dB)
k 3d_a -0.445d_b 0.435step_a 0.445step_b 0.234min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_wAzOcyBntpItrack EN_wAzOcyBntpItotal 26.5slevel spread 1.7 dBmax seam 1.7 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, moderate pitch range, light breath
(normal-paced, steady, little disfluency, monologue) English proficiency in recent immigrants came up repeatedly as among the most vulnerable when it comes to floods. Okay?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.0/10; 6.9s, EN.
EN_wAzOcyBntpI_W000115 · in -19.4 dBFS · gain -0.6 dB · emolia-02337
(concentration, contemplation · normal-paced, fairly steady, some disfluency, didactic) So (ahem) uhm, keep this (ahem) snapshot in your mind, because I'm going to come back to this. Like, we have a decent understanding of who, what types of people tend to be more vulnerable, or what types of populations tend to be more vulnerable. So (ahem) what can we do with that understanding?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, contemplation; style: didactic, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.3/10; 14.9s, EN.
EN_wAzOcyBntpI_W000116 · in -20.6 dBFS · gain +0.6 dB · emolia-02337
(doubt, contemplation · measured, fairly steady, some disfluency, casual) But (ahem) pause here, why should we even focus on the vulnerable? What's the point?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, contemplation; style: casual, monologue; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 0.3/10; 4.4s, EN.
EN_wAzOcyBntpI_W000117 · in -18.9 dBFS · gain -1.1 dB · emolia-02337
Doubt ↓  /  Reliefk-PXR-k3 · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.28.

At the same time Doubt goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.20 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · ko · emolia

hear it un-normalised (raw levels, max seam 1.0 dB)
k 3d_a -0.321d_b 0.277step_a 0.220step_b 0.195min_cos_consec 0.9366min_cos_anchor 0.8998dataset emolialang kospeaker KO__y0x1iJs6pYtrack KO__y0x1iJs6pYtotal 39.2slevel spread 1.0 dBmax seam 1.0 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, measured, normally alert, slightly relaxed, fairly steady, some disfluency
(doubt · average clarity, moderate pitch range, monologue, formal) 그래서 집을 오히려 삽니다. 다 주택자들. 자, 지금 자료 화면에 한번 보시면.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: monologue, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.4/10; 6.3s, KO.
KO__y0x1iJs6pY_W000049 · in -17.4 dBFS · gain -2.5 dB · emolia-03172
(pride · somewhat unclear, fairly narrow pitch, didactic, monologue) (low mumble) 지난해 2022년 1월에서 12월 집합 건물 다 소유 지수입니다. 그러니까 다 주택자들이 집을 얼마나 늘려 가고 있냐, 아니면 줄여 가고 있냐, 이 지수인데, 지금 그림에 보고 있는 것처럼 지난해 가을, 그죠?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: didactic, monologue; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 2.9/10; 18.9s, KO.
KO__y0x1iJs6pY_W000050 · in -17.4 dBFS · gain -2.6 dB · emolia-03172
(relief · somewhat unclear, fairly narrow pitch, didactic, monologue) 봄, 지나면서, 가을, 겨울로 오면서 계속 다 소유지수를 크게 높아집니다, 그죠? 자, 이것은 무슨 말입니까? 부동산 정책이 확 뒤집어졌고, 그전에
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: didactic, monologue; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.1/10; 13.7s, KO.
KO__y0x1iJs6pY_W000051 · in -18.4 dBFS · gain -1.6 dB · emolia-03172
Hope Enthusiasm Optimism ↓  /  Sournessk-PXR-k3 · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sourness clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.29.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 46 s · hr · eurospeech

hear it un-normalised (raw levels, max seam 0.8 dB)
k 3d_a -0.398d_b 0.289step_a 0.232step_b 0.232min_cos_consec 0.9456min_cos_anchor 0.9570dataset eurospeechlang hrspeaker croatia_20070704161554-554track croatia_20070704161554-554total 45.9slevel spread 1.1 dBmax seam 0.8 dBcos from orange-id (speaker identity)
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child somewhat masculine voice · slightly cool, fairly smooth, average recording, quiet background, moderately variable, some disfluency
(fast, energised, neutral tension, cartoonish) kakva je situacija trenutna sa samim zatvorenicima i osobama koje su tamo zaposlene? Bilo bi zaista interesantno saznati koliko je u ove tri i pol godine
full caption & clip details
A child somewhat masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; no dominant emotion; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.5/10; 13.2s, HR.
croatia_20070704161554-5545_4385568_4398752 · in -24.5 dBFS · gain +4.5 dB · eurospeech-01230
(shame, disgust, pride · brisk, normally alert, neutral tension, dramatic) zatvorenika izvršilo suicid, koliko je umrlo u zatvorskoj čeliji zbog predoziranja, štrajkova glađu, tučnjave zatvorenika. Prekapacitiranost ovdje je bila (low mumble) višestruko (low mumble) istaknuta. Nepopunjena radna mjesta, nedostatak opreme i financijskih sredstava
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, disgust, pride; style: dramatic, cartoonish; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 5.7/10; 17.4s, HR.
croatia_20070704161554-5545_4398752_4416112 · in -24.8 dBFS · gain +4.8 dB · eurospeech-01230
(sourness, contempt · measured, energised, slightly relaxed, didactic) to karakterizira zatvorski sustav danas. To je jedna loša (ahem) situacija koju, gdje je upravitelj i gdje je ravnatelj sa osobljem nemoćan u tom prilivu osuđenika.
full caption & clip details
A child feminine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, contempt; style: didactic, dramatic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.2/10; 15.0s, HR.
croatia_20070704161554-5545_4416112_4431120 · in -25.6 dBFS · gain +5.6 dB · eurospeech-01230
Concentration ↓  /  Impatience and Irritabilityk-PXR-k3 · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Impatience and Irritability around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.40.

At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · en · emolia

hear it un-normalised (raw levels, max seam 1.4 dB)
k 3d_a -0.474d_b 0.401step_a 0.241step_b 0.227min_cos_consec 0.9109min_cos_anchor 0.8707dataset emolialang enspeaker EN_B00058_S02612track EN_B00058_S02612total 38.9slevel spread 1.4 dBmax seam 1.4 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · fairly smooth, good recording, moderately variable, wide pitch range, light breath
(concentration · normal-paced, normally alert, slightly relaxed, casual) Especially in terms of (ahem) property ownership and real estate development, because real estate development in low status communities in the U.S. currently takes one of two paths. First, poverty maintenance, wherein you'll see things like
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as concentration; style: casual, conversational; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.8/10; 15.8s, EN.
EN_B00058_S02612_W000019 · in -17.6 dBFS · gain -2.4 dB · emolia-01357
(contempt, interest · brisk, energised, neutral tension, dramatic) An authentic cultural attribute of the area and thus you'll see all sorts of programs that are designed to sort of manage that, that poverty, but the communities do not improve. Money will be made, lots of it, but not for the local people.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as contempt, interest; style: dramatic, authoritative; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.1/10; 15.1s, EN.
EN_B00058_S02612_W000020 · in -17.0 dBFS · gain -3.0 dB · emolia-01357
(impatience and irritability, distress, contempt · normal-paced, normally alert, neutral tension, dramatic) Because the promising talent is encouraged to grow up and be somebody, but certainly not in their own hood.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, distress, contempt; style: dramatic, authoritative; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.2/10; 7.7s, EN.
EN_B00058_S02612_W000021 · in -18.4 dBFS · gain -1.6 dB · emolia-01357
Affection ↓  /  Contentmentk-PXR-k3 · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.25.

At the same time Affection goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · en · podcast

hear it un-normalised (raw levels, max seam 1.8 dB)
k 3d_a -0.324d_b 0.252step_a 0.187step_b 0.132min_cos_consec 0.8078min_cos_anchor 0.7867dataset podcastlang enspeaker 338960track 338960total 32.7slevel spread 2.3 dBmax seam 1.8 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, normal-paced, normally alert, some disfluency, average clarity, moderate pitch range
(affection, contemplation · relaxed, fairly steady, casual, conversational) But you know, you you focus this year on changing your relationship with alcohol,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as affection, contemplation; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 4.7/6; vocal-burst blend 2.1/10; 3.6s, EN.
338960_00114215 · in -17.1 dBFS · gain -2.9 dB · podcast-04485
(interest, hope enthusiasm optimism · neutral tension, moderately variable, casual, conversational) this isn't about being perfect, this is about progress, and there's huge progress here. So the question with that fitness result would be what would it be if you didn't smoke? Of course it would be higher, right? Because that smoke binds to a bit of the the the oxygen traveling around your bloodstream. It limits you,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism; style: casual, conversational; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 6.6/10; 15.6s, EN.
338960_00114911 · in -18.9 dBFS · gain -1.1 dB · podcast-01688
(contentment, pleasure ecstasy, teasing · neutral tension, moderately variable, casual, conversational) but equally you plus smoking equals a good fitness result. Now that fitness result becomes slightly relevant when we look at the two exercise sessions we caught. Now, you probably thought you only gave me one exercise session. I did. But
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, pleasure ecstasy, teasing; style: casual, conversational; good recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.9/10; 13.3s, EN.
338960_00116464 · in -19.4 dBFS · gain -0.6 dB · podcast-03173
Interest ↓  /  Confusionk-PXR-k3 · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.28.

At the same time Interest goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.05 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · zh · emolia

k 3d_a -0.262d_b 0.285step_a 0.223step_b 0.233min_cos_consec 0.9350min_cos_anchor 0.9350dataset emolialang zhspeaker ZH_B00078_S03809track ZH_B00078_S03809total 35.7slevel spread 1.4 dBmax seam 1.4 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, average recording, normally alert, slightly relaxed, some disfluency, moderate pitch range, light breath
(normal-paced, fairly steady, average clarity, monologue) 那部片子全程基本上就没停过,一直在跑,就发生在半天一天里边的故事吧。男主角也是怒火一直在积压,然后越来越强烈,越来越强烈,把影片的节奏往上提。我本来想看受过愤怒的海里边的黄渤。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 6.6/10; 14.8s, ZH.
ZH_B00078_S03809_W000367 · in -16.9 dBFS · gain -3.1 dB · emolia-04061
(impatience and irritability, sexual lust, anger · fast, moderately variable, average clarity, casual) 也是这样,就是愤怒值一直往上顶,一直往上顶,一直往上顶,然后到最后可能抓到李苗苗那一刻的时候,有这样巨大的抒发,疯狂的抒发。但是在这篇子子里边。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, sexual lust, anger; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 8.0/10; 11.8s, ZH.
ZH_B00078_S03809_W000368 · in -15.5 dBFS · gain -4.5 dB · emolia-04061
(confusion, contempt · fast, fairly steady, clear, storytelling) 他还想做成张弛有度,反而这个所谓的持的时候呢,就会把好不容易之前建立起来的愤怒感给往下落。
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, contempt; style: storytelling, dramatic; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 4.8/10; 8.8s, ZH.
ZH_B00078_S03809_W000369 · in -15.9 dBFS · gain -4.1 dB · emolia-04061
Emotional Numbness ↓  /  Interestk-PXR-k3 · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.33.

At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.98 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.98 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.98), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 52 s · en · emolia

k 3d_a -0.400d_b 0.328step_a 0.202step_b 0.170min_cos_consec 0.9848min_cos_anchor 0.9796dataset emolialang enspeaker EN_MYvnH6XcwaAtrack EN_MYvnH6XcwaAtotal 52.2slevel spread 0.4 dBmax seam 0.4 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness, contempt, sourness · newsreading, authoritative) One alternative spectrum offered by the conservative American Federalist Journal accounts for only the «degree of government control» without consideration for any other social or political variable and thus places «fascism», totalitarianism, at one extreme and «anarchism», no government at all, at the other extreme.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, contempt, sourness; style: newsreading, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.7s, EN.
EN_MYvnH6XcwaA_W000131 · in -15.4 dBFS · gain -4.6 dB · emolia-00870
(concentration, disgust · newsreading, formal) The Vosum chart, or Vosum cube, is based on the Nolan chart and adds a third axis for corporate issues, depicted three-dimensionally, with eight discrete categories representing eight different political ideologies. Vosum is the Russian word for
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disgust; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.4s, EN.
EN_MYvnH6XcwaA_W000132 · in -15.7 dBFS · gain -4.3 dB · emolia-00870
(interest, concentration · newsreading, formal) In 1998, political author Virginia Postral, in her book The Future and Its Enemies, offered another single-axis spectrum that measures views of the future, contrasting stacists, who allegedly fear the future and wish to control it, and dynamists, who want the future to unfold naturally and without attempts to plan and control.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.8s, EN.
EN_MYvnH6XcwaA_W000135 · in -15.3 dBFS · gain -4.7 dB · emolia-00870
Hope Enthusiasm Optimism ↓  /  Emotional Numbnessk-PXR-k3 · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it clearly present at 0.70, higher than 70 % of clips in this corpus. That is a total rise of 0.38.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.43 (lower than 57 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.13 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 16 s · en · emolia

k 3d_a -0.380d_b 0.376step_a 0.239step_b 0.245min_cos_consec 0.8007min_cos_anchor 0.7743dataset emolialang enspeaker EN_B00024_S06580track EN_B00024_S06580total 15.5slevel spread 1.0 dBmax seam 1.0 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(little disfluency, monologue, authoritative) Okay, let's get into the reading for plant adaptations.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, authoritative; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.1/10; 4.4s, EN.
EN_B00024_S06580_W000105 · in -21.4 dBFS · gain +1.4 dB · emolia-00706
(relief · little disfluency, conversational, authoritative) As I read, read along with me. You can read out loud or in your head, practice the pronunciation. Here we go.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as relief; style: conversational, authoritative; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.5/10; 6.7s, EN.
EN_B00024_S06580_W000106 · in -22.5 dBFS · gain +2.5 dB · emolia-00706
(no disfluency, monologue, formal) Plants have adapted to live in their habitats.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.0/10; 4.2s, EN.
EN_B00024_S06580_W000107 · in -22.0 dBFS · gain +2.0 dB · emolia-00706
Astonishment Surprise ↓  /  Hope Enthusiasm Optimismk-PXR-k3 · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Hope Enthusiasm Optimism around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.43.

At the same time Astonishment Surprise goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.32 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.32, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · en · podcast

k 3d_a -0.328d_b 0.428step_a 0.229step_b 0.217min_cos_consec 0.0223min_cos_anchor 0.3169dataset podcastlang enspeaker 645619track 645619total 25.9slevel spread 5.8 dBmax seam 5.8 dBcos from orange-id (speaker identity)
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, quiet background
(astonishment surprise, amusement, infatuation · normal-paced, normally alert, neutral tension, casual) And (ahem) he and then he messaged me right after. He was like, You're welcome for not selling. And I was like, You literally just need to stay alive and the the guy
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, amusement, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.8/10; 7.8s, EN.
645619_00246336 · in -33.5 dBFS · gain +13.6 dB · podcast-04128
(shame, embarrassment, intoxication altered states of consciousness · measured, very low-energy, relaxed, casual) But it's whatever. It's whatever. Hey, (low mumble) uh, Green, whatever your name is, I'm gonna drop this podcast link in your Xbox thing. I hope you listen to it. Fuck you.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, submissive, neutral openness; reads as shame, embarrassment, intoxication altered states of consciousness; style: casual, monologue; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.1/10; 11.0s, EN.
645619_00247264 · in -36.4 dBFS · gain +16.4 dB · podcast-03348
(hope enthusiasm optimism, elation, thankfulness gratitude · measured, normally alert, neutral tension, casual) Hey, I I will make (low mumble) uh I will say this right now. Cause I know the part the segment in this podcast coming up where we're all getting taste.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as hope enthusiasm optimism, elation, thankfulness gratitude; style: casual, monologue; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.0/10; 6.8s, EN.
645619_00249952 · in -30.6 dBFS · gain +10.6 dB · podcast-04122
Emotional Numbness ↓  /  Concentrationk-PXR-k3 · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.26.

At the same time Emotional Numbness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.05 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · en · emolia

k 3d_a -0.341d_b 0.261step_a 0.225step_b 0.214min_cos_consec 0.9600min_cos_anchor 0.9600dataset emolialang enspeaker EN_n4Hf6ARx8HYtrack EN_n4Hf6ARx8HYtotal 27.8slevel spread 2.1 dBmax seam 1.8 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, steady
(emotional numbness · slow, formal, monologue) Control over popular dissent — and, rationalization
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, very full; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 4.5s, EN.
EN_n4Hf6ARx8HY_W000158 · in -13.5 dBFS · gain -6.5 dB · emolia-02183
(disgust, malevolence malice · normal-paced, newsreading, formal) He further attacks within-system green initiatives like carbon trading, which he sees as a
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, malevolence malice; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.0s, EN.
EN_n4Hf6ARx8HY_W000159 · in -15.4 dBFS · gain -4.6 dB · emolia-02183
(normal-paced, newsreading, formal) Brian Tokar has further criticized carbon trading in this way, suggesting that it augments existing class inequality and gives the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 13.0s, EN.
EN_n4Hf6ARx8HY_W000160 · in -15.7 dBFS · gain -4.3 dB · emolia-02183
Doubt ↓  /  Concentrationk-PXR-k3 · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.26.

At the same time Doubt goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.19 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · zh · emolia

k 3d_a -0.256d_b 0.256step_a 0.130step_b 0.193min_cos_consec 0.9340min_cos_anchor 0.9340dataset emolialang zhspeaker ZH_B00014_S01068track ZH_B00014_S01068total 31.1slevel spread 0.3 dBmax seam 0.3 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, no background noise, measured, normally alert
(doubt · fairly steady, moderate pitch range, monologue, narration) (low mumble) 张飞摇摇头,满是无奈的说道,也不知道是哪个王八蛋,竟然将南边的毛巾给毁了个大半啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: monologue, narration; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.1/10; 9.1s, ZH.
ZH_B00014_S01068_W000029 · in -16.7 dBFS · gain -3.3 dB · emolia-03420
(steady, moderate pitch range, monologue, formal) 陈登一惊,旋即就冷静了下来。分析到两种可能,第一,曹操早就做好了丢掉河东的准备。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.8/10; 9.1s, ZH.
ZH_B00014_S01068_W000030 · in -16.7 dBFS · gain -3.3 dB · emolia-03420
(steady, narrow pitch range, monologue, formal) 所以,在看到情势危急的情况下,干脆毁了毛巾,既能挑起我们的怒火,也能拖延一定的时间,这样恰好就能解释大洋城为何如此的破。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.5/10; 12.6s, ZH.
ZH_B00014_S01068_W000031 · in -16.4 dBFS · gain -3.6 dB · emolia-03420
Fear ↓  /  Fatigue Exhaustionk-PXR-k3 · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion around average — 0.50, right about the corpus median — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.27.

At the same time Fear goes the other way, from 0.75 (higher than 75 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.07, then +0.20 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 12 s · zh · emolia

k 3d_a -0.255d_b 0.272step_a 0.184step_b 0.197min_cos_consec 0.8238min_cos_anchor 0.7653dataset emolialang zhspeaker ZH_B00065_S02960track ZH_B00065_S02960total 12.0slevel spread 2.8 dBmax seam 2.8 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(measured, steady, formal, monologue) 张常里对蔡元培此时支持请愿团的行为很是烦恼。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.6/10; 5.1s, ZH.
ZH_B00065_S02960_W000004 · in -22.9 dBFS · gain +2.9 dB · emolia-03922
(pain · normal-paced, fairly steady, formal, monologue) 本来,国会已经接受了对蔡元培的罢免提案。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.5/10; 3.4s, ZH.
ZH_B00065_S02960_W000005 · in -20.1 dBFS · gain +0.1 dB · emolia-03922
(normal-paced, fairly steady, conversational) 但眼下,国民爱国之情十分高涨。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.9/10; 3.2s, ZH.
ZH_B00065_S02960_W000006 · in -22.4 dBFS · gain +2.5 dB · emolia-03922
Relief ↓  /  Concentrationk-PXR-k3 · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.32.

At the same time Relief goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · hr · eurospeech

k 3d_a -0.266d_b 0.320step_a 0.195step_b 0.206min_cos_consec 0.9070min_cos_anchor 0.9108dataset eurospeechlang hrspeaker croatia_20150916095019-151track croatia_20150916095019-151total 40.6slevel spread 1.7 dBmax seam 1.7 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, some disfluency, somewhat unclear
(relief, thankfulness gratitude, pride · measured, normally alert, slightly relaxed, monologue) HŽ Putnički se priprema za liberalizaciju tržišta, on je većinskom vlasništvu države. Njegov plan restrukturiranja je pred Europskom komisijom.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as relief, thankfulness gratitude, pride; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 11.1s, HR.
croatia_20150916095019-15190_7534016_7545088 · in -18.8 dBFS · gain -1.2 dB · eurospeech-01441
(thankfulness gratitude, contentment, longing · normal-paced, normally alert, neutral tension, monologue) Liberalizacija će doći vrlo brzo, doći će vjerojatno 2019. godine, a vjerojatno će i kako je stvar restrukturiranja i neki drugi prijevoznici morati ili moći voziti puno ranije. Stoga je upravo nabava novih vlakova koji većinom voze na linijama Sisak
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment, longing; style: monologue, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 7.0/10; 16.1s, HR.
croatia_20150916095019-15190_7545088_7561200 · in -17.1 dBFS · gain -2.9 dB · eurospeech-01441
(concentration · measured, subdued, slightly relaxed, monologue) Karlovac – Koprivnica zapravo jedan jedini odgovor budućoj liberalizaciji i jačanju usluga što prati ulaganje u infrastrukturu, pa bar će za nekih
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, authoritative; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.1/10; 13.1s, HR.
croatia_20150916095019-15190_7561200_7574272 · in -18.2 dBFS · gain -1.8 dB · eurospeech-01441
Concentration ↓  /  Sexual Lustk-PXR-k3 · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.35.

At the same time Concentration goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · zh · emolia

k 3d_a -0.309d_b 0.353step_a 0.213step_b 0.249min_cos_consec 0.9123min_cos_anchor 0.9058dataset emolialang zhspeaker ZH_B00003_S03583track ZH_B00003_S03583total 36.0slevel spread 1.2 dBmax seam 1.2 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(concentration · measured, some disfluency, whispered, monologue) 每次按三秒重复十次左右,血液循环就会有显著的改善全身疲劳,总觉得浑身没劲儿。最近工作太辛苦了等等。
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as concentration; style: whispered, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.0/10; 12.6s, ZH.
ZH_B00003_S03583_W000015 · in -20.5 dBFS · gain +0.5 dB · emolia-03305
(thankfulness gratitude, infatuation · normal-paced, some disfluency, monologue, narration) 有两个穴位很适合这种情况,不光能促进血液循环,加快新陈代谢,还有调节肠胃的效果,经常按压全身的循环会有显著的改善。想睡个好觉的时候也可以刺激这两个穴位哦。第一,手肘附近。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, infatuation; style: monologue, narration; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.5/10; 17.0s, ZH.
ZH_B00003_S03583_W000016 · in -20.2 dBFS · gain +0.2 dB · emolia-03305
(sexual lust · measured, no disfluency, monologue, narration) 将四根手指并排放在膝盖外侧的凹陷处,靠近脚底的一侧。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 6.1s, ZH.
ZH_B00003_S03583_W000017 · in -21.4 dBFS · gain +1.4 dB · emolia-03305
Pride ↓  /  Shamek-PXR-k3 · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.40.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · ko · emolia

k 3d_a -0.273d_b 0.404step_a 0.194step_b 0.224min_cos_consec 0.9099min_cos_anchor 0.9202dataset emolialang kospeaker KO_YvywE52bBWutrack KO_YvywE52bBWutotal 39.6slevel spread 1.1 dBmax seam 1.1 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed
(pride, disgust, concentration · normal-paced, fairly steady, almost no disfluency, monologue) 그 다음에 삼보, 그 세번째로 보면은요, 보감어기의 사역자에게는 감추었던 만나와 새 이름이 기록된 힌돌을 주겠다. 이렇게 하죠. 힌돌은 예수님 의미 합니다. 그러니까 예수님과 연합하게 주겠다. 예수님과 하나가 되게 해주겠다. 니가 내 안에, 내가 니 안에 있게 해주겠다. 그런 뜻이죠.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, disgust, concentration; style: monologue, authoritative; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 4.3/10; 17.1s, KO.
KO_YvywE52bBWu_W000023 · in -13.2 dBFS · gain -6.8 dB · emolia-03178
(interest · brisk, moderately variable, some disfluency, authoritative) 그 다음에, 두하디라교의 사육자에게는 만국을 다스리는 권서와 철장 권서와 새벽 뼈를 주겠다. 여러분, 새벽자는 지저스 모닝 스타 아닙니까? 예, 그, 지금으로 하면 금성이죠. 이것은 바로 성경에서 예수님의 의미입니다. 예수님.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest; style: authoritative, dramatic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 5.7/10; 14.7s, KO.
KO_YvywE52bBWu_W000024 · in -13.2 dBFS · gain -6.8 dB · emolia-03178
(shame · brisk, moderately variable, some disfluency, conversational) 그게 아니고, 바로 예수님의 의미였죠. 그러니까 예수님과 하나 되게 해주겠다. 그런 뜻이에요. 사대교의 사역자는 흰 옷을 입게 해주겠다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame; style: conversational, dramatic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 5.5/10; 7.5s, KO.
KO_YvywE52bBWu_W000025 · in -14.3 dBFS · gain -5.7 dB · emolia-03178
Hope Enthusiasm Optimism ↓  /  Embarrassmentk-PXR-k3 · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Embarrassment clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.30.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.34 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.36 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.34, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · podcast

k 3d_a -0.319d_b 0.302step_a 0.197step_b 0.222min_cos_consec 0.3625min_cos_anchor 0.3431dataset podcastlang enspeaker 145561track 145561total 30.1slevel spread 1.5 dBmax seam 1.5 dBcos from orange-id (speaker identity)
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned
(hope enthusiasm optimism, amusement, teasing · normal-paced, normally alert, slightly relaxed, casual) advertising that you will one day laugh again, like
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, amusement, teasing; style: casual, conversational; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 3.0/10; 3.0s, EN.
145561_00317976 · in -23.2 dBFS · gain +3.2 dB · podcast-02308
(fatigue exhaustion, elation, helplessness · normal-paced, normally alert, neutral tension, casual) I have to I can only go to the store once a week now instead of like four times a week. This uh this sh this has changed my life completely.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, elation, helplessness; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.4/10; 8.4s, EN.
145561_00318592 · in -24.8 dBFS · gain +4.8 dB · podcast-02312
(embarrassment, amusement, sexual lust · measured, energised, neutral tension, casual) Just like sitting at home with my head in my hand Aaron Pauling, just like head in my (exhausted groan) hand. (cackle) (surprised gasp) You know, I never thought I'd be someone who had to buy in bulk,
full caption & clip details
A young adult masculine voice; delivery is energised, measured, neutral tension, volatile; timbre is neutral-toned, very dark, very rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is elated, neutral stance, neutral openness; reads as embarrassment, amusement, sexual lust; style: casual, playful; below-average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 0.3/10; 18.4s, EN.
145561_00319920 · in -23.8 dBFS · gain +3.8 dB · podcast-02307
Longing ↓  /  Jealousy and Envyk-PXR-k3 · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Jealousy and Envy clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.

At the same time Longing goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · en · emolia

k 3d_a -0.333d_b 0.336step_a 0.204step_b 0.240min_cos_consec 0.8822min_cos_anchor 0.8579dataset emolialang enspeaker EN_B00029_S05784track EN_B00029_S05784total 40.0slevel spread 2.9 dBmax seam 2.2 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, good recording, no background noise, measured, very low-energy, clear, wide pitch range
(longing, emotional numbness, fear · slightly relaxed, moderately variable, almost no disfluency, narration) Then he and Scout walk down the aisle to where I'm sitting. Mr. Carlson sits in the desk next to me. Scout lies down in the aisle between us.
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as longing, emotional numbness, fear; style: narration, storytelling; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.5/10; 10.4s, EN.
EN_B00029_S05784_W000025 · in -19.7 dBFS · gain -0.3 dB · emolia-00809
(teasing, affection, thankfulness gratitude · neutral tension, moderately variable, almost no disfluency, narration) Good boy, Mr. Carlson says, ruffling the fur on the dog's head. He's doing a good job of praising Scout, but I don't feel like telling him that. We need to talk, he says.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, wide pitch range, normal breath; affect is negative, neutral stance, neutral openness; reads as teasing, affection, thankfulness gratitude; style: narration, storytelling; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.8/10; 14.1s, EN.
EN_B00029_S05784_W000026 · in -21.9 dBFS · gain +1.9 dB · emolia-00809
(jealousy and envy, pain, confusion · neutral tension, variable, little disfluency, conversational) Yeah, I answer. I pick at a hangnail on my left thumb. It's not just the quiz, he continues. You didn't take any notes in class today. How did you know? I exclaim.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, neutral tension, variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, pain, confusion; style: conversational, storytelling; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.0/10; 15.3s, EN.
EN_B00029_S05784_W000027 · in -22.6 dBFS · gain +2.6 dB · emolia-00809
Doubt ↓  /  Concentrationk-PXR-k3 · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.

At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.58 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.58 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.58, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 83 s · en · podcast

k 3d_a -0.272d_b 0.258step_a 0.190step_b 0.179min_cos_consec 0.5798min_cos_anchor 0.5798dataset podcastlang enspeaker 959119track 959119total 82.7slevel spread 2.7 dBmax seam 2.7 dBcos from orange-id (speaker identity)
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, slightly relaxed, fairly steady
(doubt, jealousy and envy, hope enthusiasm optimism · subdued, some disfluency, monologue, formal) (ahem) Uh yeah, the multiple formats seems really promising for usability, Steve. And that seems like a way to get beyond the status quo to get some better constructed (ahem) um input from the public. (ahem) And now, Sarah, of course, this kind of work is gonna take a lot of outreach. (ahem) Um, what kind of engagement opportunities are we seeing OIRA offer and what stands out to you most about the learnings that they've collected so far?
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, jealousy and envy, hope enthusiasm optimism; style: monologue, formal; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 2.5/10; 28.2s, EN.
959119_00045616 · in -29.9 dBFS · gain +9.9 dB · podcast-04905
(astonishment surprise · subdued, frequent disfluency, monologue, formal) And Steve, uh (ahem) among the recommendations, you had pointed out one potential pitfall. (ahem) Uh, what is it that could lead participants to be less than satisfied with these (ahem) engagements?
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise; style: monologue, formal; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.1/10; 27.7s, EN.
959119_00057351 · in -27.2 dBFS · gain +7.2 dB · podcast-04896
(concentration · normally alert, some disfluency, casual, monologue) Yeah, I see what you mean. I see what you mean. In in any participatory process, there's going to be outcomes that (low mumble) uh didn't go the way that you wanted to, but the important thing is making sure that the input is actionable and relevant and that it's collected at the right time and that can help the agency move on its goals.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.3/10; 26.4s, EN.
959119_00089355 · in -27.4 dBFS · gain +7.4 dB · podcast-04916
Confusion ↓  /  Longingk-PXR-k3 · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Longing clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.27.

At the same time Confusion goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · zh · emolia

k 3d_a -0.268d_b 0.265step_a 0.145step_b 0.146min_cos_consec 0.9317min_cos_anchor 0.9210dataset emolialang zhspeaker ZH_B00045_S02498track ZH_B00045_S02498total 32.6slevel spread 1.3 dBmax seam 0.7 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, fairly smooth, no background noise, normally alert, slightly relaxed, no disfluency, light breath
(confusion, fatigue exhaustion · slow, steady, crisply articulate, whispered) 在这里,一即自身直接拥有的东西,才具备绝对的价值。因而。
full caption & clip details
A child feminine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, very dark, fairly smooth, thin; crisply articulate, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, fatigue exhaustion; style: whispered, ASMR; average recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.0/10; 7.1s, ZH.
ZH_B00045_S02498_W000003 · in -17.2 dBFS · gain -2.8 dB · emolia-03730
(thankfulness gratitude, contentment · measured, fairly steady, clear, whispered) 的确,名声只是某种的外在显示,名人以此证实了自己对自己抱有的高度的评价并没有错。因此,人们可以说,正如光本身是看不见的,除非它经过物体的折射。同样。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment; style: whispered, monologue; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.1/10; 18.9s, ZH.
ZH_B00045_S02498_W000004 · in -16.5 dBFS · gain -3.5 dB · emolia-03730
(longing · measured, steady, clear, monologue) 一个人所具有的卓越之处,只是通过获得名声才变得无可争议。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: monologue, whispered; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 6.3s, ZH.
ZH_B00045_S02498_W000005 · in -16.0 dBFS · gain -4.0 dB · emolia-03730
Infatuation ↓  /  Feark-PXR-k3 · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.41.

At the same time Infatuation goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · en · emolia

k 3d_a -0.335d_b 0.414step_a 0.196step_b 0.230min_cos_consec 0.9296min_cos_anchor 0.9510dataset emolialang enspeaker EN_aQOymfwMIrutrack EN_aQOymfwMIrutotal 22.8slevel spread 0.7 dBmax seam 0.7 dBcos from recomputed from spkemb_traj
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, authoritative) A gas giant planet in the same system as New California, home to 30 Edenist habitats
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.3/10; 6.2s, EN.
EN_aQOymfwMIru_W000049 · in -17.0 dBFS · gain -3.0 dB · emolia-00454
(formal, authoritative) The Capone organization attempted to take control of Yosemite after capturing New California
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.1/10; 6.4s, EN.
EN_aQOymfwMIru_W000050 · in -17.1 dBFS · gain -2.9 dB · emolia-00454
(fear · formal, newsreading) Unknown to Capone, the Yosemite habitats had turned a large part of their massive industrial capacity over to military production when the threat became apparent
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.9s, EN.
EN_aQOymfwMIru_W000051 · in -17.7 dBFS · gain -2.3 dB · emolia-00454