proxy_taillift__PXR__T0.60__C0.25__INTERNAL

Manifest tier. proxy_taillift, rule PXR, T=0.6, step cap 0.25. Population 71 chains (1 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 67.

Rule. PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable
Source. trajectories_v5lift.parquet  |  Family. the tier owner's manifest tiers -- the exact subsets used for training
Manifest tier. proxy_taillift__PXR__T0.60__C0.25__INTERNAL — population 71 chains (1 h). SHAREABLE variant: 67.
Filter. rule=='PXR' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and qmax>=0.6 and cmax<=0.25
This tier was resampled. The manifest's earlier published filter applied the one-sided test to every tier regardless of its actual rule. Tier populations were always counted under the correct rule, so the sizes quoted here were never wrong — but the earlier selection drew from a contaminated pool, of which 98.4 % were not members of this tier. These samples are drawn with the corrected filter, which tests qmax >= T and cmax <= C — per-row quantities computed under this tier's own rule. For a proxy tier no d_a/d_b predicate could be correct, because a proxy-rescued chain may legitimately exceed the cap on the named axis while its smoothness is certified on the proxy axis.
Sampled from 71 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Intoxication Altered States of Consciousness ↓  /  Emotional Numbnessproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.71.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.69 (higher than 69 % of clips in this corpus) to 0.08 (lower than 92 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.10, then +0.22, then +0.14, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 41 s · en · emolia

hear it un-normalised (raw levels, max seam 1.2 dB)
k 5d_a -0.606d_b 0.705step_a 0.237step_b 0.244min_cos_consec 0.8652min_cos_anchor 0.8652dataset emolialang enspeaker EN_SKjE_MI17TAtrack EN_SKjE_MI17TAtotal 40.5slevel spread 1.6 dBmax seam 1.2 dBqmax 0.606cmax 0.244cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, fairly steady, moderate pitch range, formal) Jam sandwiches are thought to have originated at around the 19th century in the United Kingdom.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.0/10; 5.2s, EN.
EN_SKjE_MI17TA_W000001 · in -15.0 dBFS · gain -5.0 dB · emolia-01762
(normal-paced, steady, moderate pitch range, formal) In Scotland, they are also known as pieces and jam, or geely pieces.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 4.1s, EN.
EN_SKjE_MI17TA_W000002 · in -15.5 dBFS · gain -4.5 dB · emolia-01762
(normal-paced, fairly steady, moderate pitch range, formal) The jam sandwich was an affordable food which was a major part of the diets of the lower, working class people of cities such as London and Glasgow.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 7.9s, EN.
EN_SKjE_MI17TA_W000003 · in -15.2 dBFS · gain -4.8 dB · emolia-01762
(normal-paced, fairly steady, moderate pitch range, newsreading) One plausible reason for this was that the ingredients that the jam sandwiches were made from cost little to manufacture and due to taxes being lifted on sugar in 1880, it became widely available as a cheap foodstuff.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.3s, EN.
EN_SKjE_MI17TA_W000004 · in -15.3 dBFS · gain -4.7 dB · emolia-01762
(measured, steady, fairly narrow pitch, formal) Today, jam sandwiches are mainly consumed by children. Shops do not often sell individual jam sandwiches. == Ingredients and nutrition ==
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 10.3s, EN.
EN_SKjE_MI17TA_W000005 · in -16.6 dBFS · gain -3.4 dB · emolia-01762
Pain ↓  /  Infatuationproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation barely there — 0.23, lower than 77 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.62.

At the same time Pain goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.25, then +0.20, then -0.05 — not a clean run: step 4 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 34 s · zh · emolia

hear it un-normalised (raw levels, max seam 1.5 dB)
k 5d_a -0.602d_b 0.617step_a 0.233step_b 0.248min_cos_consec 0.8813min_cos_anchor 0.9344dataset emolialang zhspeaker ZH_B00064_S03304track ZH_B00064_S03304total 33.8slevel spread 2.1 dBmax seam 1.5 dBqmax 0.602cmax 0.248cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(pain · monologue, formal) 只关心土豪书做的能量增幅,药水荒歌牌八十二年雪碧能不能用?而此时,科学家地中海告诉白毛书荒歌牌八十二年雪碧有副作用。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, formal; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.8/10; 11.8s, ZH.
ZH_B00064_S03304_W000008 · in -23.3 dBFS · gain +3.3 dB · emolia-03916
(monologue, narration) 会让人变得狂暴。白毛叔一听瞬间怒火攻了心,当即表示,如果不行的话,就不给土豪叔投钱了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.9/10; 8.0s, ZH.
ZH_B00064_S03304_W000009 · in -23.1 dBFS · gain +3.0 dB · emolia-03916
(formal, monologue) 为了拿到钱,土豪,叔决定拿自己做实验。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.8/10; 3.3s, ZH.
ZH_B00064_S03304_W000010 · in -22.0 dBFS · gain +2.0 dB · emolia-03916
(monologue, narration) 只见土豪叔喝了一瓶荒歌牌,八十二年雪碧,又闻了一阵浓烟之后。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.4/10; 5.7s, ZH.
ZH_B00064_S03304_W000011 · in -22.7 dBFS · gain +2.7 dB · emolia-03916
(monologue, formal) 土豪叔变得十分的狂暴,一耳巴子就把地中海给结了果。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.5/10; 4.4s, ZH.
ZH_B00064_S03304_W000012 · in -24.2 dBFS · gain +4.2 dB · emolia-03916
Disgust ↓  /  Concentrationproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.29, lower than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.65.

At the same time Disgust goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.07, then +0.17, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · zh · emolia

hear it un-normalised (raw levels, max seam 2.8 dB)
k 5d_a -0.605d_b 0.651step_a 0.222step_b 0.231min_cos_consec 0.7069min_cos_anchor 0.7938dataset emolialang zhspeaker ZH_B00036_S07907track ZH_B00036_S07907total 61.0slevel spread 4.5 dBmax seam 2.8 dBqmax 0.605cmax 0.231cos from recomputed from spkemb_traj
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, moderate pitch range, light breath
(disgust, intoxication altered states of consciousness · measured, slightly relaxed, fairly steady, conversational) 他他那城市老一般的男朋友is pretty glumsy,什么意思啊?笨拙的on the farm啊,在。 (ahem)
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, intoxication altered states of consciousness; style: conversational, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.1/10; 7.2s, ZH.
ZH_B00036_S07907_W000033 · in -21.2 dBFS · gain +1.2 dB · emolia-03639
(confusion, intoxication altered states of consciousness, interest · normal-paced, neutral tension, moderately variable, casual) 好,接下来再看一个look at that full city (ahem) sliker来看那个fool啊,那个傻瓜一样的city sliker怎么着或者好,yes, no idea how to see pegs啊,它是没有。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as confusion, intoxication altered states of consciousness, interest; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 7.3/10; 12.1s, ZH.
ZH_B00036_S07907_W000034 · in -18.4 dBFS · gain -1.6 dB · emolia-03639
(sexual lust, relief, intoxication altered states of consciousness · normal-paced, slightly relaxed, fairly steady, casual) (ahem) 办法去啊这个他也不知道该怎么去饲养这些啊呃猪的啊。好,那么接下来咱们来复习一下今天所讲的内容,总结一下heavy (ahem) (low mumble) (ahem) heater什么意思啊啊,大人物。 (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, relief, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.9/10; 12.8s, ZH.
ZH_B00036_S07907_W000035 · in -21.0 dBFS · gain +1.0 dB · emolia-03639
(measured, slightly relaxed, fairly steady, didactic) Keep a weether right out保持警惕。 City (ahem) (ahem) (ahem) (ahem) sldiger啊表示圆滑事故啊,表示不是特别可靠的这种城市里边的啊这种啊illy sldiger城市网。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.9/10; 11.8s, ZH.
ZH_B00036_S07907_W000036 · in -22.7 dBFS · gain +2.7 dB · emolia-03639
(concentration, interest, contemplation · brisk, slightly relaxed, fairly steady, monologue) (surprised gasp) 就是这样。好各位。那么如果大家呢想跟着我们一起练习刚才的重点句的话,那么欢迎大家进入到我们的啊练习社群之中来。那么入群方式呢,就是大家添加屏幕上的客服助角的微信,由他来去引导你入群好。那么今天咱们就先说到这儿,感谢大家,我们下次再见见。
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 6.3/10; 16.6s, ZH.
ZH_B00036_S07907_W000037 · in -22.8 dBFS · gain +2.8 dB · emolia-03639
Thankfulness Gratitude ↓  /  Concentrationproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.72.

At the same time Thankfulness Gratitude goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.67. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.19, then +0.20, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · en · emolia

hear it un-normalised (raw levels, max seam 3.5 dB)
k 5d_a -0.671d_b 0.719step_a 0.245step_b 0.215min_cos_consec 0.7852min_cos_anchor 0.7872dataset emolialang enspeaker EN_K0dRbVNBiF0track EN_K0dRbVNBiF0total 35.1slevel spread 3.5 dBmax seam 3.5 dBqmax 0.671cmax 0.245cos from recomputed from spkemb_traj
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, light breath
(thankfulness gratitude, contentment, relief · normal-paced, energised, neutral tension, casual) Yeah, thank you. Thank you, Cindy. Thank you for sharing your perspective and I can sense and understand your frustration, (low mumble) uh, as well.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as thankfulness gratitude, contentment, relief; style: casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.9/10; 7.8s, EN.
EN_K0dRbVNBiF0_W000194 · in -20.8 dBFS · gain +0.8 dB · emolia-00786
(normal-paced, normally alert, slightly relaxed, casual) And so (low mumble) uhm, grandstanding, (ahem) some people call it performative allyship.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.4/10; 4.8s, EN.
EN_K0dRbVNBiF0_W000195 · in -17.3 dBFS · gain -2.7 dB · emolia-00786
(normal-paced, normally alert, neutral tension, casual) Uh, (low mumble) virtual signaling, I mean, we're starting to see all sorts of different, different ways to describe it. But, (low mumble) uhm,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.5/10; 6.7s, EN.
EN_K0dRbVNBiF0_W000196 · in -20.2 dBFS · gain +0.2 dB · emolia-00786
(normal-paced, normally alert, slightly relaxed, conversational) Here's, (ahem) uh, you know, here's, here's what we have to understand when it, when it comes to grandstanding, when it comes to any of these perspectives, to be honest with you.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: conversational, casual; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.2/10; 7.7s, EN.
EN_K0dRbVNBiF0_W000197 · in -20.6 dBFS · gain +0.6 dB · emolia-00786
(concentration, interest · brisk, energised, neutral tension, authoritative) (low mumble) Uh, but I actually know that's, let's focus on grants, let's focus on the ability as to grandstanding. Cindy, you made a point that I think is really important.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, interest; style: authoritative, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.2/10; 7.5s, EN.
EN_K0dRbVNBiF0_W000198 · in -19.9 dBFS · gain -0.1 dB · emolia-00786
Concentration ↓  /  Jealousy and Envyproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Jealousy and Envy below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.65.

At the same time Concentration goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.23 (lower than 77 % of clips in this corpus), a change of -0.67. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.03 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 42 s · en · emolia

hear it un-normalised (raw levels, max seam 2.9 dB)
k 4d_a -0.667d_b 0.653step_a 0.232step_b 0.236min_cos_consec -0.0250min_cos_anchor 0.0221dataset emolialang enspeaker EN_xH1xo44ACnktrack EN_xH1xo44ACnktotal 42.5slevel spread 4.1 dBmax seam 2.9 dBqmax 0.653cmax 0.236cos from recomputed from spkemb_traj
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, light breath
(measured, subdued, slightly relaxed, monologue) Combine with larger wafer sizes. So (low mumble) uhm, so we, you know, we are talking about a process that
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.1/10; 8.9s, EN.
EN_xH1xo44ACnk_W000005 · in -20.9 dBFS · gain +0.9 dB · emolia-02623
(slow, normally alert, slightly relaxed, monologue) Essentially for 50 years it's been, you know, it's been driving this, (low mumble) uh, industry.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.2/10; 6.1s, EN.
EN_xH1xo44ACnk_W000006 · in -20.7 dBFS · gain +0.7 dB · emolia-02623
(intoxication altered states of consciousness, fatigue exhaustion, doubt · normal-paced, normally alert, neutral tension, conversational) Unfortunately, we are close to the end of this line. So, uh, (low mumble) you know, three nanometers. The next, uh, (ahem) the next step is, uh, (low mumble) you know, square root of three. So 1.7 nanometers. I don't know that anybody knows how to do that.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, doubt; style: conversational, casual; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 6.6/10; 16.5s, EN.
EN_xH1xo44ACnk_W000007 · in -19.7 dBFS · gain -0.3 dB · emolia-02623
(jealousy and envy, astonishment surprise, elation · normal-paced, normally alert, neutral tension, conversational) And the cost, the $250 million for a lithography machine. When you first created the process, how much did those machines? $25,000. Big scale.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, astonishment surprise, elation; style: conversational, playful; good recording, quiet background; genuineness 5.1/6; vocal-burst blend 5.9/10; 10.6s, EN.
EN_xH1xo44ACnk_W000008 · in -16.8 dBFS · gain -3.2 dB · emolia-02623
Astonishment Surprise ↓  /  Reliefproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.63.

At the same time Astonishment Surprise goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.21, then +0.11, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.54 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.58 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.54, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · zh · emolia

k 5d_a -0.620d_b 0.626step_a 0.238step_b 0.206min_cos_consec 0.5832min_cos_anchor 0.5379dataset emolialang zhspeaker ZH_B00066_S03459track ZH_B00066_S03459total 36.5slevel spread 3.9 dBmax seam 3.2 dBqmax 0.620cmax 0.238cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice
(astonishment surprise, longing, emotional numbness · measured, energised, slightly relaxed, storytelling) I could see that she was expecting a baby.
full caption & clip details
A child feminine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, longing, emotional numbness; style: storytelling, whispered; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.0/10; 3.1s, ZH.
ZH_B00066_S03459_W000493 · in -30.2 dBFS · gain +10.2 dB · emolia-03941
(distress, pain, fear · slow, very low-energy, neutral tension, storytelling) I run all the way here from weaering height, she said, goasping for breath. I couldn't countain many times. I fall en down嗯。
full caption & clip details
A child strongly feminine voice; delivery is very low-energy, slow, neutral tension, variable; timbre is slightly cool, slightly dark, smooth, slightly thin; slurred, frequent disfluency, very wide pitch range, audible breath; affect is negative, submissive, vulnerable; reads as distress, pain, fear; style: storytelling, whispered; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 2.2/10; 11.1s, ZH.
ZH_B00066_S03459_W000494 · in -29.7 dBFS · gain +9.7 dB · emolia-03941
(fear, longing, impatience and irritability · measured, energised, slightly relaxed, storytelling) Please ask me to find some dry clothes for me, and then i'll go onto the village. I'm not staying here.
full caption & clip details
A child strongly feminine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is cool, neutral-bright, smooth, thin; clear, almost no disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as fear, longing, impatience and irritability; style: storytelling, whispered; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.2/10; 6.9s, ZH.
ZH_B00066_S03459_W000495 · in -26.6 dBFS · gain +6.6 dB · emolia-03941
(affection, sexual lust, jealousy and envy · measured, normally alert, slightly relaxed, narration) First, my dear young lady, i tota you'll get warm dry, and i'll puts a bandage on that wound. Then we'll have some tea.
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is cool, slightly dark, slightly rough, thin; clear, almost no disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, slightly guarded; reads as affection, sexual lust, jealousy and envy; style: narration, storytelling; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.3/10; 9.7s, ZH.
ZH_B00066_S03459_W000496 · in -29.8 dBFS · gain +9.8 dB · emolia-03941
(relief, affection, thankfulness gratitude · measured, energised, slightly relaxed, narration) She was so exhausted, but she let me help without protesting.
full caption & clip details
An adult feminine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, neutral openness; reads as relief, affection, thankfulness gratitude; style: narration, whispered; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.8/10; 5.0s, ZH.
ZH_B00066_S03459_W000497 · in -30.5 dBFS · gain +10.5 dB · emolia-03941
Thankfulness Gratitude ↓  /  Concentrationproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.61.

At the same time Thankfulness Gratitude goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.08, then +0.20, then +0.24, then +0.25 — not a clean run: step 1 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 56 s · en · emolia

k 5d_a -0.605d_b 0.609step_a 0.198step_b 0.248min_cos_consec 0.7121min_cos_anchor 0.7651dataset emolialang enspeaker EN_WhVmSsF66mItrack EN_WhVmSsF66mItotal 55.9slevel spread 4.8 dBmax seam 4.8 dBqmax 0.605cmax 0.248cos from recomputed from spkemb_traj
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, normal-paced, slightly relaxed, moderate pitch range
(thankfulness gratitude, affection, contentment · normally alert, fairly steady, some disfluency, playful) (ahem) Three documents that Deanna shared with you all today. She gave, (low mumble) uhm, a list of
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as thankfulness gratitude, affection, contentment; style: playful, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.1/10; 5.3s, EN.
EN_WhVmSsF66mI_W000490 · in -17.2 dBFS · gain -2.8 dB · emolia-01783
(hope enthusiasm optimism, contentment, elation · normally alert, fairly steady, some disfluency, casual) You can learn how to caption videos. And of course, the distance learning team is here to help you. And so we're available to all of you so that if you want to caption your own videos, we can help you with that. (ahem) But what I will say is captioning videos is important. If you do take a video from YouTube and it has words and you feel like it's (ahem) correctly captioned,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, contentment, elation; style: casual, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.7/10; 22.9s, EN.
EN_WhVmSsF66mI_W000492 · in -19.5 dBFS · gain -0.5 dB · emolia-01783
(energised, moderately variable, some disfluency, casual) Please make sure that it is when you're looking for videos because
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, dramatic; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.3/10; 3.8s, EN.
EN_WhVmSsF66mI_W000493 · in -14.7 dBFS · gain -5.3 dB · emolia-01783
(doubt · normally alert, fairly steady, little disfluency, monologue) Just because YouTube automatically captions it, doesn't mean it's correctly captioned.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt; style: monologue, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.9/10; 4.7s, EN.
EN_WhVmSsF66mI_W000494 · in -18.9 dBFS · gain -1.1 dB · emolia-01783
(concentration · normally alert, fairly steady, some disfluency, casual) (ahem) Uhm, you would know if it's correctly captioned if the video doesn't start automatically. (ahem) Uhm, if the video has periods, commas, has uppercase lettering, uhm, (low mumble) does, does not have words misspelled, that would be correct captioning. And so, (ahem) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration; style: casual, conversational; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.4/10; 18.5s, EN.
EN_WhVmSsF66mI_W000495 · in -18.3 dBFS · gain -1.7 dB · emolia-01783
Thankfulness Gratitude ↓  /  Emotional Numbnessproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.66.

At the same time Thankfulness Gratitude goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.27 (lower than 73 % of clips in this corpus), a change of -0.66. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 34 s · en · podcast

k 4d_a -0.661d_b 0.663step_a 0.248step_b 0.239min_cos_consec 0.8024min_cos_anchor 0.8024dataset podcastlang enspeaker 65546track 65546total 33.9slevel spread 0.9 dBmax seam 0.9 dBqmax 0.661cmax 0.248cos from orange-id (speaker identity)
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, fairly steady, average clarity
(thankfulness gratitude, relief · normal-paced, neutral tension, some disfluency, conversational) they can get that help from, you know, CleanWright, or they can get that help from their distributor. Yeah. They can get that help from people in other businesses who (low mumble) uh have done similar things with restaurants, maybe. And I think that it
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, relief; style: conversational, casual; good recording, quiet background; genuineness 5.0/6; vocal-burst blend 9.9/10; 14.3s, EN.
65546_00066936 · in -19.8 dBFS · gain -0.2 dB · podcast-01153
(normal-paced, slightly relaxed, some disfluency, casual) the information is out there. You do have to do your homework,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 2.2/10; 4.2s, EN.
65546_00068360 · in -18.9 dBFS · gain -1.1 dB · podcast-01118
(anger, contempt, distress · slow, slightly relaxed, frequent disfluency, casual) (ahem) but the car wash industry has so much runway in the self serving in bay because of these car washes that are not meeting their potential.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as anger, contempt, distress; style: casual, conversational; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.9/10; 10.1s, EN.
65546_00068780 · in -19.7 dBFS · gain -0.3 dB · podcast-01120
(emotional numbness, helplessness · measured, slightly relaxed, some disfluency, casual) And the ROI is there. The operators just need the confidence
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as emotional numbness, helplessness; style: casual, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.5/10; 4.9s, EN.
65546_00069856 · in -19.5 dBFS · gain -0.5 dB · podcast-01129
Astonishment Surprise ↓  /  Disgustproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disgust barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.76.

At the same time Astonishment Surprise goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.08 (lower than 92 % of clips in this corpus), a change of -0.64. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.23, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

k 5d_a -0.640d_b 0.761step_a 0.199step_b 0.233min_cos_consec 0.9610min_cos_anchor 0.9559dataset emolialang enspeaker EN_wby6d6Ua5Lutrack EN_wby6d6Ua5Lutotal 47.8slevel spread 1.7 dBmax seam 1.6 dBqmax 0.640cmax 0.233cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, light breath, formal, authoritative) There is evidence that at least two embassies were sent to the Roman Emperor Augustus by Pandya kings
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 5.2s, EN.
EN_wby6d6Ua5Lu_W000046 · in -14.3 dBFS · gain -5.7 dB · emolia-01574
(fairly steady, light breath, formal, authoritative) Potsherds with Tamil writing have also been found in excavations on the Red Sea, suggesting the presence of Tamil merchants there
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 7.0s, EN.
EN_wby6d6Ua5Lu_W000047 · in -15.5 dBFS · gain -4.5 dB · emolia-01574
(fairly steady, light breath, formal, newsreading) An anonymous 1st century traveller's account written in Greek, Periplus Maris Arithrae, describes the ports of the Pandya and Shara kingdoms in Damarica and their commercial activity in great detail
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; mildly explicit content; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.7s, EN.
EN_wby6d6Ua5Lu_W000048 · in -15.5 dBFS · gain -4.5 dB · emolia-01574
(emotional numbness · steady, minimal breath, newsreading, formal) Peri Plus also indicates that the chief exports of the ancient Tamils were pepper, malabathrum, pearls, ivory, silk, spikenard, diamonds, sapphires, and tortoise shell.The Classical period ended around the 4th century CE with invasions by the Calabra, referred to as the Calapyrur in Tamil literature and inscriptions.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 18.7s, EN.
EN_wby6d6Ua5Lu_W000049 · in -15.5 dBFS · gain -4.5 dB · emolia-01574
(disgust · fairly steady, light breath, formal, newsreading) These invaders are described as evil kings and barbarians coming from lands to the north of the Tamil country
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 5.7s, EN.
EN_wby6d6Ua5Lu_W000050 · in -13.8 dBFS · gain -6.2 dB · emolia-01574
Disgust ↓  /  Contemplationproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.64.

At the same time Disgust goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.23, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · ko · emolia

k 4d_a -0.600d_b 0.642step_a 0.212step_b 0.231min_cos_consec 0.8834min_cos_anchor 0.8834dataset emolialang kospeaker KO_W86C06i3VxYtrack KO_W86C06i3VxYtotal 29.7slevel spread 4.1 dBmax seam 3.3 dBqmax 0.600cmax 0.231cos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(brisk, little disfluency, clear, authoritative) 또한, hdmi-264라든지, 그런 비디오 코덱 또한 gpu를 통해서 바로 디코딩이 될 수가 있습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.5/10; 5.1s, KO.
KO_W86C06i3VxY_W000027 · in -20.4 dBFS · gain +0.4 dB · emolia-03120
(normal-paced, some disfluency, average clarity, conversational) 하지만 이쪽까지는 이번 세션의 주요한 목적은 아니고요. 이쪽에 관해서 궁금하신 분들은 앞에서 설명드렸던
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.3/10; 7.7s, KO.
KO_W86C06i3VxY_W000028 · in -20.0 dBFS · gain -0.0 dB · emolia-03120
(brisk, some disfluency, average clarity, authoritative) 그 볼, 그 앞에서 적혀있던 그런 링크들을 통해서 여러 가지 메터리얼이 있으니까 그걸 보시면 될 것 같습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.0/10; 5.8s, KO.
KO_W86C06i3VxY_W000029 · in -23.3 dBFS · gain +3.3 dB · emolia-03120
(contemplation · normal-paced, some disfluency, average clarity, conversational) 이렇게 엑셀트 컴포지팅이 다 좋은데 이게 여전히 문제가 있어요. 왜냐면 여전히 느리다는 거죠. 특히 기술이 발전하고 컨텐츠들이 좋아지면 좋아질수록
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contemplation; style: conversational, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.7/10; 10.6s, KO.
KO_W86C06i3VxY_W000030 · in -24.1 dBFS · gain +4.0 dB · emolia-03120
Infatuation ↓  /  Emotional Numbnessproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness essentially absent — 0.01, lower than 99 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.77.

At the same time Infatuation goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.20 (lower than 80 % of clips in this corpus), a change of -0.64. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.09, then +0.21, then +0.25 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 51 s · en · emolia

k 5d_a -0.638d_b 0.770step_a 0.212step_b 0.249min_cos_consec 0.1887min_cos_anchor 0.1563dataset emolialang enspeaker EN_2mNWL5xJLIMtrack EN_2mNWL5xJLIMtotal 50.8slevel spread 5.8 dBmax seam 2.8 dBqmax 0.638cmax 0.249cos from recomputed from spkemb_traj
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, fairly steady, some disfluency, average clarity, moderate pitch range
(normal-paced, slightly relaxed, casual, monologue) And each kind of state having its own rules and regulations, you know, that framework that you're talking about.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.9/10; 5.9s, EN.
EN_2mNWL5xJLIM_W000059 · in -18.4 dBFS · gain -1.6 dB · emolia-00517
(doubt, fatigue exhaustion, helplessness · normal-paced, slightly relaxed, monologue, casual) Are we currently in a process where that's being discussed and worked on? Or are we still, is it still in the stages of, you know, it's, it would be a nice to have and we need to work on it.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, fatigue exhaustion, helplessness; style: monologue, casual; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.1/10; 9.6s, EN.
EN_2mNWL5xJLIM_W000060 · in -20.3 dBFS · gain +0.3 dB · emolia-00517
(jealousy and envy, pride, disappointment · brisk, neutral tension, monologue, casual) From an international perspective, the International Civil Aviation Organization is working on that with the help, obviously, from the states, because that's where the expertise relies. And, (ahem) uh, we have Nicky here with us, who is helping us a lot on that subject. But this effort is already ongoing.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, pride, disappointment; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.9/10; 18.6s, EN.
EN_2mNWL5xJLIM_W000061 · in -18.1 dBFS · gain -1.9 dB · emolia-00517
(triumph, concentration · normal-paced, slightly relaxed, formal, authoritative) It's not an easy, (low mumble) uh, effort. As you mentioned, we have a challenge because we have national security requirements. We have national culture. We have national ways of trust.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, concentration; style: formal, authoritative; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.0/10; 12.1s, EN.
EN_2mNWL5xJLIM_W000062 · in -17.4 dBFS · gain -2.6 dB · emolia-00517
(normal-paced, slightly relaxed, casual, monologue) And when we have to expand these to a global environment,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.1/10; 4.0s, EN.
EN_2mNWL5xJLIM_W000063 · in -14.6 dBFS · gain -5.4 dB · emolia-00517
Fear ↓  /  Astonishment Surpriseproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.73.

At the same time Fear goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.25 (lower than 75 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.16, then +0.13, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · en · emolia

k 5d_a -0.713d_b 0.734step_a 0.238step_b 0.224min_cos_consec 0.9157min_cos_anchor 0.9265dataset emolialang enspeaker EN_c5BYOO0j3Fytrack EN_c5BYOO0j3Fytotal 47.2slevel spread 1.4 dBmax seam 1.4 dBqmax 0.713cmax 0.238cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fear, concentration · steady, almost no disfluency, formal, newsreading) This means that the bright hemisphere is visible from Earth when Iapetus is on the western side of Saturn, and that the dark hemisphere is visible when Iapetus is on the eastern side.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, concentration; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.2s, EN.
EN_c5BYOO0j3Fy_W000008 · in -14.9 dBFS · gain -5.1 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue) The Dark Hemisphere was later named Cassini Regio in his honor
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 3.5s, EN.
EN_c5BYOO0j3Fy_W000009 · in -13.4 dBFS · gain -6.6 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue) Iopetus is named after the Titan Iopetus from Greek mythology
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.2/10; 3.6s, EN.
EN_c5BYOO0j3Fy_W000011 · in -14.4 dBFS · gain -5.6 dB · emolia-02588
(emotional numbness · fairly steady, no disfluency, newsreading, formal) The name was suggested by John Herschel – son of William Herschel, discoverer of Mimas and Enceladus – in his 1847 publication Results of Astronomical Observations Made at the Cape of Good Hope, in which he advocated naming the moons of Saturn after the Titans, brothers and sisters of the Titan Cronus – whom the Romans equated with their god Saturn.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.0s, EN.
EN_c5BYOO0j3Fy_W000012 · in -14.8 dBFS · gain -5.2 dB · emolia-02588
(fairly steady, no disfluency, authoritative, formal) When first discovered, Iapetus was among four Saturnian moons labeled the Sedera Lodoisia by their discoverer Giovanni Cassini after King Louis XIV. The other three were Tethys, Dione and Rhea
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.2s, EN.
EN_c5BYOO0j3Fy_W000013 · in -13.7 dBFS · gain -6.3 dB · emolia-02588
Doubt ↓  /  Interestproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest barely there — 0.22, lower than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.72.

At the same time Doubt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.37 (lower than 63 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.14, then +0.18, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 50 s · ja · emolia

k 5d_a -0.607d_b 0.725step_a 0.233step_b 0.210min_cos_consec 0.8801min_cos_anchor 0.8784dataset emolialang jaspeaker JA_JLiv8KUqAfytrack JA_JLiv8KUqAfytotal 50.0slevel spread 1.8 dBmax seam 1.7 dBqmax 0.607cmax 0.233cos from recomputed from spkemb_traj
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed
(doubt · normal-paced, some disfluency, average clarity, conversational) (ahem) 条件は三つで、上からまずコール ノードであること。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: conversational, casual; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 2.9/10; 4.0s, JA.
JA_JLiv8KUqAfy_W000053 · in -20.2 dBFS · gain +0.2 dB · emolia-02964
(doubt · normal-paced, some disfluency, average clarity, conversational) ファンク属性がネームでなくアトリビュートになっているときは関数じゃなくてメソッド呼び出しだったりします。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: conversational, authoritative; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.1/10; 6.5s, JA.
JA_JLiv8KUqAfy_W000054 · in -20.2 dBFS · gain +0.2 dB · emolia-02964
(measured, some disfluency, somewhat unclear, monologue) (ahem) 続いてこの判定ロジックを使ってチェッカークラスっていうのを実装してみると、こんなコードになります。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.9/10; 7.4s, JA.
JA_JLiv8KUqAfy_W000055 · in -20.3 dBFS · gain +0.3 dB · emolia-02964
(measured, frequent disfluency, somewhat unclear, monologue) (low mumble) (ahem) (ahem) (ahem) (ahem) で、このビジットメソッドの中では違反の判定を行って、もし実際に違反してた場合は、アットメッセージっていう関数、メソッドを使って違反したルールと違反しているノードを記録します。で、基本的にはこれだけでルール完成です。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 7.2/10; 21.4s, JA.
JA_JLiv8KUqAfy_W000056 · in -22.0 dBFS · gain +2.0 dB · emolia-02964
(interest, contentment · fast, some disfluency, average clarity, monologue) ASDの変換とASDの探索のところはPyLint本体が行ってくれるので、ルール自作したいなって時はチェッカークラスを実装するだけで大丈夫です。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, contentment; style: monologue, authoritative; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 5.0/10; 10.2s, JA.
JA_JLiv8KUqAfy_W000057 · in -20.8 dBFS · gain +0.8 dB · emolia-02964
Concentration ↓  /  Emotional Numbnessproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.62.

At the same time Concentration goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.16 (lower than 84 % of clips in this corpus), a change of -0.77. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.20, then +0.21, then +0.03 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.52 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.52, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · fr · emolia

k 5d_a -0.773d_b 0.616step_a 0.236step_b 0.211min_cos_consec 0.6572min_cos_anchor 0.5186dataset emolialang frspeaker FR_l_WSrm9kzxutrack FR_l_WSrm9kzxutotal 35.9slevel spread 3.5 dBmax seam 2.6 dBqmax 0.616cmax 0.236cos from recomputed from spkemb_traj
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, light breath
(concentration · normal-paced, normally alert, slightly relaxed, monologue) euh, (low mumble) par contre, avec nous, c'est un peu compliqué. Donc, on l'a bloqué, (surprised gasp) elle se retourne beaucoup, donc on la bloque avec ceci, et, (ahem) euh, pour le moment, ça fonctionne très très bien. C'est vrai qu'il n'est pas adapté pour ce lit.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 11.3s, FR.
FR_l_WSrm9kzxu_W000009 · in -19.5 dBFS · gain -0.5 dB · emolia-02843
(relief, contemplation, fatigue exhaustion · measured, subdued, slightly relaxed, whispered) Mais on a essayé de trouver des vis un peu plus longues pour, (ahem) pour le fixer et ça a bien fonctionné.
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as relief, contemplation, fatigue exhaustion; style: whispered, ASMR; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 7.3s, FR.
FR_l_WSrm9kzxu_W000010 · in -17.6 dBFS · gain -2.4 dB · emolia-02843
(measured, normally alert, slightly relaxed, whispered) Voilà sa petite table de nuit. Donc la veilleuse qui était en haut, on l'a déplacée ici en bas. Et voilà.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 6.5s, FR.
FR_l_WSrm9kzxu_W000011 · in -16.7 dBFS · gain -3.3 dB · emolia-02843
(jealousy and envy, emotional numbness · measured, normally alert, slightly relaxed, whispered) Et puis après, (low mumble) euh, voilà, de ce côté-là, il y a son petit bureau et (ahem) son petit coffre de rangement, là où elle peut mettre ses livres.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as jealousy and envy, emotional numbness; style: whispered, ASMR; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.7/10; 7.2s, FR.
FR_l_WSrm9kzxu_W000012 · in -17.7 dBFS · gain -2.3 dB · emolia-02843
(emotional numbness · slow, very low-energy, relaxed, casual) (low mumble) euh, ces, ces outils de peinture, etc.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, submissive, neutral openness; reads as emotional numbness; style: casual, ASMR; poor recording, no background noise; genuineness 4.7/6; vocal-burst blend 2.5/10; 3.0s, FR.
FR_l_WSrm9kzxu_W000013 · in -20.3 dBFS · gain +0.3 dB · emolia-02843
Thankfulness Gratitude ↓  /  Fatigue Exhaustionproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion below average — 0.27, lower than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.63.

At the same time Thankfulness Gratitude goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.11 (lower than 89 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.24, then +0.23, then +0.08 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.63 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 38 s · en · emolia

k 5d_a -0.625d_b 0.634step_a 0.212step_b 0.240min_cos_consec 0.6348min_cos_anchor 0.7004dataset emolialang enspeaker EN_WlJ2U6P_hA8track EN_WlJ2U6P_hA8total 38.2slevel spread 1.6 dBmax seam 1.1 dBqmax 0.625cmax 0.240cos from recomputed from spkemb_traj
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, quiet background
(normal-paced, normally alert, slightly relaxed, authoritative) You've got 20 questions. The pass mark is 85%. So you need 17 correct answers to pass.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, casual; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.5/10; 8.0s, EN.
EN_WlJ2U6P_hA8_W000331 · in -20.7 dBFS · gain +0.7 dB · emolia-00424
(normal-paced, normally alert, slightly relaxed, monologue) (low mumble) Out of the 80 words on this course, we found exceptions in.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.0/10; 4.4s, EN.
EN_WlJ2U6P_hA8_W000332 · in -21.2 dBFS · gain +1.2 dB · emolia-00424
(confusion, doubt, contemplation · slow, normally alert, slightly relaxed, casual) What percentage do you remember?
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, doubt, contemplation; style: casual, monologue; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 0.4/10; 3.3s, EN.
EN_WlJ2U6P_hA8_W000333 · in -22.3 dBFS · gain +2.3 dB · emolia-00424
(measured, normally alert, neutral tension, casual) Yeah, 20%. You know, if I found exceptions in, in 50% of the words of English, then I don't think I'd be doing this course. I think I've said that before.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 2.4/10; 11.6s, EN.
EN_WlJ2U6P_hA8_W000334 · in -21.3 dBFS · gain +1.3 dB · emolia-00424
(fatigue exhaustion · slow, very low-energy, relaxed, monologue) But if you've got an 80% chance of getting it right, and then just learning some sight words, I think that's a pretty good, uh, (low mumble)
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, submissive, slightly guarded; reads as fatigue exhaustion; style: monologue, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_WlJ2U6P_hA8_W000335 · in -21.4 dBFS · gain +1.4 dB · emolia-00424
Emotional Numbness ↓  /  Sexual Lustproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust below average — 0.26, lower than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.67.

At the same time Emotional Numbness goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.23 (lower than 77 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.19, then +0.25, then +0.01 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · zh · emolia

k 5d_a -0.650d_b 0.673step_a 0.246step_b 0.246min_cos_consec 0.8165min_cos_anchor 0.7863dataset emolialang zhspeaker ZH_B00030_S09886track ZH_B00030_S09886total 35.8slevel spread 3.9 dBmax seam 3.9 dBqmax 0.650cmax 0.246cos from recomputed from spkemb_traj
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · quiet background
(fast, normally alert, slightly relaxed, casual) 你你你让我禁烟就给,等于给了我这样的一个捞钱的机会。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, storytelling; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.3/10; 4.1s, ZH.
ZH_B00030_S09886_W000427 · in -16.4 dBFS · gain -3.6 dB · emolia-03578
(impatience and irritability, contempt, confusion · fast, energised, neutral tension, storytelling) 所以没有人会真的真的干。林则徐知道这一点,林则徐说我这个办法,我可以这个我可以治。
full caption & clip details
An adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, guarded; reads as impatience and irritability, contempt, confusion; style: storytelling, casual; poor recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.7/10; 5.3s, ZH.
ZH_B00030_S09886_W000428 · in -12.5 dBFS · gain -7.5 dB · emolia-03578
(impatience and irritability, bitterness, distress · fast, energised, neutral tension, casual) 我可以去改造水师,我可以改造这官僚机构。我可以去整顿他们,他确实这么做了,但是他被他放了。另外一点就是。
full caption & clip details
An elderly masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; average clarity, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, bitterness, distress; style: casual, storytelling; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 5.2/10; 7.9s, ZH.
ZH_B00030_S09886_W000429 · in -14.7 dBFS · gain -5.3 dB · emolia-03578
(intoxication altered states of consciousness, impatience and irritability, contempt · measured, energised, neutral tension, casual) 就是说你这边整顿了,你这边认真去进了那那边呢,人家抗拒不交,你怎么办啊,抗拒不交怎么办?因为。
full caption & clip details
An elderly masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness, impatience and irritability, contempt; style: casual, storytelling; below-average recording, quiet background; genuineness 5.0/6; vocal-burst blend 4.8/10; 10.1s, ZH.
ZH_B00030_S09886_W000430 · in -14.8 dBFS · gain -5.2 dB · emolia-03578
(sexual lust, malevolence malice · normal-paced, normally alert, slightly relaxed, conversational) (low mumble) (ahem) 林正席其实并不敢,就当时他那个医馆就是有一个二层三呃呃,三层小楼啊,他们住那里头。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sexual lust, malevolence malice; style: conversational, didactic; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.7/10; 7.8s, ZH.
ZH_B00030_S09886_W000431 · in -16.1 dBFS · gain -3.9 dB · emolia-03578
Infatuation ↓  /  Fatigue Exhaustionproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.65.

At the same time Infatuation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.30 (lower than 70 % of clips in this corpus), a change of -0.68. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.17, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.51 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.60 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.51, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 18 s · en · emolia

k 4d_a -0.677d_b 0.655step_a 0.248step_b 0.246min_cos_consec 0.6030min_cos_anchor 0.5063dataset emolialang enspeaker EN_QnrLXaeufAEtrack EN_QnrLXaeufAEtotal 17.8slevel spread 5.7 dBmax seam 5.7 dBqmax 0.655cmax 0.248cos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(infatuation, elation, affection · normal-paced, some disfluency, average clarity, casual) Very cool to see Lando, he's charming and manipulative as ever.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, elation, affection; style: casual, monologue; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.5/10; 3.9s, EN.
EN_QnrLXaeufAE_W000113 · in -15.7 dBFS · gain -4.3 dB · emolia-01185
(normal-paced, some disfluency, average clarity, monologue) Hira considers Chopra a member of the crew, not an object to be bargained. I appreciate that.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.8/10; 4.7s, EN.
EN_QnrLXaeufAE_W000114 · in -17.5 dBFS · gain -2.5 dB · emolia-01185
(normal-paced, some disfluency, average clarity, casual) You know, they haven't gone very far from Lothal most of the season. You know, Clone Wars would go...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.9/10; 5.4s, EN.
EN_QnrLXaeufAE_W000115 · in -11.8 dBFS · gain -8.2 dB · emolia-01185
(measured, frequent disfluency, somewhat unclear, casual) More, you know, between locations and...
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 3.9/6; vocal-burst blend 2.4/10; 3.3s, EN.
EN_QnrLXaeufAE_W000116 · in -14.0 dBFS · gain -6.0 dB · emolia-01185
Anger ↓  /  Intoxication Altered States of Consciousnessproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness barely there — 0.19, lower than 81 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.61.

At the same time Anger goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.28 (lower than 72 % of clips in this corpus), a change of -0.66. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.04, then +0.09, then +0.23 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 40 s · de · emolia

k 5d_a -0.659d_b 0.613step_a 0.246step_b 0.249min_cos_consec 0.7240min_cos_anchor 0.7742dataset emolialang despeaker DE_hmxC5jCS7CQtrack DE_hmxC5jCS7CQtotal 40.4slevel spread 4.6 dBmax seam 4.1 dBqmax 0.613cmax 0.249cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(anger, fatigue exhaustion · casual, playful) die halt irgendwie sagen, bau die Straße um oder bau so und so viele Straßen um und man hängt ganz häufig an Stellen, wo übergeordnete Gesetze einfach im Weg stehen. Daher,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, fatigue exhaustion; style: casual, playful; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 0.0/10; 10.1s, DE.
DE_hmxC5jCS7CQ_W000029 · in -18.6 dBFS · gain -1.4 dB · emolia-00200
(authoritative, casual) Okay. Okay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, casual; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 0.0/10; 6.4s, DE.
DE_hmxC5jCS7CQ_W000030 · in -19.2 dBFS · gain -0.8 dB · emolia-00200
(interest, concentration · monologue, authoritative) so im Weg stehen, gerade auch auf Landes- und Bundesebene, ein bisschen aus dem Weg zu räumen. Und in den Köpfen war da schon so ein bisschen die Idee geboren. Wir müssen das eigentlich auch nochmal auf Landesebene umsetzen. Ja, und daraus kam dann im Prinzip auch so ein bisschen diese Initiative.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: monologue, authoritative; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.1/10; 15.5s, DE.
DE_hmxC5jCS7CQ_W000031 · in -19.9 dBFS · gain -0.1 dB · emolia-00200
(astonishment surprise · conversational, casual) Das kann man vielleicht zu einem anderen Zeitpunkt noch mal vorstellen. Das läuft auch gerade erst an.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: conversational, casual; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 0.0/10; 4.8s, DE.
DE_hmxC5jCS7CQ_W000032 · in -19.4 dBFS · gain -0.6 dB · emolia-00200
(conversational, playful) Ja, und wo will man eigentlich hin und was hat man vor?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, playful; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.0/10; 3.0s, DE.
DE_hmxC5jCS7CQ_W000033 · in -15.3 dBFS · gain -4.7 dB · emolia-00200
Fatigue Exhaustion ↓  /  Astonishment Surpriseproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.61.

At the same time Fatigue Exhaustion goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.26 (lower than 74 % of clips in this corpus), a change of -0.63. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.13, then +0.09, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 72 s · en · emolia

k 5d_a -0.630d_b 0.611step_a 0.217step_b 0.224min_cos_consec 0.7025min_cos_anchor 0.7372dataset emolialang enspeaker EN_hR39MCKdgcytrack EN_hR39MCKdgcytotal 71.8slevel spread 5.4 dBmax seam 4.2 dBqmax 0.611cmax 0.224cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · slightly rough, quiet background, very low-energy, relaxed, frequent disfluency, slurred, fairly narrow pitch
(slow, steady, audible breath, whispered) So for today, we will be doing a challenge yourself step video and today's trick, we will be doing Newton's coin trick.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; no dominant emotion; style: whispered, monologue; below-average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.1/10; 7.0s, EN.
EN_hR39MCKdgcy_W000001 · in -18.1 dBFS · gain -1.9 dB · emolia-02446
(measured, fairly steady, normal breath, whispered) So for today, the materials that you will need will include a water bottle, any dollar bill, and some coins.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, submissive, neutral openness; no dominant emotion; style: whispered, ASMR; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 2.0/10; 7.6s, EN.
EN_hR39MCKdgcy_W000002 · in -17.2 dBFS · gain -2.8 dB · emolia-02446
(doubt, contemplation, jealousy and envy · slow, fairly steady, normal breath, whispered) So have you ever really wanted to challenge a friend or even make a bet against him? Well, here's something you can do. You can challenge them to take this $20 off of this water bottle. But here's the catch. You can't make these coins fall off. So, I'll give you a second to think. How could you possibly get this $20 off
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as doubt, contemplation, jealousy and envy; style: whispered, ASMR; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.2/10; 22.4s, EN.
EN_hR39MCKdgcy_W000003 · in -21.4 dBFS · gain +1.4 dB · emolia-02446
(contentment, teasing · measured, fairly steady, normal breath, whispered) Without making these coins don't fall off. So have you thought of a solution yet? Well, if not, I'll tell you the trick to it. So what you simply want to do is to put your finger right here and just quickly chop off the $20 bill. Like this.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, slightly guarded; reads as contentment, teasing; style: whispered, casual; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 5.5/10; 14.2s, EN.
EN_hR39MCKdgcy_W000004 · in -22.1 dBFS · gain +2.1 dB · emolia-02446
(astonishment surprise, interest, concentration · slow, steady, normal breath, whispered) You see, the coins didn't fall off. Okay, so now you might be wondering, how the heck did a $20 bill fall off, but not the coins? Well, there are two things that explain this. First, we have Newton's First Law, which can also be known as inertia. And second, we have friction. Okay, so...
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as astonishment surprise, interest, concentration; style: whispered, ASMR; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.9/10; 20.0s, EN.
EN_hR39MCKdgcy_W000005 · in -22.6 dBFS · gain +2.5 dB · emolia-02446
Confusion ↓  /  Triumphproxy_taillift__PXR__T0.60__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.61.

At the same time Confusion goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.34 (lower than 66 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.20, then +0.12, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.55 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.60 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.55, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · de · emolia

k 5d_a -0.650d_b 0.611step_a 0.188step_b 0.219min_cos_consec 0.6026min_cos_anchor 0.5474dataset emolialang despeaker DE_o4H1G-mk3-Utrack DE_o4H1G-mk3-Utotal 33.3slevel spread 4.1 dBmax seam 2.9 dBqmax 0.611cmax 0.219cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, balanced body, quiet background
(confusion, intoxication altered states of consciousness, doubt · slow, normally alert, slightly relaxed, casual) Oh, ja. Ja, was willst denn du? Ah, so.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as confusion, intoxication altered states of consciousness, doubt; style: casual, conversational; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.0/10; 3.2s, DE.
DE_o4H1G-mk3-U_W000000 · in -21.3 dBFS · gain +1.3 dB · emolia-00141
(longing · normal-paced, normally alert, slightly relaxed, authoritative) Nee, bis jetzt noch nicht, Manu. Außer ein 87er Grießmann Goldkarte. Heute starten wir mal mit der Eins rein, Jungs.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: authoritative, playful; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 7.6s, DE.
DE_o4H1G-mk3-U_W000001 · in -19.4 dBFS · gain -0.6 dB · emolia-00141
(relief · normal-paced, normally alert, slightly relaxed, authoritative) er sieht sogar spielbar aus, ne, sag ich dir, wie's ist. Gib dem Shadow drauf und abfahrt.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief; style: authoritative, playful; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.2/10; 5.9s, DE.
DE_o4H1G-mk3-U_W000002 · in -17.2 dBFS · gain -2.8 dB · emolia-00141
(malevolence malice, bitterness, jealousy and envy · normal-paced, energised, neutral tension, authoritative) Or, I mean, because of Sentinel. If the pace is enough for you. But you can even play it, ey. Full 3, gg. Do you want Jordi Alba, Prims?
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, bitterness, jealousy and envy; style: authoritative, casual; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.3/10; 10.7s, DE.
DE_o4H1G-mk3-U_W000003 · in -20.1 dBFS · gain +0.1 dB · emolia-00141
(triumph, disgust, anger · brisk, highly aroused, tense, cartoonish) Konnte das Spiel 10 Tage früher zocken. Konnte Content bringen. Oha! Digga!
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, guarded; reads as triumph, disgust, anger; style: cartoonish, dramatic; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.5/10; 5.3s, DE.
DE_o4H1G-mk3-U_W000004 · in -18.4 dBFS · gain -1.6 dB · emolia-00141