sc-AB2-k3

Two-sided AB2 after speaker cleaning (WavLM -id >= 0.80 on consecutive pairs AND against the first clip), k=3. This is the only source that ships min_cos_consec / min_cos_anchor.

Rule. AB2 — two-sided: emotion A falls by >=T while emotion B rises by >=T, each consecutive step <=C
Source. trajectories_v3_speaker_clean.parquet  |  Family. speaker-cleaned two-sided set (WavLM -id >= 0.80, consecutive AND anchored)
Sampled from 90,362 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Pride ↓  /  Jealousy and Envysc-AB2-k3 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Jealousy and Envy up — by at least 0.25 each.

The chain starts with Jealousy and Envy clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.

At the same time Pride goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 51 s · es · podcast

hear it un-normalised (raw levels, max seam 2.1 dB)
k 3d_a -0.274d_b 0.263step_a 0.164step_b 0.235min_cos_consec 0.8893min_cos_anchor 0.8803dataset podcastlang esspeaker 844642track 844642total 50.7slevel spread 2.1 dBmax seam 2.1 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · slightly cool, neutral-bright, quiet background, normal-paced, normally alert, neutral tension, moderately variable, some disfluency
(pride, affection, disgust · somewhat unclear, wide pitch range, casual, storytelling) Es como que nomás este dar como una plática, ¿no? Así como te ha afectado a ti, a tus seres cercanos, y tú cuentas lo mismo y lo ya después sacamos una conclusión, ¿no?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pride, affection, disgust; style: casual, storytelling; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 8.3/10; 13.2s, ES.
844642_00051584 · in -30.4 dBFS · gain +10.4 dB · podcast-00528
(affection, disappointment, helplessness · somewhat unclear, wide pitch range, casual) Pues que la gente nunca va a entender que hay que quedarnos en la casa por más grave que sea la situación. Porque pues obviamente de algo tenemos que vivir y todos los trabajadores necesitan sacar dinero para seguir con sus vidas. Entonces es obviamente que nunca vamos a poder hacer una cuarentena bien sin nadie salir. Por la economía y todo eso. Pero
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as affection, disappointment, helplessness; style: casual; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 23.7s, ES.
844642_00068856 · in -29.0 dBFS · gain +9.0 dB · podcast-00543
(jealousy and envy, sadness · slurred, moderate pitch range, casual, storytelling) que si van a salir, pues, o sea, salgan con cubrebocas, gel antibacterial, sanitizador y todo eso. O sea, para evitar los contagios y no acercarse mucho a las personas ni nada.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; slurred, some disfluency, moderate pitch range, light breath; affect is positive, neutral stance, slightly guarded; reads as jealousy and envy, sadness; style: casual, storytelling; poor recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.7/10; 13.6s, ES.
844642_00071220 · in -31.1 dBFS · gain +11.1 dB · podcast-00529
Concentration ↓  /  Interestsc-AB2-k3 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.40.

At the same time Concentration goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · en · emolia

hear it un-normalised (raw levels, max seam 2.4 dB)
k 3d_a -0.405d_b 0.401step_a 0.225step_b 0.211min_cos_consec 0.8427min_cos_anchor 0.8031dataset emolialang enspeaker EN_MVYTjoITWtUtrack EN_MVYTjoITWtUtotal 30.7slevel spread 2.4 dBmax seam 2.4 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, normally alert, slightly relaxed, fairly steady, light breath
(slow, almost no disfluency, clear, monologue) So in Daniel 7, Daniel describes the rise and fall of four world kingdoms and then the establishment of a fifth kingdom.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 9.7s, EN.
EN_MVYTjoITWtU_W000150 · in -17.7 dBFS · gain -2.3 dB · emolia-02510
(normal-paced, almost no disfluency, clear, authoritative) Daniel, however, adds two important ideas to the ones already mentioned. So Daniel talks about the four earthly kingdoms.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.1/10; 8.4s, EN.
EN_MVYTjoITWtU_W000151 · in -19.7 dBFS · gain -0.3 dB · emolia-02510
(interest, malevolence malice · measured, some disfluency, average clarity, monologue) The Babylonians, the (low mumble) Mesopotamians, (low mumble) the Persians, the Greeks, the Romans. He talks about those four with the statue and everything. And then he says a fifth kingdom will come, a stone will come.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, malevolence malice; style: monologue, authoritative; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.0/10; 12.4s, EN.
EN_MVYTjoITWtU_W000152 · in -17.4 dBFS · gain -2.6 dB · emolia-02510
Interest ↓  /  Intoxication Altered States of Consciousnesssc-AB2-k3 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness around average — 0.47, lower than 53 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.25.

At the same time Interest goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.18 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · en · emolia

hear it un-normalised (raw levels, max seam 5.2 dB)
k 3d_a -0.387d_b 0.254step_a 0.201step_b 0.179min_cos_consec 0.8696min_cos_anchor 0.8446dataset emolialang enspeaker EN_-7o_B_jU1Fgtrack EN_-7o_B_jU1Fgtotal 37.4slevel spread 5.2 dBmax seam 5.2 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · balanced body, average recording, quiet background, normal-paced, fairly steady, average clarity, moderate pitch range, light breath
(interest, concentration · normally alert, slightly relaxed, some disfluency, didactic) Quite a reasonably strong capital about the core principle which are articulating the world trading system. These are transparency, good faith.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, concentration; style: didactic, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.1/10; 8.0s, EN.
EN_-7o_B_jU1Fg_W000018 · in -23.1 dBFS · gain +3.1 dB · emolia-00448
(concentration, contemplation, doubt · very low-energy, neutral tension, frequent disfluency) (ahem) and non-discrimination. And you do not see radical contestations or radical dispute over those principles. Even in the worst situation like, (low mumble) uh, (ahem) the, uh, the one with the Trump administration, nobody has been walking out of the WPO. So you don't see really disagreement on the fact that we can cooperate on this basis. And in fact, we, we even had some (low mumble) successes (low mumble) this year.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, contemplation, doubt; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.2/10; 24.2s, EN.
EN_-7o_B_jU1Fg_W000019 · in -23.2 dBFS · gain +3.2 dB · emolia-00448
(normally alert, slightly relaxed, frequent disfluency, didactic) I may come back to it. What we face and we are confronted with is (low mumble) several trends which
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: didactic, conversational; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.8/10; 5.0s, EN.
EN_-7o_B_jU1Fg_W000020 · in -18.0 dBFS · gain -2.0 dB · emolia-00448
Impatience and Irritability ↓  /  Confusionsc-AB2-k3 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.

At the same time Impatience and Irritability goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · zh · emolia

hear it un-normalised (raw levels, max seam 2.2 dB)
k 3d_a -0.260d_b 0.265step_a 0.147step_b 0.182min_cos_consec 0.8444min_cos_anchor 0.8309dataset emolialang zhspeaker ZH_B00072_S07803track ZH_B00072_S07803total 18.0slevel spread 2.2 dBmax seam 2.2 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, no background noise, normally alert, slightly relaxed, fairly steady, moderate pitch range
(measured, no disfluency, clear, authoritative) 咱们就以这贼银起火微信号吧。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, storytelling; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.8/10; 4.1s, ZH.
ZH_B00072_S07803_W000051 · in -16.7 dBFS · gain -3.3 dB · emolia-03995
(pain, malevolence malice · fast, some disfluency, very clear, storytelling) 一切技艺妥当,除了杨文广,太君,又命杜金娥和焦廷贵、孟怀元。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; very clear, some disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as pain, malevolence malice; style: storytelling, dramatic; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.9/10; 7.3s, ZH.
ZH_B00072_S07803_W000052 · in -18.9 dBFS · gain -1.1 dB · emolia-03995
(confusion, intoxication altered states of consciousness, malevolence malice · measured, no disfluency, very clear, didactic) 蛇老太君派了一支精锐的这骑兵交付于这桂英。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; very clear, no disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, intoxication altered states of consciousness, malevolence malice; style: didactic, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.2/10; 6.3s, ZH.
ZH_B00072_S07803_W000053 · in -18.1 dBFS · gain -1.9 dB · emolia-03995
Thankfulness Gratitude ↓  /  Infatuationsc-AB2-k3 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation below average — 0.33, lower than 67 % of clips in this corpus — and ends with it clearly present at 0.65, higher than 65 % of clips in this corpus. That is a total rise of 0.32.

At the same time Thankfulness Gratitude goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 13 s · en · emolia

hear it un-normalised (raw levels, max seam 1.7 dB)
k 3d_a -0.332d_b 0.319step_a 0.178step_b 0.238min_cos_consec 0.9645min_cos_anchor 0.9645dataset emolialang enspeaker EN_SqmyliCH0K4track EN_SqmyliCH0K4total 13.3slevel spread 1.7 dBmax seam 1.7 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, fairly narrow pitch, casual, formal) Arthur Vandenbergh, Michigan, 1949–1951
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 4.3s, EN.
EN_SqmyliCH0K4_W000051 · in -14.0 dBFS · gain -6.0 dB · emolia-01368
(fairly steady, moderate pitch range, formal, casual) Arthur Kapper – Kansas, 1945–1949
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.5/10; 4.1s, EN.
EN_SqmyliCH0K4_W000052 · in -15.7 dBFS · gain -4.3 dB · emolia-01368
(fairly steady, fairly narrow pitch, casual, formal) Hiram W. Johnson, California, 1940–1945
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.4/10; 4.6s, EN.
EN_SqmyliCH0K4_W000053 · in -15.5 dBFS · gain -4.5 dB · emolia-01368
Pain ↓  /  Infatuationsc-AB2-k3 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation below average — 0.30, lower than 70 % of clips in this corpus — and ends with it clearly present at 0.70, higher than 70 % of clips in this corpus. That is a total rise of 0.39.

At the same time Pain goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · zh · emolia

k 3d_a -0.408d_b 0.392step_a 0.235step_b 0.234min_cos_consec 0.9547min_cos_anchor 0.9538dataset emolialang zhspeaker ZH_B00054_S08248track ZH_B00054_S08248total 23.5slevel spread 0.2 dBmax seam 0.2 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(pain · narration, monologue) 难道是因为激进分子在他身上装了炸弹,打算炸掉客件站吗?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: narration, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.2/10; 6.0s, ZH.
ZH_B00054_S08248_W000012 · in -18.3 dBFS · gain -1.7 dB · emolia-03820
(monologue, formal) 刚刚结束一场风波云涌,王位之争的奥斯曼居民警惕性非常之高。有人联想到这点,直接疏散人群,连忙拨打电话求助警察。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.0/10; 11.6s, ZH.
ZH_B00054_S08248_W000013 · in -18.3 dBFS · gain -1.7 dB · emolia-03820
(formal, narration) 还有些则去找了客件站站岗的安保人员说明情况,一会儿的功夫。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.1/10; 5.5s, ZH.
ZH_B00054_S08248_W000014 · in -18.5 dBFS · gain -1.5 dB · emolia-03820
Shame ↓  /  Doubtsc-AB2-k3 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Doubt up — by at least 0.25 each.

The chain starts with Doubt clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.25.

At the same time Shame goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · zh · emolia

k 3d_a -0.292d_b 0.253step_a 0.184step_b 0.147min_cos_consec 0.8949min_cos_anchor 0.9067dataset emolialang zhspeaker ZH_B00020_S00879track ZH_B00020_S00879total 27.3slevel spread 0.5 dBmax seam 0.5 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, measured, normally alert, slightly relaxed
(fairly narrow pitch, formal, narration) 如果整个年轻一代是怀着愤慨和鄙视的心情审视,先是战败,而后又获得和平的父辈。这又有什么可奇怪的呢?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.2/10; 9.3s, ZH.
ZH_B00020_S00879_W000172 · in -22.6 dBFS · gain +2.6 dB · emolia-03481
(pain · moderate pitch range, narration, formal) 难道他们不是把一切都搞砸了吗?难道他们不是什么也没遇见到吗?难道他们不是把一切都算记错了吗?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: narration, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.7/10; 10.0s, ZH.
ZH_B00020_S00879_W000173 · in -23.1 dBFS · gain +3.0 dB · emolia-03481
(fairly narrow pitch, narration, monologue) 如果年轻一代从此失去了一切尊严,因而怨恨和鄙视自己的父辈,不是很容易理解吗?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.5/10; 7.8s, ZH.
ZH_B00020_S00879_W000174 · in -22.5 dBFS · gain +2.5 dB · emolia-03481
Contentment ↓  /  Emotional Numbnesssc-AB2-k3 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Contentment down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.26.

At the same time Contentment goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · ja · emolia

k 3d_a -0.488d_b 0.261step_a 0.246step_b 0.225min_cos_consec 0.8714min_cos_anchor 0.8897dataset emolialang jaspeaker JA_LHZW4EAEoMgtrack JA_LHZW4EAEoMgtotal 36.0slevel spread 3.1 dBmax seam 3.1 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, measured, slightly relaxed, fairly steady, light breath
(contentment, relief, longing · normally alert, some disfluency, somewhat unclear, monologue) 刈り込みバサミで刈り込むと、刃先を全部切ってしまいますので、切り口が全部茶色くなります。それが見苦しいと私は思うので、あまり刈り込みバサミを使うことはありませんけれども、全く使わないかというとそういうわけでもありません。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, relief, longing; style: monologue, formal; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 4.2/10; 17.7s, JA.
JA_LHZW4EAEoMg_W000063 · in -22.2 dBFS · gain +2.2 dB · emolia-02974
(subdued, little disfluency, somewhat unclear, monologue) あまり一目につかないところや、現場の状況によっては、刈りこまずあるを得ない場合もあります。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: monologue, ASMR; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 4.2/10; 7.0s, JA.
JA_LHZW4EAEoMg_W000064 · in -19.9 dBFS · gain -0.1 dB · emolia-02974
(emotional numbness, contemplation, longing · subdued, some disfluency, slurred, monologue) ハサキを切ってしまうなら、カリゴムバサミを使った方が早いに決まっています。キバサミを使っているのに、ハサキを切ってしまうというのは、何か無駄なような気がします。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as emotional numbness, contemplation, longing; style: monologue, whispered; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 6.2/10; 11.0s, JA.
JA_LHZW4EAEoMg_W000065 · in -23.1 dBFS · gain +3.1 dB · emolia-02974
Infatuation ↓  /  Disgustsc-AB2-k3 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Disgust up — by at least 0.25 each.

The chain starts with Disgust around average — 0.53, higher than 53 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.34.

At the same time Infatuation goes the other way, from 0.67 (higher than 67 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 15 s · en · emolia

k 3d_a -0.286d_b 0.343step_a 0.211step_b 0.223min_cos_consec 0.9556min_cos_anchor 0.9393dataset emolialang enspeaker EN_5cKKtOYK_wMtrack EN_5cKKtOYK_wMtotal 15.5slevel spread 1.6 dBmax seam 1.6 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, authoritative) On the Iland Islands, a similar drink is made under the brand name Coba Libre
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.2/10; 4.3s, EN.
EN_5cKKtOYK_wM_W000196 · in -12.1 dBFS · gain -7.9 dB · emolia-02399
(fairly steady, formal, authoritative) In some European countries, especially in the Czech Republic and Germany, beet sugar is also used to make rectified spirit and vodka
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.1/10; 7.3s, EN.
EN_5cKKtOYK_wM_W000197 · in -13.8 dBFS · gain -6.2 dB · emolia-02399
(steady, formal, authoritative) An unrefined sugary syrup is produced directly from the sugar beet
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.1/10; 3.6s, EN.
EN_5cKKtOYK_wM_W000198 · in -13.1 dBFS · gain -6.9 dB · emolia-02399
Interest ↓  /  Longingsc-AB2-k3 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Longing up — by at least 0.25 each.

The chain starts with Longing clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.32.

At the same time Interest goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.09 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · ko · emolia

k 3d_a -0.325d_b 0.319step_a 0.204step_b 0.233min_cos_consec 0.9131min_cos_anchor 0.8889dataset emolialang kospeaker KO_hjGdx1mad9ytrack KO_hjGdx1mad9ytotal 31.5slevel spread 0.8 dBmax seam 0.8 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(measured, somewhat unclear, monologue, didactic) 에, 그런 것들로 여러분들이 많이 사용하셔도 됩니다. 그래서, 시간이 있을 때마다 이렇게, (low mumble) 어, 찍으시면은 유튜브에 가는 것, 유튜브도, 어, (ahem) 많은 홍보들을 할 수 있고요. 그 다음에,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.5/10; 11.7s, KO.
KO_hjGdx1mad9y_W000022 · in -18.1 dBFS · gain -1.9 dB · emolia-03241
(normal-paced, average clarity, monologue, didactic) (surprised gasp) 어, 나중에 여러분들이 강의를 한다면 강사로 활동을 할 때도 큰 도움이 되실 겁니다. 그래서 하루에 한 시간 정도 분량을 만든다면 됩니다. 그것들 얼마나 지속하냐. 예, 그게 중요하겠죠.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 7.2/10; 11.5s, KO.
KO_hjGdx1mad9y_W000023 · in -18.6 dBFS · gain -1.4 dB · emolia-03241
(longing, contemplation · normal-paced, average clarity, monologue, casual) 그 다음에 이제 하루에는 보통 책 한 권 정도 분량을 봐요. 뭐, 정말로 책을 한 권을 처음부터 끝까지 이렇게 보는데, 뭐,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, contemplation; style: monologue, casual; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 3.2/10; 8.0s, KO.
KO_hjGdx1mad9y_W000024 · in -17.8 dBFS · gain -2.2 dB · emolia-03241
Infatuation ↓  /  Concentrationsc-AB2-k3 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.33.

At the same time Infatuation goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · zh · emolia

k 3d_a -0.256d_b 0.331step_a 0.245step_b 0.190min_cos_consec 0.8051min_cos_anchor 0.8051dataset emolialang zhspeaker ZH_B00006_S01331track ZH_B00006_S01331total 18.6slevel spread 5.2 dBmax seam 5.2 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(measured, steady, formal, monologue) 大唐帝国,天可汗的雄风已经不见了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.2/10; 3.2s, ZH.
ZH_B00006_S01331_W000000 · in -21.0 dBFS · gain +1.0 dB · emolia-03333
(measured, fairly steady, formal, monologue) 他像个生病的老人,苟延残喘的活着。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.5/10; 3.3s, ZH.
ZH_B00006_S01331_W000001 · in -26.2 dBFS · gain +6.2 dB · emolia-03333
(normal-paced, fairly steady, monologue, didactic) 皇帝为了平定安史之乱,曾经向所有的叛军将领表示,只要大家投降,不但不追究过错,而且还封给他们一个称作节度使的官座。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.3/10; 11.8s, ZH.
ZH_B00006_S01331_W000002 · in -24.4 dBFS · gain +4.4 dB · emolia-03333
Concentration ↓  /  Infatuationsc-AB2-k3 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.36.

At the same time Concentration goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · zh · emolia

k 3d_a -0.257d_b 0.358step_a 0.155step_b 0.209min_cos_consec 0.8825min_cos_anchor 0.8601dataset emolialang zhspeaker ZH_B00043_S04366track ZH_B00043_S04366total 19.0slevel spread 1.2 dBmax seam 1.2 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, measured, normally alert, slightly relaxed
(steady, no disfluency, clear, monologue) 这种享受比我在真实现实中所能体会到的任何东西都真实。因此。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.1/10; 8.1s, ZH.
ZH_B00043_S04366_W000140 · in -22.6 dBFS · gain +2.6 dB · emolia-03708
(steady, no disfluency, clear, monologue) 它既可以当做是对真实界加以排斥,没有障碍的想象空间的媒介。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.6/10; 6.7s, ZH.
ZH_B00043_S04366_W000141 · in -23.9 dBFS · gain +3.9 dB · emolia-03708
(fairly steady, some disfluency, somewhat unclear, monologue) 同时也能充当你接近真实界的空间。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 3.5/10; 3.9s, ZH.
ZH_B00043_S04366_W000142 · in -23.0 dBFS · gain +3.0 dB · emolia-03708
Interest ↓  /  Fearsc-AB2-k3 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Fear up — by at least 0.25 each.

The chain starts with Fear clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.32.

At the same time Interest goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 42 s · en · emolia

k 3d_a -0.283d_b 0.319step_a 0.243step_b 0.173min_cos_consec 0.8723min_cos_anchor 0.8597dataset emolialang enspeaker EN_B00038_S08968track EN_B00038_S08968total 41.8slevel spread 5.7 dBmax seam 5.7 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, slightly relaxed, fairly steady
(interest, concentration · very low-energy, casual) Healing and (low mumble) Egyptian rituals for protection of the body and so on. They very often mixed elements that we would think of as medical, like drinking a particular potion or (low mumble) doing something else. And then things that we would classify as magic, so (low mumble) rituals and spoken incantations and so on.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, concentration; style: casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 6.2/10; 19.9s, EN.
EN_B00038_S08968_W000015 · in -26.6 dBFS · gain +6.7 dB · emolia-00987
(doubt, emotional numbness · normally alert, conversational, monologue) There is a great royal wife who is depicted and identified as such in Ramesses own inscriptions, but he doesn't give her name, so we don't actually know which of his wives this is.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, emotional numbness; style: conversational, monologue; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.0/10; 11.0s, EN.
EN_B00038_S08968_W000016 · in -20.9 dBFS · gain +0.9 dB · emolia-00987
(fear · normally alert, monologue) We do know that the Lady Isis has the title in her own tune, but she's buried much later, and so it's not clear when she was appointed as the Great Royal Wife.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear; style: monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.6/10; 10.6s, EN.
EN_B00038_S08968_W000017 · in -21.8 dBFS · gain +1.8 dB · emolia-00987
Embarrassment ↓  /  Infatuationsc-AB2-k3 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Embarrassment down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.31.

At the same time Embarrassment goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · en · podcast

k 3d_a -0.259d_b 0.310step_a 0.174step_b 0.221min_cos_consec 0.8214min_cos_anchor 0.8622dataset podcastlang enspeaker 788096track 788096total 26.0slevel spread 2.1 dBmax seam 1.2 dBcos from orange-id (speaker identity)
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, quiet background, frequent disfluency
(embarrassment, pride · normal-paced, normally alert, slightly relaxed, monologue) I believe not. (chuckle) Because (ahem) we have to make (low mumble) um huge amount of bread, so we cannot take a break.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment, pride; style: monologue, whispered; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.8/10; 7.8s, EN.
788096_00017312 · in -28.1 dBFS · gain +8.1 dB · podcast-03438
(measured, normally alert, relaxed, casual) making bread. We don't have to serve the customer. (low mumble) Um however
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.0/10; 4.5s, EN.
788096_00018624 · in -26.9 dBFS · gain +6.9 dB · podcast-03464
(infatuation, pleasure ecstasy · measured, very low-energy, relaxed, monologue) human power. Like (low mumble) um for example one bag of flour can be about twenty-five kilom kilogram and for one night we have about twenty of them for the first course. (chuckle)
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as infatuation, pleasure ecstasy; style: monologue, whispered; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.1/10; 13.4s, EN.
788096_00021144 · in -25.9 dBFS · gain +5.9 dB · podcast-03444
Concentration ↓  /  Doubtsc-AB2-k3 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Doubt up — by at least 0.25 each.

The chain starts with Doubt around average — 0.49, lower than 51 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.28.

At the same time Concentration goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.04 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 21 s · zh · emolia

k 3d_a -0.325d_b 0.282step_a 0.217step_b 0.242min_cos_consec 0.8478min_cos_anchor 0.8639dataset emolialang zhspeaker ZH_B00056_S09022track ZH_B00056_S09022total 20.8slevel spread 2.7 dBmax seam 2.7 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · slightly dark, average recording, normally alert, slightly relaxed
(measured, fairly steady, frequent disfluency, monologue) (low mumble) 但是呢一直忙活着打工啊、做生意啊、赚钱呀,还是没有剩下钱,这很明显是一个问题。但问题的答案是什么呢?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.1/10; 11.2s, ZH.
ZH_B00056_S09022_W000007 · in -20.3 dBFS · gain +0.3 dB · emolia-03834
(anger, contempt, disgust · measured, fairly steady, some disfluency, casual) 就像如果你遇到一扇被锁着的门,那你应该去哪儿找钥匙呢?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, contempt, disgust; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.8/10; 5.9s, ZH.
ZH_B00056_S09022_W000008 · in -17.6 dBFS · gain -2.4 dB · emolia-03834
(slow, steady, no disfluency, formal) 肯定不是只盯着锁去看吧。
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, whispered; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.3/10; 3.4s, ZH.
ZH_B00056_S09022_W000009 · in -19.8 dBFS · gain -0.2 dB · emolia-03834
Contemplation ↓  /  Concentrationsc-AB2-k3 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.30.

At the same time Contemplation goes the other way, from 0.95 (higher than 96 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · en · emolia

k 3d_a -0.356d_b 0.296step_a 0.227step_b 0.153min_cos_consec 0.8546min_cos_anchor 0.8229dataset emolialang enspeaker EN_yZn5hcP4Kkctrack EN_yZn5hcP4Kkctotal 31.6slevel spread 1.7 dBmax seam 0.9 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, quiet background, fairly steady, moderate pitch range, light breath
(contemplation, doubt · measured, subdued, slightly relaxed, monologue) I think, (low mumble) I think (low mumble) there is no escaping the fact that you need a big manual for
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, doubt; style: monologue, didactic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.0/10; 7.1s, EN.
EN_yZn5hcP4Kkc_W000425 · in -16.4 dBFS · gain -3.6 dB · emolia-01635
(normal-paced, normally alert, neutral tension, playful) And a lot of checks and balances have to be done as to mass balance, where are the inerts going, I mean in some of the places it's a mystery.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: playful, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.1/10; 7.8s, EN.
EN_yZn5hcP4Kkc_W000427 · in -17.1 dBFS · gain -2.9 dB · emolia-01635
(concentration, contentment · brisk, normally alert, neutral tension, monologue) We've been often going to landfills which have said that they have done it very well, but we can't find where the inerts go. So if you have a good manual for this, a step-by-step and the design procedures and (ahem) the checks and balances, I think it can be implemented, but without that, it will going to be unrestricted and free for all.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, contentment; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.8/10; 16.4s, EN.
EN_yZn5hcP4Kkc_W000428 · in -18.1 dBFS · gain -1.9 dB · emolia-01635
Concentration ↓  /  Emotional Numbnesssc-AB2-k3 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.25.

At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · emolia

k 3d_a -0.303d_b 0.254step_a 0.238step_b 0.151min_cos_consec 0.8974min_cos_anchor 0.9205dataset emolialang enspeaker EN_ZD6Le_7I7D8track EN_ZD6Le_7I7D8total 29.9slevel spread 3.3 dBmax seam 3.3 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, slightly relaxed
(concentration · measured, steady, frequent disfluency, didactic) Expand the, the, the array and put the elements inside an intermediate position. If you combine the two, you can do all sort, all sort, all sort of replacing of the elements.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.5/10; 12.3s, EN.
EN_ZD6Le_7I7D8_W000378 · in -19.3 dBFS · gain -0.7 dB · emolia-00683
(concentration · normal-paced, fairly steady, some disfluency, casual) (surprised gasp) Just don't make any confusion between slice and splice. (low mumble) Splice modifies the array in place, slice creates a new array. (ahem) You could experiment with this (ahem) very complex function.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 1.9/10; 12.0s, EN.
EN_ZD6Le_7I7D8_W000379 · in -22.3 dBFS · gain +2.3 dB · emolia-00683
(measured, fairly steady, frequent disfluency, casual) It will (ahem) change the order of the element in E, it does not create a new array.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.4/10; 5.4s, EN.
EN_ZD6Le_7I7D8_W000380 · in -19.0 dBFS · gain -1.0 dB · emolia-00683
Contemplation ↓  /  Confusionsc-AB2-k3 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.28.

At the same time Contemplation goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 56 s · zh · emolia

k 3d_a -0.342d_b 0.277step_a 0.183step_b 0.140min_cos_consec 0.8988min_cos_anchor 0.8794dataset emolialang zhspeaker ZH_B00041_S03451track ZH_B00041_S03451total 56.2slevel spread 0.5 dBmax seam 0.3 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, measured, moderate pitch range, light breath
(contemplation, sexual lust, jealousy and envy · subdued, slightly relaxed, fairly steady, monologue) 就更不容易露出马脚,只要知道工作地点的电话公司名称也能得知。你们那里一接到电话,应该直接会说,这里是十路犯罪编辑部吧,还是赤井书房?您好,连违违都不说呢,直接报上史路犯罪。如此一来,你的真实身份就被拆穿了,接下来没什么好说的。
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, sexual lust, jealousy and envy; style: monologue, narration; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.6/10; 28.5s, ZH.
ZH_B00041_S03451_W000007 · in -19.0 dBFS · gain -1.0 dB · emolia-03683
(sourness, jealousy and envy, malevolence malice · subdued, neutral tension, fairly steady, storytelling) 能顺便知道出生地更好,只要说是从老家打来的,亲切的人自然会寒暄上几句。故乡是哪儿也就曝光了,原来如此,可恶。原本以为已经很小心了,没想到还是中了他们的鬼把戏。鸟口似乎很不敢信。
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as sourness, jealousy and envy, malevolence malice; style: storytelling, narration; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 6.9/10; 20.8s, ZH.
ZH_B00041_S03451_W000008 · in -19.2 dBFS · gain -0.8 dB · emolia-03683
(normally alert, slightly relaxed, steady, monologue) 这只是因为实际发生顺序跟正常顺序不一样,所以才不容易注意到。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.5/10; 6.6s, ZH.
ZH_B00041_S03451_W000009 · in -19.4 dBFS · gain -0.6 dB · emolia-03683
Fatigue Exhaustion ↓  /  Doubtsc-AB2-k3 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Doubt up — by at least 0.25 each.

The chain starts with Doubt around average — 0.49, lower than 51 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.38.

At the same time Fatigue Exhaustion goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 17 s · zh · emolia

k 3d_a -0.313d_b 0.383step_a 0.191step_b 0.223min_cos_consec 0.8674min_cos_anchor 0.8674dataset emolialang zhspeaker ZH_B00041_S01249track ZH_B00041_S01249total 17.4slevel spread 2.7 dBmax seam 2.7 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, no background noise, normally alert, slightly relaxed
(normal-paced, narration, formal) 直到天色渐渐放亮的时候,我们才看清昨晚激战之后的战果。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.1/10; 4.9s, ZH.
ZH_B00041_S01249_W000061 · in -22.5 dBFS · gain +2.5 dB · emolia-03691
(infatuation · measured, monologue, narration) 阵地前密密麻麻的到处都是越军的尸体。这些尸体分成两个部分。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: monologue, narration; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.8/10; 5.8s, ZH.
ZH_B00041_S01249_W000062 · in -25.1 dBFS · gain +5.1 dB · emolia-03691
(measured, narration, formal) 一部分是从山脚到阵地前沿最靠近的,离我们阵地不过几米远,很显然。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.6/10; 6.4s, ZH.
ZH_B00041_S01249_W000063 · in -22.8 dBFS · gain +2.8 dB · emolia-03691
Emotional Numbness ↓  /  Infatuationsc-AB2-k3 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.28.

At the same time Emotional Numbness goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.03, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.98 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.98 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.98), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · en · emolia

k 3d_a -0.361d_b 0.276step_a 0.220step_b 0.245min_cos_consec 0.9794min_cos_anchor 0.9794dataset emolialang enspeaker EN_tZQ7dnb0nbItrack EN_tZQ7dnb0nbItotal 27.1slevel spread 2.2 dBmax seam 2.2 dBcos from trajectories_v3_speaker_clean
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness · formal, authoritative) A former Honda executive, Irimajiri had been involved with Sega of America since joining Sega in 1993
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 7.0s, EN.
EN_tZQ7dnb0nbI_W000099 · in -14.4 dBFS · gain -5.6 dB · emolia-00398
(formal, authoritative) Sega also announced that David Rosen and Nakayama had resigned from their positions as chairman and co-chairman of Sega of America, though both remained with the company
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 8.9s, EN.
EN_tZQ7dnb0nbI_W000100 · in -13.7 dBFS · gain -6.3 dB · emolia-00398
(newsreading, formal) Bernie Stoller, a former executive at Sony Computer Entertainment of America, was named Sega of America's executive vice president in charge of product development and third-party relations
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 10.9s, EN.
EN_tZQ7dnb0nbI_W000101 · in -15.9 dBFS · gain -4.1 dB · emolia-00398