k-VN1-k3

VN1 at chain length k=3, all corpora, at the mining floor.

Rule. VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C
Source. trajectories_v5.parquet  |  Family. rule x chain length
Sampled from 672,000 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
DARC — darkness of timbrek-VN1-k3 · #1

This is a VoiceNet dimension, not an emotion: darkness of timbre (DARC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with darkness of timbre (DARC) low — 0.25, lower than 75 % of clips in this corpus — and ends with it above average at 0.61, higher than 61 % of clips in this corpus. That is a total rise of 0.37.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · zh · emolia

hear it un-normalised (raw levels, max seam 1.7 dB)
k 3d_a 0.366d_b 0.366step_a 0.211step_b 0.211min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00078_S08667track ZH_B00078_S08667total 31.3slevel spread 1.7 dBmax seam 1.7 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(thankfulness gratitude, contemplation · measured, some disfluency, somewhat unclear, conversational) (ahem) 嗯,那说到这个阿那亚,我们就来具体聊聊今天的主题啊,就是我其实第一次知道阿那亚喜剧节。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as thankfulness gratitude, contemplation; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.3/10; 8.9s, ZH.
ZH_B00078_S08667_W000014 · in -19.1 dBFS · gain -0.9 dB · emolia-04059
(triumph, interest, astonishment surprise · normal-paced, some disfluency, average clarity, monologue) (ahem) 而且是在那个这个节日刚开始做的那段时间,密集的朋友圈被刷屏。就是有很多朋友在夏天的时候啊,然后去那边你能给大家简单科普一下,比如说这个安那亚戏剧节的缘起,它是一个什么样子的活动,然后这个活动大概一个游玩的方向是什么样子。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, interest, astonishment surprise; style: monologue, authoritative; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 7.6/10; 18.6s, ZH.
ZH_B00078_S08667_W000015 · in -20.0 dBFS · gain -0.0 dB · emolia-04059
(normal-paced, no disfluency, clear, authoritative) 小县城,秦皇岛下面还分几个县。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.2/10; 3.5s, ZH.
ZH_B00078_S08667_W000016 · in -18.2 dBFS · gain -1.8 dB · emolia-04059
S_CONV — style: conversationalk-VN1-k3 · #2

This is a VoiceNet dimension, not an emotion: style: conversational (S_CONV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: conversational (S_CONV) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to above average at 0.64, higher than 64 % of clips in this corpus. That is a total fall of 0.34.

It takes 3 clips to get there. Clip to clip the moves are -0.13, then -0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · en · emolia

hear it un-normalised (raw levels, max seam 1.3 dB)
k 3d_a -0.337d_b -0.337step_a 0.203step_b 0.203min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_XKWJjBC16Ootrack EN_XKWJjBC16Oototal 41.4slevel spread 1.6 dBmax seam 1.3 dB
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · slightly bright, fairly smooth, average recording, quiet background, moderately variable, some disfluency, average clarity, wide pitch range
(sexual lust, embarrassment, pain · normal-paced, normally alert, slightly relaxed, conversational) It's, it's my way (ahem) of, if I can anticipate by talking with a client.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, embarrassment, pain; style: conversational, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.1/10; 4.3s, EN.
EN_XKWJjBC16Oo_W000085 · in -15.9 dBFS · gain -4.1 dB · emolia-02245
(fear, confusion, impatience and irritability · brisk, energised, neutral tension, casual) What, what they're afraid of, are they afraid they're going to forget to do something? Okay. Then we talk about tickler files, which I know you and I are going to talk about later. If what they're afraid of is that they won't be prompted to do something that the paper won't, you (low mumble) know, I worked with a client yesterday.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, confusion, impatience and irritability; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 8.6/10; 17.9s, EN.
EN_XKWJjBC16Oo_W000086 · in -17.2 dBFS · gain -2.8 dB · emolia-02245
(affection, interest, sourness · normal-paced, energised, neutral tension, casual) Who said, you know, it's time for her to look at the, (low mumble) uhm, the new Medicare Part D, the, the, the pharmacy, (low mumble) uh, information because, you know, you re-up that every, every October and she's like, she keeps getting all of this mail and she just sort of tosses it in the pile because what she's afraid of
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as affection, interest, sourness; style: casual, dramatic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 6.0/10; 18.9s, EN.
EN_XKWJjBC16Oo_W000087 · in -17.5 dBFS · gain -2.5 dB · emolia-02245
RESP — audible breath / respirationk-VN1-k3 · #3

This is a VoiceNet dimension, not an emotion: audible breath / respiration (RESP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with audible breath / respiration (RESP) around average — 0.46, lower than 54 % of clips in this corpus — and works its way down to low at 0.14, lower than 86 % of clips in this corpus. That is a total fall of 0.32.

It takes 3 clips to get there. Clip to clip the moves are -0.23, then -0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.22 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.22 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 16 s · ko · emolia

hear it un-normalised (raw levels, max seam 0.9 dB)
k 3d_a -0.319d_b -0.319step_a 0.225step_b 0.225min_cos_consec min_cos_anchor dataset emolialang kospeaker KO_d6bZr1-iwhEtrack KO_d6bZr1-iwhEtotal 16.4slevel spread 1.2 dBmax seam 0.9 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, light breath
(normal-paced, fairly steady, no disfluency, narration) 지금까지 6천회가 넘는 여진이 이어지면서 트리키의 남동부 10개 지역이 초토화됐습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.7/10; 6.5s, KO.
KO_d6bZr1-iwhE_W000010 · in -17.9 dBFS · gain -2.1 dB · emolia-03095
(impatience and irritability, confusion, anger · fast, moderately variable, some disfluency, dramatic) 이 밑에 지금 다 무너져 있습니다. 다 무너져 있어. 지금 차 밑으로도 다 깔려 있고.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, confusion, anger; style: dramatic, storytelling; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.1/10; 5.1s, KO.
KO_d6bZr1-iwhE_W000011 · in -17.5 dBFS · gain -2.5 dB · emolia-03095
(pride · normal-paced, fairly steady, no disfluency, authoritative) 지진이 발생한 건 동의특이도전인 새벽 4시 17분경.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: authoritative, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.1/10; 4.5s, KO.
KO_d6bZr1-iwhE_W000012 · in -16.7 dBFS · gain -3.3 dB · emolia-03095
ROUG — roughnessk-VN1-k3 · #4

This is a VoiceNet dimension, not an emotion: roughness (ROUG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with roughness (ROUG) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to low at 0.24, lower than 76 % of clips in this corpus. That is a total fall of 0.41.

It takes 3 clips to get there. Clip to clip the moves are -0.25, then -0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.21 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.21 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 17 s · en · emolia

hear it un-normalised (raw levels, max seam 5.1 dB)
k 3d_a -0.406d_b -0.406step_a 0.248step_b 0.248min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_QHEOpBf3_yAtrack EN_QHEOpBf3_yAtotal 17.1slevel spread 5.2 dBmax seam 5.1 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, some disfluency
(embarrassment, shame · steady, clear, monologue, formal) To the main repository, which basically says here, I've done this change. Do you want it?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, shame; style: monologue, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.7/10; 5.2s, EN.
EN_QHEOpBf3_yA_W000055 · in -22.8 dBFS · gain +2.8 dB · emolia-01537
(doubt · fairly steady, average clarity, playful, authoritative) And even if it's still under development, uh, (low mumble) so while it's still in the stage of a pull request, others can already use it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: playful, authoritative; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.4/10; 6.5s, EN.
EN_QHEOpBf3_yA_W000056 · in -17.7 dBFS · gain -2.3 dB · emolia-01537
(fairly steady, average clarity, casual, playful) So if there's, uh, (low mumble) so also occasionally we'll have on the forums and someone with the same problem.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.2/10; 5.1s, EN.
EN_QHEOpBf3_yA_W000057 · in -17.6 dBFS · gain -2.4 dB · emolia-01537
S_CONV — style: conversationalk-VN1-k3 · #5

This is a VoiceNet dimension, not an emotion: style: conversational (S_CONV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: conversational (S_CONV) below average — 0.40, lower than 60 % of clips in this corpus — and works its way down to low at 0.08, lower than 92 % of clips in this corpus. That is a total fall of 0.32.

It takes 3 clips to get there. Clip to clip the moves are -0.18, then -0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.08 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.08 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · emolia

hear it un-normalised (raw levels, max seam 7.5 dB)
k 3d_a -0.316d_b -0.316step_a 0.181step_b 0.181min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_mnLBLckMmPytrack EN_mnLBLckMmPytotal 30.4slevel spread 7.5 dBmax seam 7.5 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, slightly relaxed, fairly steady
(measured, normally alert, frequent disfluency, casual) (low mumble) Uh, very rundown building and there, the buildings around it weren't much better.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.2/10; 5.2s, EN.
EN_mnLBLckMmPy_W000008 · in -23.0 dBFS · gain +3.0 dB · emolia-01331
(contentment, thankfulness gratitude · measured, very low-energy, frequent disfluency, monologue) But one of our, uh, (low mumble) United Properties taglines is we build communities. And we've really viewed this as the start of a, (low mumble) of a community. It's a large building, Ford Center, and, (low mumble) uh, we knew we could re-tenant it, (low mumble) uhm, you know, with some good tenants and hopefully expand from there.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude; style: monologue, casual; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 0.7/10; 18.5s, EN.
EN_mnLBLckMmPy_W000009 · in -25.0 dBFS · gain +5.0 dB · emolia-01331
(triumph, astonishment surprise · normal-paced, normally alert, no disfluency, narration) With no track record in historic renovation, turning old into new was a risky but exciting undertaking.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as triumph, astonishment surprise; style: narration, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.5s, EN.
EN_mnLBLckMmPy_W000010 · in -17.5 dBFS · gain -2.5 dB · emolia-01331
ATCK — attack / onset sharpnessk-VN1-k3 · #6

This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with attack / onset sharpness (ATCK) high — 0.88, higher than 88 % of clips in this corpus — and works its way down to around average at 0.52, higher than 52 % of clips in this corpus. That is a total fall of 0.36.

It takes 3 clips to get there. Clip to clip the moves are -0.15, then -0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · de · emolia

k 3d_a -0.359d_b -0.359step_a 0.214step_b 0.214min_cos_consec min_cos_anchor dataset emolialang despeaker DE_Q1avxQI72lytrack DE_Q1avxQI72lytotal 18.0slevel spread 1.5 dBmax seam 1.5 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, newsreading) Regionale Menschenrechtssysteme, vor allem in Afrika und Amerika, müssen gestärkt werden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 3.9s, DE.
DE_Q1avxQI72ly_W000030 · in -17.6 dBFS · gain -2.4 dB · emolia-00255
(disgust · formal, newsreading) In Weltregionen wie Asien oder dem Nahen Osten gibt es bisher überhaupt keine regionalen Menschenrechtssysteme. Diese müssen gegründet werden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 6.2s, DE.
DE_Q1avxQI72ly_W000031 · in -19.1 dBFS · gain -0.9 dB · emolia-00255
(disappointment · formal, monologue) Die einzelnen Länder und die internationale Gemeinschaft müssen sich mehr anstrengen Menschenrechtsverletzungen zu verhindern, die durch nichtstaatliche Akteure verursacht werden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 7.5s, DE.
DE_Q1avxQI72ly_W000032 · in -17.9 dBFS · gain -2.1 dB · emolia-00255
CLRT — clarity / intelligibilityk-VN1-k3 · #7

This is a VoiceNet dimension, not an emotion: clarity / intelligibility (CLRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with clarity / intelligibility (CLRT) around average — 0.53, higher than 53 % of clips in this corpus — and works its way down to low at 0.20, lower than 80 % of clips in this corpus. That is a total fall of 0.33.

It takes 3 clips to get there. Clip to clip the moves are -0.19, then -0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 43 s · en · emolia

k 3d_a -0.331d_b -0.331step_a 0.193step_b 0.193min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_B00066_S05348track EN_B00066_S05348total 42.6slevel spread 0.7 dBmax seam 0.6 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, balanced body, average recording, quiet background, slightly relaxed, fairly steady
(normal-paced, normally alert, some disfluency, didactic) A (ahem) parameter that is been created essentially a thermodynamic parameter to stand for this pressure ratio.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.6/10; 6.6s, EN.
EN_B00066_S05348_W000006 · in -18.2 dBFS · gain -1.8 dB · emolia-01517
(concentration · measured, subdued, frequent disfluency, didactic) So we will be talking a little about multi spooling also in today's class as part of multi staging and we shall also talk a little about how this multi spool arrangement of turbines
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.7/10; 16.4s, EN.
EN_B00066_S05348_W000007 · in -18.8 dBFS · gain -1.2 dB · emolia-01517
(concentration · measured, subdued, frequent disfluency, monologue) (ahem) (low mumble) Let us look at some of these (ahem) fundamental considerations that go into, (low mumble) uh, adoption or selection of multistage, (low mumble) uh, configuration for axial flow turbines.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.8/10; 19.3s, EN.
EN_B00066_S05348_W000008 · in -18.9 dBFS · gain -1.1 dB · emolia-01517
S_TECH — style: technicalk-VN1-k3 · #8

This is a VoiceNet dimension, not an emotion: style: technical (S_TECH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: technical (S_TECH) at the very bottom of the range — 0.03, lower than 97 % of clips in this corpus — and ends with it below average at 0.37, lower than 63 % of clips in this corpus. That is a total rise of 0.34.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.33 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.33 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 12 s · zh · emolia

k 3d_a 0.343d_b 0.343step_a 0.226step_b 0.226min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00010_S03614track ZH_B00010_S03614total 12.2slevel spread 1.7 dBmax seam 1.7 dB
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, average clarity, light breath
(sexual lust, amusement, teasing · normal-paced, moderately variable, some disfluency, casual) 浪费啊,但我不喜欢你来帮我浪费。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, amusement, teasing; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 0.6/10; 3.5s, ZH.
ZH_B00010_S03614_W000007 · in -17.7 dBFS · gain -2.3 dB · emolia-03376
(normal-paced, moderately variable, frequent disfluency, casual) (chuckle) 明白明白了,我记得我们就是那次聊天的时候。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.4/6; vocal-burst blend 4.6/10; 4.9s, ZH.
ZH_B00010_S03614_W000008 · in -19.4 dBFS · gain -0.6 dB · emolia-03376
(fast, fairly steady, no disfluency, casual) 你给了我一个,我不知道你记不记得了啊嗯就一七年那次。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.1/10; 3.5s, ZH.
ZH_B00010_S03614_W000009 · in -18.0 dBFS · gain -2.0 dB · emolia-03376
R_MIXD — resonance: mixedk-VN1-k3 · #9

This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: mixed (R_MIXD) around average — 0.49, lower than 51 % of clips in this corpus — and ends with it high at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.34.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · en · emolia

k 3d_a 0.341d_b 0.341step_a 0.229step_b 0.229min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_vMotqa789tgtrack EN_vMotqa789tgtotal 22.4slevel spread 5.1 dBmax seam 4.2 dB
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, normally alert, fairly steady, average clarity, light breath
(infatuation · brisk, neutral tension, some disfluency, casual) You may also create a (low mumble) hello application and you may also include it and ship it, but that will create a problem. Then directory listing is not disabled on the server.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as infatuation; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.3/10; 9.9s, EN.
EN_vMotqa789tg_W000090 · in -18.4 dBFS · gain -1.6 dB · emolia-00504
(fast, slightly relaxed, some disfluency, formal) Then the application's servers configuration allows detailed error messages.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.3/10; 3.8s, EN.
EN_vMotqa789tg_W000091 · in -19.2 dBFS · gain -0.8 dB · emolia-00504
(measured, slightly relaxed, frequent disfluency, didactic) Then a cloud service provider has default, uh, (low mumble) sharing permission open to the internet by other CSP users.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.0/10; 8.3s, EN.
EN_vMotqa789tg_W000092 · in -23.5 dBFS · gain +3.5 dB · emolia-00504
R_MIXD — resonance: mixedk-VN1-k3 · #10

This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: mixed (R_MIXD) at the very bottom of the range — 0.08, lower than 92 % of clips in this corpus — and ends with it around average at 0.53, higher than 53 % of clips in this corpus. That is a total rise of 0.45.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 24 s · en · podcast

k 3d_a 0.450d_b 0.450step_a 0.242step_b 0.242min_cos_consec 0.8075min_cos_anchor 0.8035dataset podcastlang enspeaker 840502track 840502total 23.7slevel spread 3.3 dBmax seam 3.3 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, fairly smooth, balanced body, moderately variable, wide pitch range, light breath
(intoxication altered states of consciousness, confusion, doubt · slow, normally alert, slightly relaxed, conversational) 6 a.m. Do I stay up? Do I take a nap? I
full caption & clip details
A child masculine voice; delivery is normally alert, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, confusion, doubt; style: conversational, casual; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.9/10; 5.0s, EN.
840502_00050384 · in -31.6 dBFS · gain +11.6 dB · podcast-02221
(confusion, doubt, amusement · normal-paced, normally alert, neutral tension, casual) don't know. Like it all depends on the day. It all depends on how I feel. Like maybe I will take a nap. And
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as confusion, doubt, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 5.4/10; 5.8s, EN.
840502_00050880 · in -28.3 dBFS · gain +8.3 dB · podcast-02231
(doubt, fatigue exhaustion, infatuation · brisk, energised, neutral tension, casual) if I take a nap, where does that put me at for the rest of the day? Does that is that gonna make it hard for me to fall asleep at what we call the normal time, which is night, you know, where everyone else goes to sleep? Or am I gonna
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as doubt, fatigue exhaustion, infatuation; style: casual, conversational; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 6.9/10; 12.6s, EN.
840502_00051464 · in -30.2 dBFS · gain +10.2 dB · podcast-02219
SMTH — smoothnessk-VN1-k3 · #11

This is a VoiceNet dimension, not an emotion: smoothness (SMTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with smoothness (SMTH) high — 0.77, higher than 77 % of clips in this corpus — and works its way down to below average at 0.38, lower than 62 % of clips in this corpus. That is a total fall of 0.39.

It takes 3 clips to get there. Clip to clip the moves are -0.20, then -0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.64 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.65 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.64, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · es · podcast

k 3d_a -0.395d_b -0.395step_a 0.203step_b 0.203min_cos_consec 0.6459min_cos_anchor 0.6374dataset podcastlang esspeaker 4097track 4097total 41.0slevel spread 0.6 dBmax seam 0.6 dBcos from orange-id (speaker identity)
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, slightly relaxed, fairly steady, average clarity, moderate pitch range
(fear, doubt · brisk, normally alert, some disfluency, monologue) Y podemos hablar también de un suceso histórico que le tocó vivir, podríamos decir, que fue el llamado motín del pan de mil setecientos sesenta y seis de Zaragoza, que bueno, que fue bastante importante, ¿verdad, Santiago?
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, doubt; style: monologue, authoritative; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 5.0/10; 10.3s, ES.
4097_00121572 · in -25.0 dBFS · gain +5.0 dB · podcast-05917
(infatuation, triumph, sexual lust · normal-paced, subdued, some disfluency, monologue) Sí, sí. Vamos a explicarlo esto un poquito en detalle, ¿no? Que yo creo que lo merece para que se pueda entender bien. En esa etapa de deformación, y aunque no sabemos si a Francisco de Goya, pues le pilló en esos momentos en Zaragoza o no. El caso es que en mil setecientos sesenta y seis se vivió en la capital aragonesa uno de los momentos más sangrientos de su historia, que es el motín del pan, como decía Sergio, de ese año 176.
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, triumph, sexual lust; style: monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 9.9/10; 26.2s, ES.
4097_00122632 · in -25.6 dBFS · gain +5.6 dB · podcast-05924
(normal-paced, normally alert, no disfluency, conversational) Motín del pan, o de los broqueleros, que también se llama aquí, ¿no?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, playful; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.3/10; 4.2s, ES.
4097_00125252 · in -25.5 dBFS · gain +5.5 dB · podcast-05927
COGL — cognitive loadk-VN1-k3 · #12

This is a VoiceNet dimension, not an emotion: cognitive load (COGL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with cognitive load (COGL) below average — 0.40, lower than 60 % of clips in this corpus — and ends with it above average at 0.64, higher than 64 % of clips in this corpus. That is a total rise of 0.25.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · en · eurospeech

k 3d_a 0.248d_b 0.248step_a 0.128step_b 0.128min_cos_consec min_cos_anchor dataset eurospeechlang enspeaker uk_uk_29_07032016track uk_uk_29_07032016total 40.1slevel spread 0.7 dBmax seam 0.7 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged feminine voice · slightly cool, fairly smooth, balanced body, normal-paced, clear, wide pitch range, light breath
(energised, slightly relaxed, moderately variable, dramatic) and by the ability to escalate the decision to the Secretary of State if there is disagreement, with an independent review panel assessing the business case.
full caption & clip details
A middle-aged feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.4/10; 11.3s, EN.
uk_uk_29_07032016_16397248_16408576 · in -24.4 dBFS · gain +4.4 dB · eurospeech-00996
(concentration, infatuation · normally alert, slightly relaxed, moderately variable, formal) When the Minister winds up, I would be interested to hear in more detail how those two processes will work in practice. Although I accept that there is a need to co-operate, the processes need to be genuinely robust to address the underlying resistance to change.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, infatuation; style: formal, authoritative; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.6/10; 18.2s, EN.
uk_uk_29_07032016_16408576_16426735 · in -23.6 dBFS · gain +3.6 dB · eurospeech-00996
(doubt · energised, neutral tension, fairly steady, authoritative) I would also be interested to know how frequently the reviews could be undertaken, should there be a need to revisit a business case. I have
full caption & clip details
A middle-aged feminine voice; delivery is energised, normal-paced, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as doubt; style: authoritative, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.4/10; 10.3s, EN.
uk_uk_29_07032016_16426735_16437024 · in -23.7 dBFS · gain +3.7 dB · eurospeech-00996
SMTH — smoothnessk-VN1-k3 · #13

This is a VoiceNet dimension, not an emotion: smoothness (SMTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with smoothness (SMTH) low — 0.21, lower than 79 % of clips in this corpus — and ends with it around average at 0.53, higher than 53 % of clips in this corpus. That is a total rise of 0.32.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.48 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.49 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.48, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · en · podcast

k 3d_a 0.317d_b 0.317step_a 0.190step_b 0.190min_cos_consec 0.4930min_cos_anchor 0.4834dataset podcastlang enspeaker 413599track 413599total 18.0slevel spread 4.3 dBmax seam 2.2 dBcos from orange-id (speaker identity)
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · slightly bright, quiet background, neutral tension
(teasing, elation, doubt · fast, energised, volatile, casual) They're like, Are you sure we want to do this? Yeah, keep filming. I know, but it's the season finale. No.
full caption & clip details
A child feminine voice; delivery is energised, fast, neutral tension, volatile; timbre is slightly cool, slightly bright, very rough, thin; slurred, some disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, very vulnerable; reads as teasing, elation, doubt; style: casual, playful; below-average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 4.5/10; 5.5s, EN.
413599_00095072 · in -27.5 dBFS · gain +7.5 dB · podcast-05368
(embarrassment, intoxication altered states of consciousness, elation · brisk, energised, volatile, casual) Like every five minutes. Should I be doing this? (breathy giggle) Keep going. It's great. We'll we'll fix it in post. But then I don't think they did. Alright, so
full caption & clip details
A young adult somewhat masculine voice; delivery is energised, brisk, neutral tension, volatile; timbre is slightly cool, slightly bright, very rough, slightly thin; slurred, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, vulnerable; reads as embarrassment, intoxication altered states of consciousness, elation; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 3.8/6; vocal-burst blend 1.5/10; 8.7s, EN.
413599_00095976 · in -29.6 dBFS · gain +9.6 dB · podcast-05361
(relief, pleasure ecstasy, triumph · normal-paced, normally alert, moderately variable, casual) Oh, finally it ends. I wrote Buns, Buns, Buns, thank God it's over. Okay, yeah.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as relief, pleasure ecstasy, triumph; style: casual, conversational; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.6/10; 3.5s, EN.
413599_00097024 · in -31.9 dBFS · gain +11.9 dB · podcast-05373
S_STRY — style: storytellingk-VN1-k3 · #14

This is a VoiceNet dimension, not an emotion: style: storytelling (S_STRY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: storytelling (S_STRY) low — 0.10, lower than 90 % of clips in this corpus — and ends with it around average at 0.44, lower than 56 % of clips in this corpus. That is a total rise of 0.34.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · en · emolia

k 3d_a 0.344d_b 0.344step_a 0.217step_b 0.217min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_0X0QvauOToMtrack EN_0X0QvauOToMtotal 29.0slevel spread 1.2 dBmax seam 0.9 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(contemplation, confusion, concentration · little disfluency, monologue, authoritative) and goes on to explain that the external form does not always indicate the truth of the internal being a person's nature cannot be known by their appearance
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, confusion, concentration; style: monologue, authoritative; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.2/10; 8.5s, EN.
EN_0X0QvauOToM_W000017 · in -18.2 dBFS · gain -1.8 dB · emolia-00461
(almost no disfluency, dramatic, monologue) This is the first of many distinctions between the external and the internal form that will be made in this quest. So put that in your pocket for now.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: dramatic, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 7.9s, EN.
EN_0X0QvauOToM_W000018 · in -19.1 dBFS · gain -0.9 dB · emolia-00461
(sadness, disappointment, fear · almost no disfluency, formal, monologue) We learn that additionally to the void left in the elemental beings after the purge of forbidden knowledge, their home had also been destroyed. Those who survived the apocalypse adapted their external form and became fungi.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sadness, disappointment, fear; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 12.3s, EN.
EN_0X0QvauOToM_W000019 · in -19.4 dBFS · gain -0.6 dB · emolia-00461
BRGT — brightness of timbrek-VN1-k3 · #15

This is a VoiceNet dimension, not an emotion: brightness of timbre (BRGT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with brightness of timbre (BRGT) above average — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the range at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.31.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 25 s · en · emolia

k 3d_a 0.314d_b 0.314step_a 0.173step_b 0.173min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_UtYlXz0zZjctrack EN_UtYlXz0zZjctotal 25.1slevel spread 1.4 dBmax seam 1.4 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed, clear
(steady, almost no disfluency, formal, newsreading) Although a relatively new technology, PEF has been successfully used in both food decontamination processes as well as wastewater treatments.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 9.4s, EN.
EN_UtYlXz0zZjc_W000442 · in -15.1 dBFS · gain -4.9 dB · emolia-01017
(fairly steady, no disfluency, formal, newsreading) Origin Oils Inc has been researching a revolutionary method called the Helix Bioreactor, altering the common closed-loop growth system
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.0s, EN.
EN_UtYlXz0zZjc_W000444 · in -15.3 dBFS · gain -4.7 dB · emolia-01017
(emotional numbness · steady, no disfluency, formal, monologue) This system utilizes low-energy lights in a helical pattern, enabling each algal cell to obtain the required amount of light
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.5s, EN.
EN_UtYlXz0zZjc_W000445 · in -13.9 dBFS · gain -6.1 dB · emolia-01017
R_NASL — resonance: nasalk-VN1-k3 · #16

This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with resonance: nasal (R_NASL) above average — 0.67, higher than 67 % of clips in this corpus — and works its way down to below average at 0.34, lower than 66 % of clips in this corpus. That is a total fall of 0.33.

It takes 3 clips to get there. Clip to clip the moves are -0.21, then -0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · en · emolia

k 3d_a -0.328d_b -0.328step_a 0.214step_b 0.214min_cos_consec min_cos_anchor dataset emolialang enspeaker EN_doxSLEGp-bUtrack EN_doxSLEGp-bUtotal 26.8slevel spread 1.1 dBmax seam 1.1 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(doubt · casual, conversational) That they're actually in the process of spreading apart. Now I'm, I'm wildly extrapolating from what I read here, by the way, but they could well be starting to sort of fly apart. Maybe.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.6/10; 9.1s, EN.
EN_doxSLEGp-bU_W000268 · in -16.8 dBFS · gain -3.2 dB · emolia-00822
(confusion, helplessness · casual, conversational) It's dark matter has been stolen. Maybe it's dark matter has been removed somehow or.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion, helplessness; style: casual, conversational; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.7/10; 4.4s, EN.
EN_doxSLEGp-bU_W000269 · in -17.5 dBFS · gain -2.5 dB · emolia-00822
(doubt, confusion, contemplation · conversational, casual) We don't know, I mean we don't really know what it is, so we can't tell you how it would come or go, but maybe what we're seeing matches to some extent the logic, because if the dark matter weren't there, then we would expect it to be diffuse, because it would actually be
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as doubt, confusion, contemplation; style: conversational, casual; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.1/10; 13.0s, EN.
EN_doxSLEGp-bU_W000270 · in -16.5 dBFS · gain -3.5 dB · emolia-00822
TEMP — tempok-VN1-k3 · #17

This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with tempo (TEMP) below average — 0.37, lower than 63 % of clips in this corpus — and works its way down to low at 0.09, lower than 91 % of clips in this corpus. That is a total fall of 0.28.

It takes 3 clips to get there. Clip to clip the moves are -0.18, then -0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · zh · emolia

k 3d_a -0.275d_b -0.275step_a 0.176step_b 0.176min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00046_S00573track ZH_B00046_S00573total 22.3slevel spread 3.0 dBmax seam 3.0 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · no background noise, very low-energy
(affection, pleasure ecstasy, contentment · slow, relaxed, moderately variable, storytelling) 门上挂着一个大大的黑框牌子,上面写着。
full caption & clip details
A child feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, bright, smooth, balanced body; slurred, no disfluency, very wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as affection, pleasure ecstasy, contentment; style: storytelling, cartoonish; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 3.3/10; 5.3s, ZH.
ZH_B00046_S00573_W000020 · in -17.4 dBFS · gain -2.6 dB · emolia-03735
(pain · measured, slightly relaxed, fairly steady, formal) 卡斯佩尔的肚子里一阵绞痛。
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.9/10; 3.4s, ZH.
ZH_B00046_S00573_W000021 · in -15.9 dBFS · gain -4.0 dB · emolia-03735
(confusion, longing, astonishment surprise · measured, relaxed, moderately variable, whispered) 这是因为害怕呢,还是因为刚才吃的酸黄瓜和脱脂牛奶在作怪,我是不是应该回头呢?这时候。
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly warm, slightly dark, fairly smooth, slightly thin; clear, little disfluency, very wide pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as confusion, longing, astonishment surprise; style: whispered, narration; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.0/10; 13.3s, ZH.
ZH_B00046_S00573_W000022 · in -19.0 dBFS · gain -1.0 dB · emolia-03735
METL — metallic qualityk-VN1-k3 · #18

This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with metallic quality (METL) around average — 0.57, higher than 57 % of clips in this corpus — and works its way down to low at 0.24, lower than 76 % of clips in this corpus. That is a total fall of 0.33.

It takes 3 clips to get there. Clip to clip the moves are -0.16, then -0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · zh · emolia

k 3d_a -0.329d_b -0.329step_a 0.174step_b 0.174min_cos_consec min_cos_anchor dataset emolialang zhspeaker ZH_B00062_S06690track ZH_B00062_S06690total 23.5slevel spread 1.6 dBmax seam 1.2 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(moderate pitch range, monologue, formal) 他的言语为太后的旧情人隆科多刷了一波存在感。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.2/10; 4.8s, ZH.
ZH_B00062_S06690_W000019 · in -17.3 dBFS · gain -2.7 dB · emolia-03892
(sexual lust, infatuation, concentration · fairly narrow pitch, whispered, monologue) 但是太后并不知道,皇帝对于太后和隆科多之间的那点事情是心知肚明,所以皇帝一声也没吭。太后此行是来干嘛呢?她是来催着皇帝敲定选秀的事情。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, infatuation, concentration; style: whispered, monologue; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.8/10; 15.3s, ZH.
ZH_B00062_S06690_W000020 · in -18.5 dBFS · gain -1.5 dB · emolia-03892
(moderate pitch range, formal, monologue) 太后的语言艺术也是非常值得我们学习的。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 4.4/10; 3.1s, ZH.
ZH_B00062_S06690_W000021 · in -18.9 dBFS · gain -1.1 dB · emolia-03892
DFLU — disfluencyk-VN1-k3 · #19

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with disfluency (DFLU) around average — 0.43, lower than 57 % of clips in this corpus — and ends with it high at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.39.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.22 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.22, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · en · podcast

k 3d_a 0.390d_b 0.390step_a 0.212step_b 0.212min_cos_consec 0.1930min_cos_anchor 0.2202dataset podcastlang enspeaker 115808track 115808total 50.1slevel spread 2.9 dBmax seam 1.5 dBcos from orange-id (speaker identity)
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, average clarity
(interest, contentment, hope enthusiasm optimism · normal-paced, normally alert, slightly relaxed, casual) Yes, absolutely. (ahem) Um, but it's definitely a glimpse of how Jesus was presenting God to the world. Gonna move on to the next theme that we like to look at. Is this a promise fulfilled or a promise being made in this passage? And (ahem) um I'm gonna jump right in here and talk go.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, contentment, hope enthusiasm optimism; style: casual, monologue; good recording, no background noise; mildly explicit content; genuineness 2.8/6; vocal-burst blend 3.2/10; 16.9s, EN.
115808_00066096 · in -30.8 dBFS · gain +10.8 dB · podcast-03911
(fear, infatuation, contemplation · normal-paced, very low-energy, slightly relaxed, casual) Wait, wait, what is it? Is it a promise fulfilled or being (ahem) made?
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as fear, infatuation, contemplation; style: casual, conversational; good recording, quiet background; mildly explicit content; genuineness 3.2/6; vocal-burst blend 3.4/10; 26.7s, EN.
115808_00067840 · in -29.4 dBFS · gain +9.4 dB · podcast-06451
(doubt, embarrassment, confusion · measured, normally alert, relaxed, casual) Well said. I don't know if I see another specific promise in here, but this is kind of the big one.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, embarrassment, confusion; style: casual, conversational; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 3.0/10; 6.2s, EN.
115808_00083028 · in -27.9 dBFS · gain +7.9 dB · podcast-03896
S_DRAM — style: dramatick-VN1-k3 · #20

This is a VoiceNet dimension, not an emotion: style: dramatic (S_DRAM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.

The chain starts with style: dramatic (S_DRAM) above average — 0.72, higher than 72 % of clips in this corpus — and works its way down to below average at 0.32, lower than 68 % of clips in this corpus. That is a total fall of 0.40.

It takes 3 clips to get there. Clip to clip the moves are -0.15, then -0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.53 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 76 s · en · podcast

k 3d_a -0.402d_b -0.402step_a 0.248step_b 0.248min_cos_consec 0.5274min_cos_anchor 0.5274dataset podcastlang enspeaker 494937track 494937total 75.8slevel spread 1.4 dBmax seam 1.4 dBcos from orange-id (speaker identity)
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, light breath
(sexual lust, contentment, amusement · brisk, normally alert, slightly relaxed, casual) My lips are sealed, unfortunately. (low mumble) Uh well, yeah, contractually obligated to keep that quiet. But (ahem) um no, that's uh okay, that's good to hear that you've got most of this settled down. I guess just to finally get you on the record and to say overall how confident do you feel and and how closely you've nailed this and how much you think you're gonna impress Steve with your response. I'll say I'll take a grade or a percentage or anything (chuckle) else.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as sexual lust, contentment, amusement; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.3/10; 28.4s, EN.
494937_00081640 · in -16.3 dBFS · gain -3.7 dB · podcast-04382
(elation, amusement, infatuation · normal-paced, very low-energy, relaxed, casual) A (childlike giggle) (low mumble) very (low mumble) lawyery (low mumble) response (low mumble) from (chuckle) all of you, caveating and qualifying your your (low mumble) uh your responses.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as elation, amusement, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 4.8/10; 25.0s, EN.
494937_00085524 · in -17.6 dBFS · gain -2.4 dB · podcast-00397
(interest, thankfulness gratitude, pleasure ecstasy · brisk, normally alert, slightly relaxed, casual) But (ahem) um no, I think it's time to to to bring Steve in and for you guys to have a bit of a discussion as to the the the the themes and the topics about this transaction and obviously for you guys to hear what he's been telling me and and his perspective as well. Okay, Steve, welcome back. I've just had a great discussion (low mumble) uh with the trainees. They've been sharing their their thoughts about this case, and it was really interesting to hear your perspective on that and how closely or not they've they've sort of mirrored each other in the thought
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as interest, thankfulness gratitude, pleasure ecstasy; style: casual, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 9.1/10; 22.1s, EN.
494937_00088024 · in -17.5 dBFS · gain -2.5 dB · podcast-04374