Manifest tier. voicenet, rule VN1, T=0.2, step cap 0.25. Population 430,993,195 chains (4,330,452 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 330,029,403.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.voicenet__VN1__T0.20__C0.25__INTERNAL — population 430,993,195 chains (4,330,452 h). SHAREABLE variant: 330,029,403. Filter.rule=='VN1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.2 and step_b<=0.25 Sampled from 2,688,000 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is a VoiceNet dimension, not an emotion: style: cartoonish (S_CART) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: cartoonish (S_CART) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to around average at 0.43, lower than 57 % of clips in this corpus. That is a total fall of 0.55.
It takes 5 clips to get there. Clip to clip the moves are -0.09, then -0.11, then -0.18, then -0.17 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.48 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.48 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 37 s · en · emolia
hear it un-normalised (raw levels, max seam 4.1 dB)
k 5d_a -0.549d_b -0.549step_a 0.179step_b 0.179min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_mljhpQksEJQtrack EN_mljhpQksEJQtotal 37.4slevel spread 4.2 dBmax seam 4.1 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice
(distress, astonishment surprise, impatience and irritability · brisk, highly aroused, tense, casual)What the fuck did you do? You must be Johnny. Who are you? My name is not important.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, volatile; timbre is slightly cool, neutral-bright, very rough, thin; average clarity, some disfluency, wide pitch range, audible breath; affect is deeply negative, very dominant, guarded; reads as distress, astonishment surprise, impatience and irritability; style: casual, dramatic; below-average recording, some background noise; mildly explicit content; genuineness 3.4/6; vocal-burst blend 1.8/10; 6.3s, EN.
EN_mljhpQksEJQ_W000207 · in -13.7 dBFS · gain -6.3 dB · emolia-00656
(impatience and irritability, jealousy and envy, contempt·measured, subdued, slightly tense, storytelling)And I'm not in the dude business, dude. You either do it or junkie gets killed.
full caption & clip details
A middle-aged strongly masculine voice; delivery is subdued, measured, slightly tense, steady; timbre is warm, slightly dark, rough, very full; very clear, almost no disfluency, fairly narrow pitch, normal breath; affect is negative, dominant, fairly guarded; reads as impatience and irritability, jealousy and envy, contempt; style: storytelling, authoritative; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.1/10; 6.2s, EN.
EN_mljhpQksEJQ_W000210 · in -17.8 dBFS · gain -2.2 dB · emolia-00656
(triumph·normal-paced, normally alert, slightly relaxed, storytelling)Not difficult decision, even for a man stuck in 1960s time warp.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph; style: storytelling, authoritative; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.8/10; 4.4s, EN.
EN_mljhpQksEJQ_W000211 · in -17.9 dBFS · gain -2.1 dB · emolia-00656
(teasing, bitterness, disappointment·measured, very low-energy, neutral tension, casual)And this will pay off our debts? Well, (exhausted groan) it pays off interest. Wonderful. The name of the man we want is Roman Balik. Yep.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as teasing, bitterness, disappointment; style: casual, monologue; below-average recording, quiet background; mildly explicit content; genuineness 3.3/6; vocal-burst blend 2.2/10; 16.8s, EN.
EN_mljhpQksEJQ_W000213 · in -17.1 dBFS · gain -2.9 dB · emolia-00656
(emotional numbness· measured, normally alert, slightly relaxed, casual)He runs a cab business, but hangs around some.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: casual, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.8/10; 3.0s, EN.
EN_mljhpQksEJQ_W000214 · in -17.1 dBFS · gain -2.9 dB · emolia-00656
This is a VoiceNet dimension, not an emotion: roughness (ROUG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with roughness (ROUG) below average — 0.41, lower than 59 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.36.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.18 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 60 s · en · podcast
hear it un-normalised (raw levels, max seam 2.9 dB)
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, quiet background, light breath
(hope enthusiasm optimism · normal-paced, subdued, slightly relaxed, casual)So I I think I want to call it now. (low mumble) Um, but I think that while all us Eagles fans think they're gonna end up taking a receiver. That Howie is going to try to galaxy brain some shit and end up drafting another quarterback after trading Wentz, and it's going to be a QB competition going into the year, or they're going to trade Hertz to some weird fucking situation because that's just how he is. He thinks he's smarter than everyone else, and he thinks
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 8.1/10; 25.2s, EN.
903150_00258608 · in -29.0 dBFS · gain +9.0 dB · podcast-06373
(triumph, intoxication altered states of consciousness, contentment· normal-paced, subdued, neutral tension, casual)that like going against the grain is just going to win him everything. When most times, if you just took the guy who was mocked to us in the latest mock draft, would be a much better team than we are now. (ahem) Um, so I just want that on the record. But yeah, I mean, look, I'm not I'm not gonna defend our drafting of wide receivers, but I think that's part of why we wanted a top pick because it's like the closer you get to the top of the draft, the less
full caption & clip details
An adolescent masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as triumph, intoxication altered states of consciousness, contentment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 10.0/10; 27.3s, EN.
903150_00261119 · in -26.8 dBFS · gain +6.8 dB · podcast-06366
(impatience and irritability·brisk, energised, neutral tension, casual)I think it's the opposite. There's gotta be the higher the draft pick, the more pressure, the more magnification on the pick.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as impatience and irritability; style: casual, dramatic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.3/10; 6.8s, EN.
903150_00264200 · in -23.9 dBFS · gain +3.9 dB · podcast-01190
CHNK — chunking / phrasing density ↑voicenet__VN1__T0.20__C0.25__INTERNAL · #3
This is a VoiceNet dimension, not an emotion: chunking / phrasing density (CHNK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with chunking / phrasing density (CHNK) below average — 0.26, lower than 74 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.42.
It takes 4 clips to get there. Clip to clip the moves are +0.08, then +0.11, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.95 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.95 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 24 s · en · emolia
hear it un-normalised (raw levels, max seam 5.5 dB)
k 4d_a 0.423d_b 0.423step_a 0.240step_b 0.240min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_-ijxINwNuC8track EN_-ijxINwNuC8total 23.8slevel spread 5.5 dBmax seam 5.5 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(doubt, fear · some disfluency, monologue, didactic)(ahem) It is especially a big problem if the parents are low socio-economic status or if they are very high.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, fear; style: monologue, didactic; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.0/10; 7.3s, EN.
EN_-ijxINwNuC8_W000018 · in -21.5 dBFS · gain +1.5 dB · emolia-02382
(frequent disfluency, casual, conversational)(low mumble) Uh, with very high socio-economic status, there is also this kind of, uh, (ahem)
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.2/10; 5.2s, EN.
EN_-ijxINwNuC8_W000019 · in -16.5 dBFS · gain -3.5 dB · emolia-02382
(emotional numbness, distress, helplessness·some disfluency, monologue, formal)Bad feelings on the side of the teacher that they are kind of inferior to them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, distress, helplessness; style: monologue, formal; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 0.5/10; 4.8s, EN.
EN_-ijxINwNuC8_W000020 · in -22.0 dBFS · gain +2.0 dB · emolia-02382
(some disfluency, casual, monologue)(ahem) And that's an area that we usually neglect when we talk about (low mumble) parental engagement.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.7/10; 6.0s, EN.
EN_-ijxINwNuC8_W000021 · in -19.2 dBFS · gain -0.8 dB · emolia-02382
This is a VoiceNet dimension, not an emotion: resonance: mixed (R_MIXD) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: mixed (R_MIXD) around average — 0.55, higher than 55 % of clips in this corpus — and works its way down to low at 0.16, lower than 84 % of clips in this corpus. That is a total fall of 0.39.
It takes 3 clips to get there. Clip to clip the moves are -0.20, then -0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 21 s · en · emolia
hear it un-normalised (raw levels, max seam 1.9 dB)
k 3d_a -0.394d_b -0.394step_a 0.201step_b 0.201min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_nui2wv8w9w0track EN_nui2wv8w9w0total 20.9slevel spread 1.9 dBmax seam 1.9 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged somewhat feminine voice · neutral-bright, fairly smooth, quiet background, normally alert, slightly relaxed, average clarity, light breath
(pain · measured, moderately variable, some disfluency, monologue)Will be from this charge to that charge and it will be repaired. And if the charge happens to be.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as pain; style: monologue, formal; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.9/10; 6.6s, EN.
EN_nui2wv8w9w0_W000053 · in -16.4 dBFS · gain -3.6 dB · emolia-01124
(normal-paced, fairly steady, little disfluency, formal)A region which has a positive charge and a negative charge.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, playful; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.3/10; 3.6s, EN.
EN_nui2wv8w9w0_W000055 · in -15.6 dBFS · gain -4.4 dB · emolia-01124
(concentration·measured, moderately variable, frequent disfluency, didactic)Now these lines are so close, the lines along a particular direction are so close that we normally connect them by a continuous curve.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: didactic, playful; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.9/10; 10.3s, EN.
EN_nui2wv8w9w0_W000056 · in -17.5 dBFS · gain -2.5 dB · emolia-01124
This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with warmth (WARM) above average — 0.59, higher than 59 % of clips in this corpus — and works its way down to below average at 0.28, lower than 72 % of clips in this corpus. That is a total fall of 0.31.
It takes 4 clips to get there. Clip to clip the moves are -0.19, then +0.10, then -0.22 — not a clean run: step 2 moves back the other way by 0.10 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.48 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.53 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.48, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 69 s · en · podcast
hear it un-normalised (raw levels, max seam 2.2 dB)
Unchanged across all 4 clips: a young adult somewhat masculine voice
(embarrassment, relief, fear · fast, normally alert, neutral tension, casual)Not scared to send someone out. I'm gonna be I was nodding at the back of the room and go, yeah, perfect. I'd just written that that student did need to go, and
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, vulnerable; reads as embarrassment, relief, fear; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.1/6; vocal-burst blend 4.0/10; 6.2s, EN.
907719_00087072 · in -34.8 dBFS · gain +14.8 dB · podcast-02055
(shame, affection·normal-paced, normally alert, slightly relaxed, whispered)just to follow off on what James said, I think that's really good advice for (low mumble) um NQTs and trainees to follow things up and not to promise anything you can't do if you can't think of a sanction there, and then rather than promising that you're gonna phone 15 children
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, dark, fairly smooth, slightly thin; slurred, some disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as shame, affection; style: whispered, narration; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.8/10; 13.4s, EN.
907719_00087792 · in -35.6 dBFS · gain +15.6 dB · podcast-00957
(contemplation, embarrassment, contentment·fast, normally alert, neutral tension, casual)parents or something, maybe say, I'm not sure what I'm gonna do at the moment, I'm not happy with what you've done, and I'll come back to you when I've thought about it and spoken to your head of year or head of department, buys you some time and means you don't make a rash decision. And that could just be then a conversation the next lesson. I've decided actually what I'm gonna do is just talk to you outside for a few moments, and that's that dealt with. But at least you did follow it up, but you didn't promise something crazy.
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, embarrassment, contentment; style: casual, monologue; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 6.8/10; 21.0s, EN.
907719_00089168 · in -34.9 dBFS · gain +14.9 dB · podcast-00956
(contentment, relief, concentration·normal-paced, very low-energy, slightly relaxed, casual)Well, let's move on. (ahem) Um let's now just take some time to perhaps reflect on our own practice. So, in terms of BFL, what areas do you think (ahem) um you've mastered? An area that you've mastered, and perhaps an area that you still feel you need to work on. Is there an area that you would like to see yourself (low mumble) developing? James, we'll (low mumble) start with you.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, relief, concentration; style: casual, conversational; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 6.1/10; 28.2s, EN.
907719_00091520 · in -32.6 dBFS · gain +12.6 dB · podcast-00955
This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: authoritative (S_AUTH) around average — 0.44, lower than 56 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.01, lower than 99 % of clips in this corpus. That is a total fall of 0.43.
It takes 5 clips to get there. Clip to clip the moves are -0.14, then -0.03, then -0.08, then -0.17 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 39 s · zh · emolia
k 5d_a -0.425d_b -0.425step_a 0.174step_b 0.174min_cos_consec —min_cos_anchor —dataset emolialang zhspeaker ZH_B00014_S02585track ZH_B00014_S02585total 39.1slevel spread 2.4 dBmax seam 2.4 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · no background noise, slow, audible breath
This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with perceived gender (GEND) high — 0.77, higher than 77 % of clips in this corpus — and works its way down to around average at 0.55, higher than 55 % of clips in this corpus. That is a total fall of 0.23.
It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.21 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.21 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 8 s · en · emolia
k 2d_a -0.230d_b -0.230step_a 0.230step_b 0.230min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_CClBH7Sel_ctrack EN_CClBH7Sel_ctotal 7.7slevel spread 4.2 dBmax seam 4.2 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · slightly cool, slightly bright, very rough, thin, noisy background, brisk, highly aroused, moderately variable
(hope enthusiasm optimism, elation, intoxication altered states of consciousness · tense, almost no disfluency, very clear, ranting)She's got the whole WWE universe rallying behind her.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, slightly bright, very rough, thin; very clear, almost no disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, fairly guarded; reads as hope enthusiasm optimism, elation, intoxication altered states of consciousness; style: ranting, casual; poor recording, noisy background; genuineness 3.4/6; vocal-burst blend 4.4/10; 4.4s, EN.
EN_CClBH7Sel_c_W000032 · in -14.8 dBFS · gain -5.2 dB · emolia-00894
(distress, impatience and irritability, relief·slightly tense, no disfluency, clear, casual)She's taking this outside. This one cannot be lost by Countdown.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, very rough, thin; clear, no disfluency, wide pitch range, normal breath; affect is deeply negative, slightly dominant, guarded; reads as distress, impatience and irritability, relief; style: casual, authoritative; below-average recording, noisy background; genuineness 2.2/6; vocal-burst blend 1.9/10; 3.1s, EN.
EN_CClBH7Sel_c_W000033 · in -19.0 dBFS · gain -1.0 dB · emolia-00894
This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: monologue (S_MONO) low — 0.12, lower than 88 % of clips in this corpus — and ends with it below average at 0.39, lower than 61 % of clips in this corpus. That is a total rise of 0.27.
It takes 5 clips to get there. Clip to clip the moves are -0.11, then +0.03, then +0.18, then +0.17 — not a clean run: step 1 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.01 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.01 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 21 s · en · emolia
k 5d_a 0.271d_b 0.271step_a 0.185step_b 0.185min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_X2v_hyNn8kItrack EN_X2v_hyNn8kItotal 21.4slevel spread 4.2 dBmax seam 4.0 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · thin
(infatuation, longing, jealousy and envy · normal-paced, energised, neutral tension, storytelling)He wants you for his wife. He loves beautiful women.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, rough, thin; average clarity, almost no disfluency, wide pitch range, audible breath; affect is mildly negative, dominant, fairly guarded; reads as infatuation, longing, jealousy and envy; style: storytelling, cartoonish; poor recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.5/10; 3.8s, EN.
EN_X2v_hyNn8kI_W000101 · in -19.1 dBFS · gain -0.9 dB · emolia-02540
(distress, helplessness, pain·fast, frantic, very tense, dramatic)Where in the world can he be? Mike! Mike!
full caption & clip details
A child strongly feminine voice; delivery is frantic, fast, very tense, volatile; timbre is cool, bright, gravelly, thin; slurred, heavy disfluency, extreme pitch range, heavy breath; affect is elated, slightly dominant, very vulnerable; reads as distress, helplessness, pain; style: dramatic, casual; below-average recording, some background noise; genuineness 2.0/6; vocal-burst blend 3.5/10; 4.0s, EN.
EN_X2v_hyNn8kI_W000105 · in -15.1 dBFS · gain -4.9 dB · emolia-02540
(jealousy and envy, impatience and irritability, distress ·brisk, energised, tense, storytelling)Forgive you? Just wait till I tell my husband!
full caption & clip details
A child strongly feminine voice; delivery is energised, brisk, tense, volatile; timbre is cool, bright, very rough, thin; slurred, some disfluency, very wide pitch range, heavy breath; affect is deeply negative, slightly dominant, guarded; reads as jealousy and envy, impatience and irritability, distress; style: storytelling, casual; below-average recording, some background noise; genuineness 3.4/6; vocal-burst blend 3.7/10; 3.5s, EN.
EN_X2v_hyNn8kI_W000109 · in -17.4 dBFS · gain -2.6 dB · emolia-02540
An elderly strongly masculine voice; delivery is highly aroused, slow, tense, volatile; timbre is cool, dark, rough, thin; slurred, almost no disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, guarded; reads as malevolence malice, infatuation, emotional numbness; style: storytelling, cartoonish; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.0/10; 5.7s, EN.
EN_X2v_hyNn8kI_W000110 · in -19.3 dBFS · gain -0.7 dB · emolia-02540
(distress, sadness, helplessness·measured, very low-energy, relaxed, storytelling)Very well. I won't tell my husband. Now let me out.
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, dark, smooth, thin; slurred, little disfluency, wide pitch range, audible breath; affect is negative, submissive, vulnerable; reads as distress, sadness, helplessness; style: storytelling, whispered; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.5/10; 3.8s, EN.
EN_X2v_hyNn8kI_W000111 · in -18.5 dBFS · gain -1.5 dB · emolia-02540
DARC — darkness of timbre ↓voicenet__VN1__T0.20__C0.25__INTERNAL · #9
This is a VoiceNet dimension, not an emotion: darkness of timbre (DARC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with darkness of timbre (DARC) high — 0.78, higher than 78 % of clips in this corpus — and works its way down to below average at 0.38, lower than 62 % of clips in this corpus. That is a total fall of 0.40.
It takes 5 clips to get there. Clip to clip the moves are -0.17, then +0.13, then -0.21, then -0.15 — not a clean run: step 2 moves back the other way by 0.13 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 54 s · en · emolia
k 5d_a -0.402d_b -0.402step_a 0.208step_b 0.208min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_uq4xpdYL700track EN_uq4xpdYL700total 54.3slevel spread 2.7 dBmax seam 2.1 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency, moderate pitch range
(interest · normal-paced, average clarity, casual, conversational)You may also notice that there are closed captions or subtitles running on your screen. (ahem) Uh, if they are running on your screen and you wish to turn them off, you can go to the bottom of your screen, click on the live transcript button and click, (low mumble) uh, hide subtitles, (low mumble) uh, to get those off your screen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest; style: casual, conversational; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 3.3/10; 19.9s, EN.
EN_uq4xpdYL700_W000008 · in -17.3 dBFS · gain -2.7 dB · emolia-00804
(hope enthusiasm optimism, infatuation· normal-paced, average clarity, casual, conversational)(ahem) Before we jump in, I do want to mention we do have another virtual event coming up in a month.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, infatuation; style: casual, conversational; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.2/10; 5.1s, EN.
EN_uq4xpdYL700_W000009 · in -15.5 dBFS · gain -4.5 dB · emolia-00804
(normal-paced, average clarity, casual, monologue)That's Streaming Media West Connect and that will be happening November 1st through the 5th.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.9/10; 3.8s, EN.
EN_uq4xpdYL700_W000010 · in -14.6 dBFS · gain -5.4 dB · emolia-00804
(hope enthusiasm optimism, elation·brisk, average clarity, casual, conversational)And once again, if you can't make it live to that, or if the timings just don't work out for you, all those sessions will be available on our YouTube channel. (ahem) Uh, Steve Nathan's Kelly, our producer, just popped the link to streaming media connect in the chat.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation; style: casual, conversational; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.2/10; 12.3s, EN.
EN_uq4xpdYL700_W000011 · in -16.7 dBFS · gain -3.3 dB · emolia-00804
(hope enthusiasm optimism, elation · brisk, somewhat unclear, casual, conversational)And he will probably also pop the link to our YouTube channel in there as well. You can also find all our conference videos by going to streamingmedia.com or europe.streamingmedia.com and clicking on videos.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 5.3/10; 12.8s, EN.
EN_uq4xpdYL700_W000012 · in -15.6 dBFS · gain -4.4 dB · emolia-00804
This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: monologue (S_MONO) low — 0.22, lower than 78 % of clips in this corpus — and ends with it above average at 0.60, higher than 60 % of clips in this corpus. That is a total rise of 0.38.
It takes 5 clips to get there. Clip to clip the moves are -0.09, then +0.16, then +0.09, then +0.22 — not a clean run: step 1 moves back the other way by 0.09 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 68 s · en · emolia
k 5d_a 0.381d_b 0.381step_a 0.222step_b 0.222min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00008_S01409track EN_B00008_S01409total 67.9slevel spread 1.5 dBmax seam 1.2 dB
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, good recording, some disfluency
(hope enthusiasm optimism, relief, interest · brisk, energised, neutral tension, casual)Don't worry about, don't worry about who sees it. But sometimes all of a sudden you're thinking, well, if we know, if the boss can't see what we're doing, we're not going to get these extra funds. We're not going to get this extra personnel. How do I set this up so my boss can actually see so I can take care of the team better?
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, relief, interest; style: casual, conversational; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 8.0/10; 12.1s, EN.
EN_B00008_S01409_W000182 · in -15.7 dBFS · gain -4.3 dB · emolia-00425
(teasing, anger, malevolence malice·normal-paced, normally alert, neutral tension, casual)So there are times when showmanship is going to play a role in you being able to take care of your team better. That happens. How much of, how much showmanship did you do flying an F-18 at the Top Gun school?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, full; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as teasing, anger, malevolence malice; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.2/6; vocal-burst blend 3.8/10; 13.9s, EN.
EN_B00008_S01409_W000183 · in -15.7 dBFS · gain -4.3 dB · emolia-00425
(doubt, concentration, confusion· normal-paced, normally alert, slightly relaxed, casual)I'm not gonna even rely on faking it. What I'm gonna do is I'm gonna figure out in my own head how this makes sense. How, (ahem) how this makes sense for me to be able to do what I'm gonna do to the nth degree. Like, I can tell you, here's one. When I was going to college,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as doubt, concentration, confusion; style: casual, conversational; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 7.2/10; 17.6s, EN.
EN_B00008_S01409_W000184 · in -16.9 dBFS · gain -3.1 dB · emolia-00425
(shame, fear, distress· normal-paced, normally alert, neutral tension, casual)I didn't want to work, I didn't want to do these, read this book and write this, do this, all this stupid stuff. I turned it into a competition in my head against the teacher.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as shame, fear, distress; style: casual, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 5.7/10; 8.1s, EN.
EN_B00008_S01409_W000185 · in -17.2 dBFS · gain -2.8 dB · emolia-00425
(triumph, relief, contemplation· normal-paced, energised, neutral tension, conversational)So I think that was my tactic and my technique was to truly find what was good about some situation and completely focus on that 100%. (ahem) Because then I'm not trying to, because you know I always say you can smell, you can smell, uh, (low mumble)
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as triumph, relief, contemplation; style: conversational, casual; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 9.8/10; 15.6s, EN.
EN_B00008_S01409_W000186 · in -16.1 dBFS · gain -3.9 dB · emolia-00425
This is a VoiceNet dimension, not an emotion: style: formal (S_FORM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: formal (S_FORM) around average — 0.49, right about the corpus median — and ends with it high at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.29.
It takes 5 clips to get there. Clip to clip the moves are +0.10, then +0.04, then +0.01, then +0.13 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 50 s · en · emolia
k 5d_a 0.295d_b 0.295step_a 0.135step_b 0.135min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_sHs053gQR4Ytrack EN_sHs053gQR4Ytotal 50.0slevel spread 2.8 dBmax seam 2.8 dB
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(concentration · normal-paced, some disfluency, average clarity, casual)In Java you need the semicolon, in this language we do not. Sorry, I apologize for that. So no semicolon there. (low mumble) Uh, we're gonna write the value
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.2/10; 10.6s, EN.
EN_sHs053gQR4Y_W000057 · in -19.1 dBFS · gain -0.8 dB · emolia-02284
(concentration · normal-paced, frequent disfluency, average clarity, monologue)Of each of the values in the for loop 0 through 255 to our pin. So 0 is going to be really really dim and
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.3/10; 8.8s, EN.
EN_sHs053gQR4Y_W000058 · in -21.9 dBFS · gain +1.9 dB · emolia-02284
(brisk, some disfluency, average clarity, casual)Up to 255, which will be as bright as it can be. Remember, we're using analog write, and analog write can only take the values between 0
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, dramatic; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 2.7/10; 7.9s, EN.
EN_sHs053gQR4Y_W000059 · in -20.8 dBFS · gain +0.8 dB · emolia-02284
(concentration·normal-paced, some disfluency, average clarity, monologue)And 255. Whereas analog read could read the values between 0 and 1023. So there's a bit of a difference there when we're using with our analog write with our analog read. And it's important to remember. Then we're gonna delay before we go to the next. So it's gonna
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration; style: monologue, casual; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 2.0/10; 16.8s, EN.
EN_sHs053gQR4Y_W000060 · in -21.6 dBFS · gain +1.6 dB · emolia-02284
(brisk, little disfluency, clear, authoritative)Delay a little bit so we can actually see the change. If we make it go too fast, then it's really hard to see the change.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.1/10; 5.3s, EN.
EN_sHs053gQR4Y_W000061 · in -20.9 dBFS · gain +0.9 dB · emolia-02284
This is a VoiceNet dimension, not an emotion: resonance: chest (R_CHST) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: chest (R_CHST) high — 0.84, higher than 84 % of clips in this corpus — and works its way down to around average at 0.53, higher than 53 % of clips in this corpus. That is a total fall of 0.31.
It takes 3 clips to get there. Clip to clip the moves are -0.12, then -0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, average clarity, wide pitch range
(confusion, doubt, contemplation · measured, neutral tension, fairly steady, casual)Mirror not a window, what do you mean? Why does that grade against you? Why are you standing on a platform of of an identity and that statement came and
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as confusion, doubt, contemplation; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.1/10; 13.5s, EN.
612621_00507863 · in -24.3 dBFS · gain +4.3 dB · podcast-04093
(sexual lust, teasing·brisk, neutral tension, moderately variable, conversational)legs that you're standing on. You know, is that the coffee mug that you're holding a little too tight to a little too tight,
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as sexual lust, teasing; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 4.5/6; vocal-burst blend 4.2/10; 5.8s, EN.
612621_00509632 · in -21.6 dBFS · gain +1.6 dB · podcast-00535
(embarrassment, teasing, intoxication altered states of consciousness·normal-paced, relaxed, moderately variable, casual)you can quit any time you like. (low mumble) Um I get (low mumble) it. Uh sure you can. Sure you can. No, I I I totally believe you. Right.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, teasing, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 2.4/10; 7.6s, EN.
612621_00510687 · in -22.3 dBFS · gain +2.3 dB · podcast-01439
This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: authoritative (S_AUTH) below average — 0.38, lower than 62 % of clips in this corpus — and works its way down to low at 0.14, lower than 86 % of clips in this corpus. That is a total fall of 0.24.
It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, frequent disfluency, average clarity
(contemplation, interest, concentration · measured, relaxed, fairly steady, casual)So I I think those experiences are even so much more meaningful if you think about it in that way. (low mumble) Um that's the A, you know, my A B test. (low mumble) Um and then the other one is and and this is like going back to V1. (low mumble) Um separate the church as a corporation from your life of fee.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contemplation, interest, concentration; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 5.5/10; 21.9s, EN.
148981_00192704 · in -20.3 dBFS · gain +0.3 dB · podcast-03076
(bitterness, impatience and irritability, sourness·normal-paced, neutral tension, moderately variable, casual)same, you're gonna have a hard time, bad time, right? Because church as a corporation is run by people like me.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as bitterness, impatience and irritability, sourness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.2/10; 7.7s, EN.
148981_00195376 · in -19.4 dBFS · gain -0.7 dB · podcast-04750
This is a VoiceNet dimension, not an emotion: resonance: oral (R_ORAL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with resonance: oral (R_ORAL) above average — 0.69, higher than 69 % of clips in this corpus — and works its way down to low at 0.24, lower than 76 % of clips in this corpus. That is a total fall of 0.45.
It takes 5 clips to get there. Clip to clip the moves are -0.24, then -0.23, then +0.24, then -0.21 — not a clean run: step 3 moves back the other way by 0.24 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 45 s · en · emolia
k 5d_a -0.447d_b -0.447step_a 0.244step_b 0.244min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_PLSiCAhWUyItrack EN_PLSiCAhWUyItotal 45.0slevel spread 2.3 dBmax seam 1.8 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, no disfluency
(measured, steady, crisply articulate, casual)== Techniques of molecular biology ==
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, thin; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.2/10; 3.7s, EN.
EN_PLSiCAhWUyI_W000024 · in -14.1 dBFS · gain -5.9 dB · emolia-01183
(normal-paced, steady, clear, formal)One of the most basic techniques of molecular biology to study protein function is molecular cloning
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 7.1s, EN.
EN_PLSiCAhWUyI_W000026 · in -14.7 dBFS · gain -5.3 dB · emolia-01183
(emotional numbness, concentration·measured, steady, clear, formal)In this technique, DNA coding for a protein of interest is cloned using polymerase chain reaction, PCR, and or restriction enzymes into a plasmid' expression vector
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.3s, EN.
EN_PLSiCAhWUyI_W000027 · in -14.7 dBFS · gain -5.3 dB · emolia-01183
(emotional numbness · measured, steady, clear, formal)A vector has three distinctive features, an origin of replication, a multiple cloning site Mcs, and a selective marker usually antibiotic resistance
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_PLSiCAhWUyI_W000028 · in -16.4 dBFS · gain -3.6 dB · emolia-01183
(emotional numbness · measured, fairly steady, clear, formal)Located upstream of the multiple cloning site are the promoter regions and the transcription start site which regulate the expression of cloned gene
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.8s, EN.
EN_PLSiCAhWUyI_W000029 · in -15.4 dBFS · gain -4.6 dB · emolia-01183
This is a VoiceNet dimension, not an emotion: style: storytelling (S_STRY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: storytelling (S_STRY) around average — 0.45, lower than 55 % of clips in this corpus — and works its way down to low at 0.23, lower than 77 % of clips in this corpus. That is a total fall of 0.23.
It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 12 s · en · emolia
k 2d_a -0.226d_b -0.226step_a 0.226step_b 0.226min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_TpZoZaR-HMctrack EN_TpZoZaR-HMctotal 12.0slevel spread 0.1 dBmax seam 0.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(relief, sadness · monologue, casual)Of all the problems that developing countries have fallen into after 20 or 30 years of growth beforehand.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, sadness; style: monologue, casual; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.9/10; 6.5s, EN.
EN_TpZoZaR-HMc_W000283 · in -16.2 dBFS · gain -3.8 dB · emolia-02086
(contemplation, doubt· monologue, didactic)Maybe, maybe when it seems like this crisis, the sensible thing to do is just to wait.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, doubt; style: monologue, didactic; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.5/10; 5.3s, EN.
EN_TpZoZaR-HMc_W000284 · in -16.3 dBFS · gain -3.7 dB · emolia-02086
AGEV — perceived age ↓voicenet__VN1__T0.20__C0.25__INTERNAL · #16
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with perceived age (AGEV) above average — 0.70, higher than 70 % of clips in this corpus — and works its way down to around average at 0.46, lower than 54 % of clips in this corpus. That is a total fall of 0.24.
It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 10 s · snippets
k 2d_a -0.242d_b -0.242step_a 0.242step_b 0.242min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch82_part1_batch82_parttrack batch82_part1_batch82_parttotal 10.2slevel spread 0.2 dBmax seam 0.2 dB
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(impatience and irritability, sourness, anger · normal-paced, almost no disfluency, clear, storytelling)the system is such, (ahem) it just doesn't reward that kind of behavior.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as impatience and irritability, sourness, anger; style: storytelling, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.5/10; 5.1s.
batch82_part1_batch82_part1_chunk_1738_1_2013763 · in -29.7 dBFS · gain +9.7 dB · snippets-01314
(measured, some disfluency, slurred, casual)we did a nationwide survey of a total (ahem) of 6,500 families.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.5/10; 5.0s.
batch82_part1_batch82_part1_chunk_1738_1_2013778 · in -29.9 dBFS · gain +9.9 dB · snippets-01314
This is a VoiceNet dimension, not an emotion: register (speaking pitch region) (REGS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with register (speaking pitch region) (REGS) low — 0.17, lower than 83 % of clips in this corpus — and ends with it around average at 0.48, lower than 52 % of clips in this corpus. That is a total rise of 0.30.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.05, then +0.09 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.76 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.76 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 45 s · en · emolia
k 4d_a 0.304d_b 0.304step_a 0.170step_b 0.170min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_udPzw5U_N3Qtrack EN_udPzw5U_N3Qtotal 44.6slevel spread 2.0 dBmax seam 1.1 dB
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · slightly relaxed
(pleasure ecstasy, contentment, interest · normal-paced, normally alert, fairly steady, casual)Some fish rice, which I decided to go with. Comes with a bowl of fruits, (ahem) some ice cream as well. And in terms of presentation, it looks really good. Cruised real quick, (ahem) very efficient. So now let's dig in and let's see what it tastes like.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pleasure ecstasy, contentment, interest; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.8/10; 16.1s, EN.
EN_udPzw5U_N3Q_W000046 · in -19.3 dBFS · gain -0.7 dB · emolia-00861
(pleasure ecstasy, hope enthusiasm optimism, elation· normal-paced, normally alert, moderately variable, casual)So, and of course this video wouldn't be complete with a loo review. (low mumble) Uhm, the meal, by the way, was amazing, especially the ice cream. Yum. Very, very yummy. And this is what a brand new Airbus 321 (low mumble) loo looks like.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as pleasure ecstasy, hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 8.7/10; 14.7s, EN.
EN_udPzw5U_N3Q_W000047 · in -20.4 dBFS · gain +0.4 dB · emolia-00861
(awe, intoxication altered states of consciousness·measured, very low-energy, fairly steady, whispered)Very clean, (ahem) very organized, very fresh. I mean, look at that basin here. I mean, look at the sink.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as awe, intoxication altered states of consciousness; style: whispered, casual; below-average recording, some background noise; genuineness 2.3/6; vocal-burst blend 1.2/10; 8.7s, EN.
EN_udPzw5U_N3Q_W000048 · in -20.3 dBFS · gain +0.3 dB · emolia-00861
(normal-paced, normally alert, fairly steady, casual)Got a couple of amenities here as well. So, this is an A-class loo.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.7/10; 4.7s, EN.
EN_udPzw5U_N3Q_W000049 · in -21.3 dBFS · gain +1.3 dB · emolia-00861
DARC — darkness of timbre ↓voicenet__VN1__T0.20__C0.25__INTERNAL · #18
This is a VoiceNet dimension, not an emotion: darkness of timbre (DARC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with darkness of timbre (DARC) below average — 0.33, lower than 67 % of clips in this corpus — and works its way down to low at 0.10, lower than 90 % of clips in this corpus. That is a total fall of 0.22.
It takes 2 clips to get there. Clip to clip the moves are -0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, fairly steady, moderate pitch range
(thankfulness gratitude · normal-paced, neutral tension, some disfluency, conversational)And I think, well, um thank you for asking this question because I think most entrepreneurs or business owners will probably agree. When I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude; style: conversational, casual; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.8/10; 6.3s, EN.
8961_00064328 · in -28.1 dBFS · gain +8.1 dB · podcast-04976
(doubt·slow, relaxed, frequent disfluency, casual)biggest struggle is staff and employment HR. (low mumble) Um, you know,
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.5/10; 5.8s, EN.
8961_00065640 · in -28.1 dBFS · gain +8.1 dB · podcast-04971
This is a VoiceNet dimension, not an emotion: style: ranting (S_RANT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: ranting (S_RANT) around average — 0.45, lower than 55 % of clips in this corpus — and ends with it high at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.35.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then -0.06, then +0.22 — not a clean run: step 2 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 19 s · en · emolia
k 4d_a 0.350d_b 0.350step_a 0.221step_b 0.221min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00006_S06870track EN_B00006_S06870total 19.0slevel spread 1.9 dBmax seam 1.9 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, some disfluency, average clarity, light breath
(teasing, pleasure ecstasy · normally alert, slightly relaxed, fairly steady, casual)But you know who I feel bad for in terms of music? Who? Pete Best.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as teasing, pleasure ecstasy; style: casual, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.1/10; 3.2s, EN.
EN_B00006_S06870_W000138 · in -21.5 dBFS · gain +1.5 dB · emolia-00374
(confusion, impatience and irritability, doubt·energised, slightly relaxed, moderately variable, conversational)Why? You know what that is? Why do you feel bad for Pete Best?
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as confusion, impatience and irritability, doubt; style: conversational, casual; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 0.0/10; 3.3s, EN.
EN_B00006_S06870_W000139 · in -21.0 dBFS · gain +1.0 dB · emolia-00374
(teasing·normally alert, slightly relaxed, moderately variable, conversational)Laughing all the way to the bank. Oh wow, killing it. So let me ask you something. Yea. So Pete Best, right, was the original drummer for the Beatles.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as teasing; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 0.9/10; 7.1s, EN.
EN_B00006_S06870_W000140 · in -20.1 dBFS · gain +0.1 dB · emolia-00374
(pride, elation· normally alert, neutral tension, moderately variable, casual)Before Ringo, and he, he best goes, guys, I'm gonna go to art school.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as pride, elation; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.1/10; 5.0s, EN.
EN_B00006_S06870_W000141 · in -22.1 dBFS · gain +2.1 dB · emolia-00374
This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with disfluency (DFLU) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to above average at 0.66, higher than 66 % of clips in this corpus. That is a total fall of 0.28.
It takes 3 clips to get there. Clip to clip the moves are -0.15, then -0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: an elderly masculine voice · neutral-toned, slightly rough, balanced body, frequent disfluency
(pride, teasing, triumph · slow, very low-energy, relaxed, monologue)(low mumble) Um so that is our week one preview. I will say I made two random bets.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as pride, teasing, triumph; style: monologue, whispered; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 2.5/10; 8.8s, EN.
775551_00332096 · in -22.3 dBFS · gain +2.3 dB · podcast-02053
(shame, bitterness, contemplation· slow, subdued, relaxed, whispered)On this one, because who would I be if I didn't? I gotta keep the 7k pick traditional live.
full caption & clip details
A young adult masculine voice; delivery is subdued, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as shame, bitterness, contemplation; style: whispered, monologue; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.6/10; 8.4s, EN.
775551_00333056 · in -25.3 dBFS · gain +5.3 dB · podcast-02053
(amusement, embarrassment, pleasure ecstasy·normal-paced, normally alert, neutral tension, casual)So literally, I said to my friend that if that if the three dollar one hits, then the 7k pick podcast, I am paying for (ahem) all inclusive trip. I'm not paying for their drinks.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, embarrassment, pleasure ecstasy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 3.5/10; 15.6s, EN.
775551_00337024 · in -20.4 dBFS · gain +0.4 dB · podcast-03749