Manifest tier. voicenet, rule VN1, T=0.5, step cap 0.25. Population 17,672,060 chains (233,454 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 12,494,744.
Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.voicenet__VN1__T0.50__C0.25__INTERNAL — population 17,672,060 chains (233,454 h). SHAREABLE variant: 12,494,744. Filter.rule=='VN1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and qmax>=0.5 and cmax<=0.25 This tier was resampled. The manifest's earlier published filter applied the one-sided test to every tier regardless of its actual rule. Tier populations were always counted under the correct rule, so the sizes quoted here were never wrong — but the earlier selection drew from a contaminated pool. These samples are drawn with the corrected filter, which tests qmax >= T and cmax <= C — per-row quantities computed under this tier's own rule. For a proxy tier no d_a/d_b predicate could be correct, because a proxy-rescued chain may legitimately exceed the cap on the named axis while its smoothness is certified on the proxy axis. Sampled from 116,099 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is a VoiceNet dimension, not an emotion: clarity / intelligibility (CLRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with clarity / intelligibility (CLRT) above average — 0.70, higher than 70 % of clips in this corpus — and works its way down to low at 0.17, lower than 83 % of clips in this corpus. That is a total fall of 0.53.
It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.21, then -0.03, then -0.09 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · zh · emolia
hear it un-normalised (raw levels, max seam 5.4 dB)
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, slightly relaxed, fairly steady, some disfluency, moderate pitch range, light breath
This is a VoiceNet dimension, not an emotion: roughness (ROUG) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with roughness (ROUG) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.53.
It takes 5 clips to get there. Clip to clip the moves are -0.24, then -0.17, then +0.02, then -0.14 — not a clean run: step 3 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 29 s · snippets
hear it un-normalised (raw levels, max seam 5.1 dB)
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, some disfluency, average clarity, light breath
(pride, interest · fast, energised, neutral tension, dramatic)he says, (ahem) in this article, I want to discuss the possibility that the goal of theoretical physics might be achieved.
full caption & clip details
An adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, interest; style: dramatic, storytelling; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.3/10; 5.3s.
batch208_part0_batch208_part0_chunk_342_1_179488 · in -23.9 dBFS · gain +3.9 dB · snippets-00559
(impatience and irritability, bitterness, contempt·measured, normally alert, slightly relaxed, casual)not the opposite. So you can't run those films backwards.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as impatience and irritability, bitterness, contempt; style: casual, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.5/10; 3.8s.
batch208_part0_batch208_part0_chunk_342_1_179628 · in -22.0 dBFS · gain +2.0 dB · snippets-00559
(shame, embarrassment, pain·normal-paced, normally alert, slightly relaxed, conversational)One argument which I'm actually quite in favor of is that we're not going to get a fear of everything unless we bring in that last
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, embarrassment, pain; style: conversational, casual; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.2/10; 8.9s.
batch208_part0_batch208_part0_chunk_342_1_179642 · in -20.9 dBFS · gain +0.9 dB · snippets-00559
(normal-paced, normally alert, slightly relaxed, casual)We know what electricity is, we know what magnetism is. They're actually
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.7/10; 3.5s.
batch208_part0_batch208_part0_chunk_342_1_179943 · in -26.1 dBFS · gain +6.0 dB · snippets-00559
(awe, confusion, astonishment surprise· normal-paced, normally alert, neutral tension, conversational)the planets in orbit around the sun and the moon in orbit around the earth. I mean until then no one thought that was why would that be obvious.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as awe, confusion, astonishment surprise; style: conversational, casual; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.3/10; 6.4s.
batch208_part0_batch208_part0_chunk_342_1_179981 · in -25.7 dBFS · gain +5.7 dB · snippets-00559
This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with vocal focus (FOCS) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to low at 0.10, lower than 90 % of clips in this corpus. That is a total fall of 0.55.
It takes 4 clips to get there. Clip to clip the moves are -0.15, then -0.22, then -0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.07 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.07 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 47 s · en · emolia
hear it un-normalised (raw levels, max seam 2.3 dB)
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, quiet background, slightly relaxed, fairly steady, light breath
(normal-paced, normally alert, some disfluency, casual)(ahem) We'd also like to look, if possible, at Google Analytics or whatever open source equivalent we may have on the website and that'll help us determine recommendations for segmenting the site as well.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.9/10; 12.5s, EN.
EN_UcxEpUNUOZ0_W000046 · in -17.7 dBFS · gain -2.3 dB · emolia-01135
(doubt, interest, confusion· normal-paced, subdued, some disfluency, conversational)(low mumble) Uhm, one question because I, I, I don't understand these things particularly well from a technical side. I noticed your reaction when we said Google Analytics. (ahem) Uhm, and I was just wondering whether or not, (low mumble) uhm, you wanted to elaborate on, on what kind of analytics are actually available to us, what kind of information about visitors to the website and, and about the learners.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, interest, confusion; style: conversational, casual; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 19.6s, EN.
EN_UcxEpUNUOZ0_W000048 · in -18.7 dBFS · gain -1.3 dB · emolia-01135
(embarrassment·measured, normally alert, frequent disfluency, monologue)I can just tell you that (low mumble) uhm, at the moment there is Google Analytics (low mumble) as much as we are, have trepidation about using it. (low mumble) Uhm, we also have been looking at (low mumble) uhm, Google (low mumble) uhm,
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment; style: monologue, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 2.1/10; 10.5s, EN.
EN_UcxEpUNUOZ0_W000049 · in -16.4 dBFS · gain -3.6 dB · emolia-01135
(emotional numbness· measured, subdued, frequent disfluency, casual)The tagging functionality that Google has as well, which, uh, (low mumble)
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: casual, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.5/10; 4.3s, EN.
EN_UcxEpUNUOZ0_W000050 · in -17.7 dBFS · gain -2.3 dB · emolia-01135
This is a VoiceNet dimension, not an emotion: recording quality (RCQL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with recording quality (RCQL) above average — 0.74, higher than 74 % of clips in this corpus — and works its way down to low at 0.20, lower than 80 % of clips in this corpus. That is a total fall of 0.54.
It takes 5 clips to get there. Clip to clip the moves are -0.22, then -0.06, then -0.11, then -0.15 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 38 s · en · podcast
hear it un-normalised (raw levels, max seam 4.3 dB)
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body
(affection, contentment · normal-paced, normally alert, neutral tension, casual)Yeah, exactly. Your your rookies are really your key to having a better team. So it's important that they're all playing, they're all there to make money.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as affection, contentment; style: casual, conversational; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 5.0/10; 10.3s, EN.
955506_00130562 · in -27.9 dBFS · gain +7.8 dB · podcast-06060
(normal-paced, normally alert, slightly relaxed, casual)Yep. All right. So this next strategy we have is one of our cousins. It's sort of one of our personal strategies that we've adopted and we love it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.0/10; 8.4s, EN.
955506_00131616 · in -29.0 dBFS · gain +9.0 dB · podcast-02075
(fatigue exhaustion·slow, very low-energy, relaxed, casual)cousin's tip. Yep. Save a bank. So of our what, fifteen point five million starting salary.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion; style: casual, monologue; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.3/10; 9.2s, EN.
955506_00132496 · in -24.7 dBFS · gain +4.7 dB · podcast-02093
(slow, normally alert, relaxed, casual)you could spend every last dollar (low mumble) um your starting squad.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 2.9/6; vocal-burst blend 0.6/10; 6.1s, EN.
955506_00133600 · in -28.1 dBFS · gain +8.1 dB · podcast-02070
(embarrassment·normal-paced, normally alert, slightly relaxed, casual)(chuckle) mistake. 'Cause come round one. (low mumble) Um
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; style: casual, conversational; average recording, no background noise; reads as embarrassment; genuineness 4.5/6; vocal-burst blend 0.7/10; 3.4s, EN.
955506_00134328 · in -23.9 dBFS · gain +3.9 dB · podcast-00880
This is a VoiceNet dimension, not an emotion: style: newsreading (S_NEWS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: newsreading (S_NEWS) below average — 0.34, lower than 66 % of clips in this corpus — and ends with it high at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.53.
It takes 5 clips to get there. Clip to clip the moves are +0.07, then +0.19, then +0.05, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.39 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.39 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 58 s · en · emolia
hear it un-normalised (raw levels, max seam 2.9 dB)
Unchanged across all 5 clips: a young adult feminine voice · slightly bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed
(interest, concentration, doubt · normal-paced, fairly steady, some disfluency, casual)Helps confirm that you did your best, right? You don't just want to ignore these differing data or emerging issues, (ahem) but I think that there's a good way to to dive into them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, concentration, doubt; style: casual, conversational; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 2.1/10; 11.9s, EN.
EN_BwKGXXUqgmk_W000227 · in -20.5 dBFS · gain +0.5 dB · emolia-01074
(interest ·brisk, moderately variable, some disfluency, casual)(ahem) If they don't fit within the original objectives of your evaluation, I think you should talk with the project staff and say, you know, this seems really important or sensitive. And if it's sensitive, you can share back with them in more of an oral report instead of a written report.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest; style: casual, conversational; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 2.3/10; 16.4s, EN.
EN_BwKGXXUqgmk_W000228 · in -20.4 dBFS · gain +0.4 dB · emolia-01074
(concentration·normal-paced, fairly steady, some disfluency, casual)You could also talk to them about actually creating another mini evaluation report that talks about some of these issues that may not fit within the original structure of the initial evaluation plan.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: casual, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.6/10; 11.5s, EN.
EN_BwKGXXUqgmk_W000229 · in -18.3 dBFS · gain -1.7 dB · emolia-01074
(concentration ·brisk, fairly steady, some disfluency, playful)Sorry about that. So we are going to dig into section three and explore strategies for reporting qualitative data now.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: playful, casual; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.5/10; 6.6s, EN.
EN_BwKGXXUqgmk_W000231 · in -21.1 dBFS · gain +1.1 dB · emolia-01074
(normal-paced, fairly steady, little disfluency, formal)So here we have outlined eight different strategies, including word clouds, call out boxes, highlight quotes, tables, annotated graphs, photos, icons, and journey maps.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.7/10; 11.1s, EN.
EN_BwKGXXUqgmk_W000232 · in -20.6 dBFS · gain +0.7 dB · emolia-01074
AGEV — perceived age ↑voicenet__VN1__T0.50__C0.25__INTERNAL · #6
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with perceived age (AGEV) at the very bottom of the range — 0.06, lower than 94 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.63.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.14, then +0.13, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice
(normal-paced, normally alert, relaxed, casual)(surprised gasp) okay, let's go backward. What is what current business do you have?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, some background noise; genuineness 5.8/6; vocal-burst blend 1.6/10; 4.2s, EN.
534524_00008504 · in -25.4 dBFS · gain +5.5 dB · podcast-02420
(contentment, relief, sexual lust·measured, normally alert, slightly relaxed, whispered)Currently, I have one business, which is my therapy practice.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, relief, sexual lust; style: whispered, casual; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.9/10; 4.4s, EN.
534524_00008952 · in -29.9 dBFS · gain +9.9 dB · podcast-02397
(doubt, contemplation, confusion· measured, normally alert, relaxed, casual)So I'm Defon. (low mumble) Um I'm not really sure that which uh (low mumble) which of my title I should go first. What do you
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, contemplation, confusion; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.2/10; 7.1s, EN.
534524_00010664 · in -29.1 dBFS · gain +9.1 dB · podcast-02392
(pride, shame, thankfulness gratitude·normal-paced, very low-energy, neutral tension, conversational)think? (low mumble) (ahem) Professor.
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pride, shame, thankfulness gratitude; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 7.0/10; 22.2s, EN.
534524_00011368 · in -29.0 dBFS · gain +9.0 dB · podcast-05206
(shame, helplessness, pride ·measured, very low-energy, relaxed, casual)Well do we just want to start (low mumble) (low mumble) with (low mumble) um maybe how long we've been married?
full caption & clip details
An elderly somewhat masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, helplessness, pride; style: casual, ASMR; below-average recording, some background noise; genuineness 5.4/6; vocal-burst blend 9.0/10; 27.5s, EN.
534524_00013584 · in -31.2 dBFS · gain +11.2 dB · podcast-05201
This is a VoiceNet dimension, not an emotion: emphasis / stress strength (EMPH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with emphasis / stress strength (EMPH) at the very bottom of the range — 0.07, lower than 93 % of clips in this corpus — and ends with it above average at 0.61, higher than 61 % of clips in this corpus. That is a total rise of 0.54.
It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.06, then +0.20, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.31 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.31, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth
(disgust, jealousy and envy · normal-paced, very low-energy, slightly relaxed, monologue)purpose to demonstrate like yeah, she dresses proper, but she still has this quote unquote wild side to her.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as disgust, jealousy and envy; style: monologue, casual; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.0/10; 7.1s, EN.
159735_00094280 · in -42.4 dBFS · gain +22.4 dB · podcast-01506
(hope enthusiasm optimism, sexual lust, teasing·measured, normally alert, slightly relaxed, whispered)Now we can start getting to the guys. I mean, we can also talk about like
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, sexual lust, teasing; style: whispered, storytelling; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.3/10; 4.5s, EN.
159735_00095056 · in -37.6 dBFS · gain +17.6 dB · podcast-01527
(infatuation, sexual lust, longing·normal-paced, normally alert, relaxed, conversational)Oh yeah. Well, I was gonna talk about Julian and Zoya together.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as infatuation, sexual lust, longing; style: conversational, storytelling; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.1/10; 4.2s, EN.
159735_00096040 · in -32.4 dBFS · gain +12.4 dB · podcast-01516
(embarrassment, doubt, infatuation · normal-paced, normally alert, relaxed, casual)Just cause yeah. (low mumble) Um, so Nate, to be honest, I think Nate is the most boring character ever.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment, doubt, infatuation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 3.8/10; 5.6s, EN.
159735_00096544 · in -33.6 dBFS · gain +13.6 dB · podcast-01513
(amusement, sourness, jealousy and envy·brisk, energised, neutral tension, casual)Where was the storyline? Fashion wise, too. Cause he, you know, he was just the typical like Prep school jock boy, you know, like I can't really describe a style for him. Even once they graduated high school and once he started working for the newspaper his grandfather's newspaper, like he just wore like suit and ties the whole time. But
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, sourness, jealousy and envy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 6.9/10; 18.8s, EN.
159735_00097320 · in -32.0 dBFS · gain +12.0 dB · podcast-05842
This is a VoiceNet dimension, not an emotion: aesthetic pleasantness (ESTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with aesthetic pleasantness (ESTH) low — 0.08, lower than 92 % of clips in this corpus — and ends with it above average at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.58.
It takes 4 clips to get there. Clip to clip the moves are +0.13, then +0.21, then +0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.24 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.24, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording
(intoxication altered states of consciousness, confusion, astonishment surprise · normal-paced, normally alert, neutral tension, casual)was a lot of. Whoa, whoa, whoa,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as intoxication altered states of consciousness, confusion, astonishment surprise; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 5.8/6; vocal-burst blend 2.2/10; 5.4s, EN.
568215_00024240 · in -25.9 dBFS · gain +5.9 dB · podcast-05332
(concentration, triumph, hope enthusiasm optimism·measured, subdued, slightly relaxed, monologue)gonna say the word. I'm gonna try my best. Okay. I'm gonna try my best to pronounce the word correctly. I might pronounce incorrectly. Oh well, you guys signed up for this. (low mumble) Um you guys can ask for definition and use in a sentence. And then that's it. So I can repeat one more time. Definition and using a sentence.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, triumph, hope enthusiasm optimism; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.3/10; 23.9s, EN.
568215_00026888 · in -33.4 dBFS · gain +13.4 dB · podcast-03778
(doubt, embarrassment·normal-paced, normally alert, slightly relaxed, casual)No, yeah. I mean, (ahem) uh it was on Nets declassified.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; style: casual, conversational; average recording, quiet background; reads as doubt, embarrassment; genuineness 4.8/6; vocal-burst blend 1.1/10; 3.2s, EN.
568215_00029568 · in -30.3 dBFS · gain +10.3 dB · podcast-05328
(astonishment surprise· normal-paced, normally alert, relaxed, casual)Side note of this. If you haven't seen like videos (low mumble) of them together, like in podcasts, and you should watch it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as astonishment surprise; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.1/10; 6.3s, EN.
568215_00031568 · in -29.1 dBFS · gain +9.1 dB · podcast-05328
This is a VoiceNet dimension, not an emotion: resonance: oral (R_ORAL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: oral (R_ORAL) low — 0.12, lower than 88 % of clips in this corpus — and ends with it above average at 0.62, higher than 62 % of clips in this corpus. That is a total rise of 0.50.
It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.08, then +0.17 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.13 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.13, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, normal-paced, some disfluency, light breath
(pleasure ecstasy, triumph, hope enthusiasm optimism · normally alert, neutral tension, fairly steady, casual)I completely agree. Yeah. I mean, you can look at Ezekiel Elliott in any matchup. He's you know, on FanDuel, he's 8700, so you get (low mumble) um a little bit of a discount from Le'Veon Bell and David Johnson, and we talk about how these guys' ceilings are kind of all the same regardless of the matchup. So like that would be a decent place to pivot, whereas you're not gonna have that option on the (low mumble) uh Sunday slate. So
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, triumph, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.5/10; 24.1s, EN.
190701_00295008 · in -24.6 dBFS · gain +4.6 dB · podcast-02233
(interest, hope enthusiasm optimism, teasing·subdued, slightly relaxed, fairly steady, casual)that's the question I was just leaning to. You are an ownership expert and GPP game theory expert like myself. So is this the great day to go all in Lev Bell David Johnson teams on Thursday? Because they will be lower owned due to the Ezekiel bump on Thursday night. What do you think Ezekiel's ownership were gonna be and how would you play it if you had one lineup and a GPP on Thursday on FanDuel?
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism, teasing; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 7.0/10; 23.8s, EN.
190701_00297448 · in -21.4 dBFS · gain +1.4 dB · podcast-06212
(intoxication altered states of consciousness, fatigue exhaustion, elation·normally alert, neutral tension, moderately variable, casual)One lineup (low mumble) uh for Thursday. Oh man, I probably am gonna have one lineup on Thursday. So I haven't decided exactly what my lineup's gonna be, but I would absolutely consider playing Elliott, not the notes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as intoxication altered states of consciousness, fatigue exhaustion, elation; style: casual, conversational; good recording, some background noise; genuineness 4.8/6; vocal-burst blend 7.2/10; 13.1s, EN.
190701_00299864 · in -22.3 dBFS · gain +2.3 dB · podcast-02243
(astonishment surprise, confusion, teasing·energised, neutral tension, moderately variable, casual)thing. I think you're hint I thought you were hinting at it being the other way to go. No, I was. I didn't know. I'm
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, confusion, teasing; style: casual, conversational; average recording, quiet background; genuineness 6.0/6; vocal-burst blend 3.9/10; 4.4s, EN.
190701_00301520 · in -19.4 dBFS · gain -0.6 dB · podcast-02236
This is a VoiceNet dimension, not an emotion: style: asmr (S_ASMR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: asmr (S_ASMR) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.50.
It takes 4 clips to get there. Clip to clip the moves are -0.25, then -0.03, then -0.22 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: an elderly somewhat feminine voice · average recording, quiet background, normally alert
(disgust · normal-paced, neutral tension, moderately variable, casual)Earlier this decade, (childlike giggle) the First Minister announced that, in all hospitals in Scotland apart from those built (ahem) under private
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is negative, neutral stance, slightly guarded; reads as disgust; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.6/10; 15.6s, EN.
uk_uk_13_21112018_3799056_3814672 · in -26.3 dBFS · gain +6.3 dB · eurospeech-00872
(concentration, disappointment·measured, slightly relaxed, fairly steady, formal)parking charges would be withdrawn. That has been carried out. That simple measure can help nurses, and I urge the Minister most sincerely to consider it, as well as some of the other practices that we have taken on board to increase nursing numbers in Scotland.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; clear, some disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as concentration, disappointment; style: formal, monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.3/10; 19.4s, EN.
uk_uk_13_21112018_3814672_3834071 · in -28.4 dBFS · gain +8.4 dB · eurospeech-00872
(thankfulness gratitude, affection, shame·normal-paced, neutral tension, fairly steady, monologue)(low mumble) from great personal experience, and I (low mumble) thank her and everyone else who has worked in the NHS for their contribution over many years to make it an institution of which we are all rightly proud.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, affection, shame; style: monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.2/10; 12.8s, EN.
uk_uk_13_21112018_3853024_3865856 · in -24.6 dBFS · gain +4.7 dB · eurospeech-00872
(concentration, sourness, disgust·measured, slightly relaxed, steady, formal)hon. Friend comprehensively dismantled the Government’s arguments on the merits of removing the bursary. As she said, it is indisputable that the number of applications and the numbers of people starting courses have fallen, and that the age profile of students has changed.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, sourness, disgust; style: formal, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.8/10; 16.4s, EN.
uk_uk_13_21112018_3865856_3882304 · in -26.3 dBFS · gain +6.3 dB · eurospeech-00872
This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with warmth (WARM) above average — 0.66, higher than 66 % of clips in this corpus — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.54.
It takes 5 clips to get there. Clip to clip the moves are -0.22, then +0.02, then -0.16, then -0.19 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.26 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, fairly steady, moderate pitch range, light breath
(jealousy and envy, contentment, hope enthusiasm optimism · normal-paced, slightly relaxed, some disfluency, monologue)wat Ronald zegt, ik denk dat hij gewoon geslachtofferd is om de duidelijkheid van de regels aan te stippen. En dat de UCI niet met zich laat spotten over die nieuwe regels. Maar ik ben er wel van overtuigd dat als je (low mumble) die regels in acht neem, dat je natuurlijk ook nog wel een beetje kan kijken naar wanneer het gevaarlijk is en wanneer het gewoon in dienst is van de koers.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, contentment, hope enthusiasm optimism; style: monologue, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 6.8/10; 29.5s, NL.
522989_00163928 · in -20.7 dBFS · gain +0.7 dB · podcast-05986
(confusion, shame· normal-paced, neutral tension, some disfluency, casual)Ik had die foto gezien, maar dat was dus niet een verzorger of zo. Ik denk het wel op.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, shame; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 9.8/10; 16.0s, NL.
522989_00166876 · in -19.2 dBFS · gain -0.8 dB · podcast-05951
(doubt, fatigue exhaustion, relief· normal-paced, neutral tension, some disfluency, casual)zoals Van Vleut ook zei, ja, we hadden hoge snelheid, ik zag een blauwe jas, ik denk, dat is van mijn ploeg. Maar dat was niet zo. Dus ja, als je gewoon heel streng bent volgens de UCI, had ze gewoon uit koers moeten genomen worden of gedisqualificeerd na afgelopen. Maar dat doe je niet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, fatigue exhaustion, relief; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 7.3/10; 15.5s, NL.
522989_00169336 · in -19.1 dBFS · gain -0.9 dB · podcast-01821
(fast, slightly relaxed, some disfluency, casual)Of een tijdstraf, dat zou nog (ahem) pijnlijker zijn dat je net zoveel tijdszaf krijgt dat je tweede
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 1.0/10; 4.2s, NL.
522989_00170904 · in -21.6 dBFS · gain +1.6 dB · podcast-02197
(intoxication altered states of consciousness·normal-paced, slightly relaxed, frequent disfluency, conversational)wordt. Ze hebben die regel ingevoerd. En direct bij de eerste koers werken ze ook al meteen met twee maten.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: conversational, casual; average recording, some background noise; genuineness 4.1/6; vocal-burst blend 0.5/10; 9.4s, NL.
522989_00171320 · in -19.1 dBFS · gain -0.9 dB · podcast-02178
AGEV — perceived age ↓voicenet__VN1__T0.50__C0.25__INTERNAL · #12
This is a VoiceNet dimension, not an emotion: perceived age (AGEV) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with perceived age (AGEV) high — 0.90, higher than 90 % of clips in this corpus — and works its way down to low at 0.20, lower than 80 % of clips in this corpus. That is a total fall of 0.70.
It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.10, then -0.18, then -0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.44 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, average clarity
(elation, amusement, hope enthusiasm optimism · very low-energy, neutral tension, moderately variable, casual)I (low mumble) literally am buzzing for them. (surprised gasp)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as elation, amusement, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.7/10; 23.0s, EN.
540848_00151416 · in -22.8 dBFS · gain +2.8 dB · podcast-05080
(teasing, shame, affection· very low-energy, neutral tension, moderately variable, casual)(ahem) (low mumble) Or even do a like if (wistful sigh) (ahem) you search like a masterclass in threads, there's just really nothing.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing, shame, affection; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.5/10; 26.0s, EN.
540848_00153720 · in -22.9 dBFS · gain +2.9 dB · podcast-05093
(affection, disgust, infatuation·normally alert, neutral tension, fairly steady, casual)And she is one of the best trainers for PDO threads I've ever come across. Love that. Yeah, she's insane. (low mumble) Um, and she's don't this person on stage will look like they've had a stroke. That's how much she's going to lose.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, disgust, infatuation; style: casual, conversational; good recording, no background noise; genuineness 4.7/6; vocal-burst blend 6.5/10; 13.0s, EN.
540848_00156664 · in -22.6 dBFS · gain +2.6 dB · podcast-04322
(jealousy and envy· normally alert, slightly relaxed, moderately variable, conversational)It'll be like, have you just cut skin away and literally sewed it up back there? Because that's it, the reality of the look you can get. You know, non-surgical facelift. It is literally the surgical results
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy; style: conversational, casual; good recording, no background noise; genuineness 3.7/6; vocal-burst blend 6.6/10; 10.6s, EN.
540848_00158096 · in -21.8 dBFS · gain +1.8 dB · podcast-04359
(relief, confusion, astonishment surprise· normally alert, relaxed, moderately variable, conversational)without the knife. Exactly. It's just wild. I know. And again, who else have we got?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, confusion, astonishment surprise; style: conversational, casual; good recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.7/10; 4.7s, EN.
540848_00159176 · in -23.5 dBFS · gain +3.5 dB · podcast-04340
This is a VoiceNet dimension, not an emotion: resonance: oral (R_ORAL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: oral (R_ORAL) low — 0.09, lower than 91 % of clips in this corpus — and ends with it above average at 0.63, higher than 63 % of clips in this corpus. That is a total rise of 0.54.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.15, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.07 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a child strongly feminine voice · fairly smooth
(confusion, contemplation, embarrassment · measured, very low-energy, relaxed, casual)What have I connect with my audience? Like (low mumble) uh like I said, like I don't know, like I already was dreaming of wanting to do social media if that makes sense. So it was I don't know. It just turned into it just being myself, like
full caption & clip details
A child strongly feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as confusion, contemplation, embarrassment; style: casual, whispered; below-average recording, some background noise; genuineness 3.8/6; vocal-burst blend 3.3/10; 16.1s, EN.
840134_00189504 · in -41.0 dBFS · gain +21.0 dB · podcast-05674
(embarrassment, pleasure ecstasy, amusement·fast, energised, neutral tension, casual)that makes sense. I don't know. It was just going with the flow. Because I just liked making I was just doing it for fun.
full caption & clip details
A child feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, bright, fairly smooth, thin; slurred, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, pleasure ecstasy, amusement; style: casual, playful; below-average recording, quiet background; genuineness 5.0/6; vocal-burst blend 6.3/10; 4.5s, EN.
840134_00191240 · in -35.5 dBFS · gain +15.5 dB · podcast-05666
(doubt·measured, normally alert, slightly relaxed, ASMR)Yeah. Do you recommend that they do that with any post or just certain posts? Or what do you think?
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt; style: ASMR, casual; average recording, no background noise; genuineness 3.7/6; vocal-burst blend 3.2/10; 4.8s, EN.
840134_00193840 · in -44.2 dBFS · gain +24.2 dB · podcast-05668
(doubt, contentment, sourness·normal-paced, normally alert, slightly relaxed, casual)When you're posting, do you always try to ask a question so that they can engage? Or th or those just questions saved for certain posts? Like if you're just doing certain posts? Like every lash post, does it have a question? Or are you just saving those questions to engage? Like when you said, (wistful sigh) Oh, uh, I asked them like, hey guys, what do you guys think? Or you know, stuff like that. Do you ask on every post? Like, hey guys.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, contentment, sourness; style: casual, playful; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.4/10; 24.6s, EN.
840134_00194591 · in -40.5 dBFS · gain +20.5 dB · podcast-05676
TEMP — tempo ↓voicenet__VN1__T0.50__C0.25__INTERNAL · #14
This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with tempo (TEMP) above average — 0.72, higher than 72 % of clips in this corpus — and works its way down to low at 0.18, lower than 82 % of clips in this corpus. That is a total fall of 0.54.
It takes 4 clips to get there. Clip to clip the moves are -0.25, then -0.06, then -0.23 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, slightly rough, balanced body, quiet background, fairly steady, somewhat unclear
(thankfulness gratitude, shame, disappointment · normal-paced, normally alert, slightly relaxed, monologue)u jedinici Hrvatskog zavoda za zapošljavanje evidentirano je 315 tisuća 438 nezaposlenih osoba. Zadnji podatak koji se nalazi na web stranici
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, shame, disappointment; style: monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.7/10; 12.0s, HR.
croatia_20120222164809-7106_125456_137456 · in -18.1 dBFS · gain -1.9 dB · eurospeech-01345
(thankfulness gratitude, contentment, infatuation·measured, normally alert, slightly relaxed, monologue)Hrvatskog zavoda za zapošljavanje je negdje oko 339 tisuća. Kad se to usporedi sa brojkom od prije godinu dana ona je otprilike u tom obliku jer u Hrvatskoj mi imamo tzv. "biciklički"
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment, infatuation; style: monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.1/10; 14.6s, HR.
croatia_20120222164809-7106_137456_152096 · in -18.0 dBFS · gain -2.0 dB · eurospeech-01345
(elation, thankfulness gratitude, fatigue exhaustion· measured, subdued, neutral tension, monologue)(ahem) "biciklički" raspored nezaposlenosti koja raste negdje veljača, ožujak nakon toga tijekom proljeća i ljeta pada da bi opet počela rasti krajem godine. I to se događa iz godine u godinu.
croatia_20120222164809-7106_152096_169456 · in -17.6 dBFS · gain -2.4 dB · eurospeech-01345
(shame, fatigue exhaustion, sadness· measured, normally alert, slightly relaxed, monologue)Ono što je problem, problem je da u Hrvatskoj imamo samu strukturu nezaposlenih izrazito nepovoljnu,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as shame, fatigue exhaustion, sadness; style: monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.5/10; 10.4s, HR.
croatia_20120222164809-7106_169456_179887 · in -15.7 dBFS · gain -4.3 dB · eurospeech-01345
This is a VoiceNet dimension, not an emotion: style: whispered (S_WHIS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: whispered (S_WHIS) below average — 0.28, lower than 72 % of clips in this corpus — and ends with it high at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.51.
It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.23, then +0.12 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · balanced body, normally alert, moderate pitch range
(contemplation, interest, relief · brisk, neutral tension, fairly steady, casual)That's right. So one thing that I was realizing is that those people that tend to be against progression and against change, if you think about it, the areas that they tend to live in are areas that are kind of closed off. Back in the day, it was okay to be that way. Everybody around you was that way. You never really came across outsiders because not a lot of people would visit you, but now with the advent of the internet, everybody's connected.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contemplation, interest, relief; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.9/6; vocal-burst blend 10.0/10; 20.6s, EN.
340280_00093232 · in -24.3 dBFS · gain +4.3 dB · podcast-02335
(thankfulness gratitude, elation, jealousy and envy·normal-paced, neutral tension, fairly steady, casual)everyone's opinions are exactly. So now, you know, you're being brought onto all these ideas of other people who already have progressed, (low mumble) um, as far as societal views, and now you're kind of being hit with those new ideas like a truck, and it's scaring you because you're seeing the world change so quickly than what you were used to
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, elation, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 10.0/10; 16.4s, EN.
340280_00095376 · in -24.3 dBFS · gain +4.3 dB · podcast-00605
(sourness, distress, disgust· normal-paced, relaxed, moderately variable, casual)going around you. Which it's bad it's it's bad to be that way, to fear so much that you're hating on other people, but
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly submissive, neutral openness; reads as sourness, distress, disgust; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.1/6; vocal-burst blend 3.9/10; 8.0s, EN.
340280_00097176 · in -24.4 dBFS · gain +4.3 dB · podcast-03152
(impatience and irritability, relief· normal-paced, slightly relaxed, fairly steady, casual)you have to manage it and work through it and understand that to some degree you have to accept it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as impatience and irritability, relief; style: casual, conversational; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.0/10; 5.8s, EN.
340280_00098496 · in -23.3 dBFS · gain +3.3 dB · podcast-00133
This is a VoiceNet dimension, not an emotion: style: ranting (S_RANT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: ranting (S_RANT) below average — 0.25, lower than 75 % of clips in this corpus — and ends with it high at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.55.
It takes 5 clips to get there. Clip to clip the moves are -0.05, then +0.18, then +0.25, then +0.17 — not a clean run: step 1 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.24 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.24, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, normal-paced, normally alert, light breath
(pain · slightly relaxed, steady, no disfluency, formal)off our shores by the end of the year.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.9/10; 3.3s, EN.
331990_00026032 · in -29.1 dBFS · gain +9.1 dB · podcast-05049
(contentment, hope enthusiasm optimism· slightly relaxed, fairly steady, some disfluency, casual)We take a look at news in Melbourne. Uh (low mumble) for those of you who are not from Melbourne or not from Victoria. This weekend we saw some skyrocketing temperatures coming through. Uh (low mumble) for individuals like me who suffer from hay fever, it was a terrible two days. Uh
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, hope enthusiasm optimism; style: casual, whispered; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 13.9s, EN.
331990_00026360 · in -25.9 dBFS · gain +5.9 dB · podcast-05035
(fully relaxed, fairly steady, some disfluency, casual)(low mumble) Uh unfortunately. And obviously you're only allowed to take one tablet a day.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; style: casual, playful; average recording, quiet background; no dominant emotion; genuineness 4.1/6; vocal-burst blend 5.9/10; 3.5s, EN.
331990_00028408 · in -24.4 dBFS · gain +4.5 dB · podcast-01696
(sexual lust, teasing, embarrassment· fully relaxed, moderately variable, some disfluency, casual)should have taken something stronger, like me. And then knock yourself out
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, teasing, embarrassment; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 5.3/6; vocal-burst blend 6.7/10; 3.2s, EN.
331990_00029192 · in -20.2 dBFS · gain +0.2 dB · podcast-01703
(astonishment surprise, embarrassment, elation·neutral tension, moderately variable, some disfluency, casual)for another six hours. I was gonna say the only other option to handle the project is sleeping pill and just sleeping through the whole thing.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, embarrassment, elation; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 7.6/10; 4.8s, EN.
331990_00029527 · in -20.4 dBFS · gain +0.4 dB · podcast-00373
This is a VoiceNet dimension, not an emotion: resonance: nasal (R_NASL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: nasal (R_NASL) low — 0.23, lower than 77 % of clips in this corpus — and ends with it at the very top of the range at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.71.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.09, then +0.23, then +0.16 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an elderly strongly masculine voice
(thankfulness gratitude, pride, triumph · slow, lethargic, fully relaxed, whispered)of your careers, the success
full caption & clip details
An elderly strongly masculine voice; delivery is lethargic, slow, fully relaxed, steady; timbre is slightly warm, dark, rough, very full; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as thankfulness gratitude, pride, triumph; style: whispered, monologue; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 5.8/10; 3.3s.
batch65_part0_batch65_part0_chunk_1582_1_1810379 · in -32.2 dBFS · gain +12.2 dB · snippets-01220
(thankfulness gratitude, relief· slow, normally alert, slightly relaxed, casual)who have given me counsel. The anointing
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as thankfulness gratitude, relief; style: casual, storytelling; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.2/10; 3.2s.
batch65_part0_batch65_part0_chunk_1582_1_1810406 · in -28.5 dBFS · gain +8.5 dB · snippets-01220
(emotional numbness·measured, normally alert, slightly relaxed, casual)talking about bilocation, which is to being at two places at the same time.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: casual, monologue; average recording, no background noise; genuineness 3.8/6; vocal-burst blend 2.7/10; 4.5s.
batch65_part0_batch65_part0_chunk_1582_1_1810426 · in -34.4 dBFS · gain +14.4 dB · snippets-01220
(awe, thankfulness gratitude, affection·brisk, normally alert, neutral tension, casual)in the vision that you've seen which was given by God because you set the Lord before you.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, thankfulness gratitude, affection; style: casual, storytelling; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 3.4/10; 5.6s.
batch65_part0_batch65_part0_chunk_1582_1_1810522 · in -27.7 dBFS · gain +7.7 dB · snippets-01220
(awe, relief, thankfulness gratitude ·measured, normally alert, slightly relaxed, casual)And God lifted him into a place where the where the spirit
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, relief, thankfulness gratitude; style: casual, storytelling; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.5/10; 4.5s.
batch65_part0_batch65_part0_chunk_1582_1_1810547 · in -27.2 dBFS · gain +7.2 dB · snippets-01220
This is a VoiceNet dimension, not an emotion: aesthetic pleasantness (ESTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with aesthetic pleasantness (ESTH) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to below average at 0.40, lower than 60 % of clips in this corpus. That is a total fall of 0.58.
It takes 5 clips to get there. Clip to clip the moves are -0.06, then -0.08, then -0.21, then -0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, slightly relaxed, fairly steady, moderate pitch range, light breath
(normal-paced, normally alert, no disfluency, formal)Its accuracy was therefore a proportiononate to its volume.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.0/10; 3.1s, ZH.
ZH_B00067_S02547_W000327 · in -20.4 dBFS · gain +0.4 dB · emolia-03949
(longing· normal-paced, normally alert, some disfluency, monologue)Her film is always critical of tolsttoys warn peace. It's like the wrote could hardly be faulted.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: monologue, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.3/10; 5.7s, ZH.
ZH_B00067_S02547_W000328 · in -20.0 dBFS · gain +0.0 dB · emolia-03949
(interest, disgust, sourness·measured, very low-energy, some disfluency, monologue)Yet fafailure incorporate what he termed cosmic ideas, gave him an impression of inco heits. We can take cosmic ideas to mean sterreheharfielelfafavor book was a dog of flenders. Can you believe he is quoted best eangg that dog would really give up his life for a painting? Her fuel do does ask the following in a newspaper interview?
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, disgust, sourness; style: monologue, whispered; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 2.3/10; 21.3s, ZH.
ZH_B00067_S02547_W000329 · in -20.5 dBFS · gain +0.5 dB · emolia-03949
(normal-paced, normally alert, some disfluency, casual)Can your most most recent novel, your hero waldo dies twice on mars? And then once again on venus then't that contraictory.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.3/10; 6.8s, ZH.
ZH_B00067_S02547_W000330 · in -16.7 dBFS · gain -3.3 dB · emolia-03949
(confusion, fear· normal-paced, normally alert, some disfluency, monologue)Are you familiar art field replied with how time closing cosmic space? No, the reporter answered, but then no one else is either.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, fear; style: monologue, authoritative; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.8/10; 8.1s, ZH.
ZH_B00067_S02547_W000331 · in -18.5 dBFS · gain -1.5 dB · emolia-03949
This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: monologue (S_MONO) around average — 0.58, higher than 58 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.07, lower than 93 % of clips in this corpus. That is a total fall of 0.51.
It takes 5 clips to get there. Clip to clip the moves are -0.17, then -0.13, then -0.24, then +0.04 — not a clean run: step 4 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.24 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.14 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.24, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth
(astonishment surprise, sourness, awe · normal-paced, normally alert, neutral tension, casual)the smallest town, like literally a one horse fucking town that was 30 minutes away that worked in agriculture, and then also kids that were kind of
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, sourness, awe; style: casual, conversational; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 4.7/10; 11.8s, EN.
851462_00167920 · in -24.6 dBFS · gain +4.6 dB · podcast-03864
(normal-paced, normally alert, slightly relaxed, casual)more part of that working class economy.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; no dominant emotion; style: casual, storytelling; good recording, no background noise; mildly explicit content; genuineness 1.5/6; vocal-burst blend 5.1/10; 3.4s, EN.
851462_00169104 · in -29.0 dBFS · gain +9.0 dB · podcast-03895
(amusement, embarrassment, astonishment surprise· normal-paced, normally alert, neutral tension, casual)(low mumble) Um, and nobody knew what to do with us because we are just outsiders, so we didn't fit in.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as amusement, embarrassment, astonishment surprise; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 2.9/10; 5.7s, EN.
851462_00169444 · in -25.3 dBFS · gain +5.3 dB · podcast-03857
(longing, pleasure ecstasy, embarrassment ·brisk, normally alert, neutral tension, casual)Well, and then after we went back to high school, I feel like we would just pop up every summer. We nobody knew
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as longing, pleasure ecstasy, embarrassment; style: casual, playful; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 7.0/10; 5.7s, EN.
851462_00170048 · in -24.9 dBFS · gain +4.9 dB · podcast-05014
(interest, amusement, pleasure ecstasy · brisk, energised, neutral tension, casual)what was going on. They were like, Oh, they're back. Yeah. Like, hey, let's go drive the four wheeler around and not tell anyone where we're going. And then like, you know, (breathy giggle) it's fine. No, we loved it. We loved it. We're like, this is great. You guys just drive your four wheeler to the gas station. Cool, let's go. Oh yeah.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as interest, amusement, pleasure ecstasy; style: casual, conversational; average recording, some background noise; genuineness 5.5/6; vocal-burst blend 7.5/10; 25.2s, EN.
851462_00170616 · in -24.0 dBFS · gain +4.0 dB · podcast-03858
This is a VoiceNet dimension, not an emotion: recording quality (RCQL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with recording quality (RCQL) around average — 0.55, higher than 55 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.05, lower than 95 % of clips in this corpus. That is a total fall of 0.50.
It takes 4 clips to get there. Clip to clip the moves are -0.19, then -0.10, then -0.21 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult masculine voice · quiet background, neutral tension
(hope enthusiasm optimism, elation, interest · brisk, energised, moderately variable, casual)In Jesus' name. I just bless you guys today. I'm going to encourage you to sow a seed today. I have the information there. You can go into the thing, you can go into the drop box right there. There's a link that goes directly to (low mumble) uh uh to the cash app that I have. So you can just press it, it'll go, it'll go right to the cash app. You can go to the blue fire ministries (ahem) uh dot com slash give and you can give there as well.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, interest; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 9.2/10; 26.2s, EN.
1096_00461524 · in -21.0 dBFS · gain +1.0 dB · podcast-04605
(affection, hope enthusiasm optimism, awe· brisk, energised, moderately variable, dramatic)And just know you can also partner with us monthly. (ahem) Um, if you if you if you (ahem) are blessed and want to bless us that way, just know that that option is available too. But be blessed and be encouraged and be strengthened by the Lord and walk in freedom, walk in righteousness, walk in holiness and allow him to continue to walk (low mumble) uh over your over or
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as affection, hope enthusiasm optimism, awe; style: dramatic, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 9.4/10; 21.8s, EN.
1096_00464135 · in -21.2 dBFS · gain +1.2 dB · podcast-04606
(hope enthusiasm optimism, interest, contentment· brisk, energised, moderately variable, dramatic)(ahem) walk in your life and walk through the things that are in your life to restore it, to restore it. It's like let him do a walkthrough in your house. What is it that you need me to still have here? Does this belong here? Does this please you? This is belong here, Lord. Let him begin to clean house and restore you and restore your family, which I believe the more and more we all do
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, full; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as hope enthusiasm optimism, interest, contentment; style: dramatic, storytelling; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 4.4/10; 23.3s, EN.
1096_00466311 · in -21.2 dBFS · gain +1.2 dB · podcast-04625
(contentment, thankfulness gratitude, pleasure ecstasy·measured, very low-energy, fairly steady, casual)this, the more people will be restored. The people of God will be in their rightful place in all that God wants to restore us in. So be blessed, be blessed. (ahem)
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, thankfulness gratitude, pleasure ecstasy; style: casual, monologue; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.7/10; 23.8s, EN.
1096_00468637 · in -23.1 dBFS · gain +3.1 dB · podcast-05151