Rule.VN1 — VoiceNet: one of 57 voice-descriptor dimensions sweeps by >=T, each consecutive step <=C Source. trajectories_v5.parquet | Family. one corpus in isolation Sampled from 1,152,000 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is a VoiceNet dimension, not an emotion: register (speaking pitch region) (REGS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with register (speaking pitch region) (REGS) around average — 0.51, right about the corpus median — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.39.
It takes 3 clips to get there. Clip to clip the moves are -0.19, then -0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 23 s · en · podcast
hear it un-normalised (raw levels, max seam 2.3 dB)
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, average recording, normal-paced, normally alert, moderately variable, frequent disfluency, light breath
(doubt, shame, contemplation · neutral tension, somewhat unclear, wide pitch range, casual)I don't want that responsibility of when I'm out of the country or when I'm traveling or if I'm working, to have that strong like responsibility of like raising a child. I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as doubt, shame, contemplation; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 5.8/6; vocal-burst blend 9.2/10; 11.4s, EN.
559362_00106600 · in -26.2 dBFS · gain +6.2 dB · podcast-03518
(helplessness, distress, sadness·relaxed, average clarity, moderate pitch range, casual)don't see myself. My parents are bec are gonna become old at some point. They're gonna need help.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as helplessness, distress, sadness; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 3.9/10; 6.4s, EN.
559362_00107752 · in -25.7 dBFS · gain +5.7 dB · podcast-03518
(sexual lust, jealousy and envy, longing·fully relaxed, somewhat unclear, wide pitch range, casual)and I would much rather I would much rather take care of them
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as sexual lust, jealousy and envy, longing; style: casual, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.9/10; 5.2s, EN.
559362_00108488 · in -23.3 dBFS · gain +3.3 dB · podcast-03536
S_PLAY — style: playful ↓c-podcast-VN1 · #2
This is a VoiceNet dimension, not an emotion: style: playful (S_PLAY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: playful (S_PLAY) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to above average at 0.61, higher than 61 % of clips in this corpus. That is a total fall of 0.33.
It takes 5 clips to get there. Clip to clip the moves are -0.18, then -0.19, then -0.10, then +0.14 — not a clean run: step 4 moves back the other way by 0.14 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 62 s · en · podcast
hear it un-normalised (raw levels, max seam 2.3 dB)
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, average clarity
(amusement, embarrassment, teasing · brisk, energised, neutral tension, casual)the first move, you know. But I did slide up on his DMs, so he's gonna always claim a fame that's your girl.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, embarrassment, teasing; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 8.0/10; 5.8s, EN.
269919_00310264 · in -18.3 dBFS · gain -1.7 dB · podcast-05594
(infatuation, embarrassment, longing·normal-paced, normally alert, slightly relaxed, conversational)I was the one that, you know, did the first move when really I feel like it's the person that follows. But anyway.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, embarrassment, longing; style: conversational, casual; good recording, no background noise; mildly explicit content; genuineness 3.9/6; vocal-burst blend 5.6/10; 5.0s, EN.
269919_00310880 · in -20.1 dBFS · gain +0.1 dB · podcast-05563
(contentment, affection, infatuation · normal-paced, very low-energy, slightly relaxed, casual)Yeah, just find that out there. (low mumble) Um, but we met, we had some mutual friends. He was in a freshman Bible study class with (ahem) uh the church that I attended. (ahem) Um, and so, but then he was the student intern pastor at the church. We then went to in Tuscaloosa and just fell in love with (ahem) um a really special family that he actually lived with on campus, and they're now our disciples and mentors.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as contentment, affection, infatuation; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.5/10; 25.4s, EN.
269919_00311720 · in -17.8 dBFS · gain -2.2 dB · podcast-05541
(contentment, affection, disappointment· normal-paced, very low-energy, relaxed, casual)just kind of knew that this wasn't just us, you know, a fluke thing. I knew that this was a sign and Will just became my best friend, and I knew that that was that was the biggest thing. Like you want to find your best friend, and within that is your soulmate. And so
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as contentment, affection, disappointment; style: casual, monologue; average recording, no background noise; genuineness 4.4/6; vocal-burst blend 6.3/10; 14.7s, EN.
269919_00314440 · in -16.9 dBFS · gain -3.1 dB · podcast-05539
(affection, infatuation, contentment · normal-paced, very low-energy, relaxed, casual)We just started hanging out, and (ahem) um, I knew that there was something different about him. He really celebrated all of my dreams and all the things that I wanted to do, and
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, infatuation, contentment; style: casual, conversational; average recording, no background noise; genuineness 3.9/6; vocal-burst blend 5.0/10; 10.6s, EN.
269919_00316088 · in -16.8 dBFS · gain -3.2 dB · podcast-05541
METL — metallic quality ↓c-podcast-VN1 · #3
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with metallic quality (METL) high — 0.87, higher than 87 % of clips in this corpus — and works its way down to low at 0.17, lower than 83 % of clips in this corpus. That is a total fall of 0.70.
It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.23, then -0.07, then -0.20 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.41 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 66 s · en · podcast
hear it un-normalised (raw levels, max seam 5.4 dB)
k 5d_a -0.701d_b -0.701step_a 0.226step_b 0.226min_cos_consec 0.4150min_cos_anchor 0.5303dataset podcastlang enspeaker 912629track 912629total 65.6slevel spread 5.8 dBmax seam 5.4 dBcos from recomputed from spkemb_traj
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, quiet background
(doubt, fear, contemplation · normal-paced, normally alert, neutral tension, casual)I I would I'd be surprised if either one of us could get at least two predictions, right? 'Cause I mean, I thought Cincinnati without (low mumble) uh going into Dallas, Dallas.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, fear, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.5/10; 9.3s, EN.
912629_00392680 · in -34.2 dBFS · gain +14.2 dB · podcast-04135
(fatigue exhaustion, jealousy and envy, longing· normal-paced, normally alert, neutral tension, casual)if you if you go Yeah. But
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as fatigue exhaustion, jealousy and envy, longing; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 6.9/10; 16.1s, EN.
912629_00397888 · in -35.7 dBFS · gain +15.7 dB · podcast-04136
(emotional numbness·measured, normally alert, slightly relaxed, casual)there's only one bye now. So jump into the MLB.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness; style: casual, conversational; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 0.8/10; 3.6s, EN.
912629_00399560 · in -34.2 dBFS · gain +14.2 dB · podcast-04136
(jealousy and envy, sourness, amusement·normal-paced, very low-energy, slightly relaxed, monologue)Next topic. (low mumble) Big news for the Braves this weekend. Braves clinched last night when the Mets beat the Brewers, which we don't like, but it's okay because the Braves won like they always do, except for today because they suck in day games. But with those day games, the Mets did get trounced by the Bluers. Blewers. The Brewers 6 0. So every time the Braves lose, you know who else loses?
full caption & clip details
A child masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, sourness, amusement; style: monologue, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.0/10; 25.0s, EN.
912629_00399936 · in -34.5 dBFS · gain +14.5 dB · podcast-06064
(pride, triumph·measured, subdued, slightly relaxed, whispered)It's all about the Mets. Yeah, so Spencer Strider set a MLB record recording his two hundredth strikeout of the season in a hundred and thirty innings pitched.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, triumph; style: whispered, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.1/10; 11.0s, EN.
912629_00402720 · in -40.0 dBFS · gain +20.0 dB · podcast-04136
GEND — perceived gender ↓c-podcast-VN1 · #4
This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with perceived gender (GEND) at the very top of the range — 0.93, higher than 93 % of clips in this corpus — and works its way down to above average at 0.68, higher than 68 % of clips in this corpus. That is a total fall of 0.26.
It takes 5 clips to get there. Clip to clip the moves are -0.22, then +0.01, then +0.04, then -0.08 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.17 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 50 s · en · podcast
hear it un-normalised (raw levels, max seam 2.9 dB)
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, normal-paced, light breath
(fatigue exhaustion · normally alert, neutral tension, moderately variable, casual)that could be our new so episode three now is is gonna be we're gonna be talking about crime. We're gonna be talking about like
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 4.0/10; 6.4s, EN.
345840_00147576 · in -20.2 dBFS · gain +0.2 dB · podcast-04317
(normally alert, relaxed, moderately variable, casual)Yeah, we'll talk about it. We'll talk about it. We'll put that on our Instagram. Yeah, on right on the right on the gram.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, some background noise; genuineness 5.9/6; vocal-burst blend 7.1/10; 4.6s, EN.
345840_00149488 · in -17.7 dBFS · gain -2.3 dB · podcast-00161
(intoxication altered states of consciousness, contentment, pleasure ecstasy·very low-energy, neutral tension, fairly steady, casual)(low mumble) that's uh it should be there right next to episode two. And uh, (ahem) you know, we would just like to thank our sponsor, Flaming Hot Cheetos, (low mumble) uh, you know, for this episode, letting us do that. (low mumble) Um they actually don't sponsor us yet, but we'd like them to. So if anybody in your family works for Flaming Hot Cheetos, (low mumble) um, you know, give us give us a link us up. Link us
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as intoxication altered states of consciousness, contentment, pleasure ecstasy; style: casual, conversational; below-average recording, quiet background; genuineness 5.6/6; vocal-burst blend 10.0/10; 19.7s, EN.
345840_00152400 · in -20.6 dBFS · gain +0.6 dB · podcast-02397
(energised, neutral tension, moderately variable, casual)it's it's only fans dot com slash forward slash H Baronholtz, and then and then mine is forward slash M civilians
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; no dominant emotion; style: casual, playful; below-average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 2.7/10; 7.7s, EN.
345840_00155624 · in -20.3 dBFS · gain +0.3 dB · podcast-00156
(contentment, pleasure ecstasy, elation·normally alert, relaxed, moderately variable, casual)straightforward. I don't got any I don't really got much. But yeah, we we appreciate your donations for the podcasts. All proceeds get recycled somehow. And um Yeah, (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly guarded; reads as contentment, pleasure ecstasy, elation; style: casual, conversational; below-average recording, some background noise; genuineness 5.5/6; vocal-burst blend 10.0/10; 10.7s, EN.
345840_00156424 · in -17.7 dBFS · gain -2.3 dB · podcast-00460
This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with attack / onset sharpness (ATCK) high — 0.78, higher than 78 % of clips in this corpus — and works its way down to around average at 0.43, lower than 57 % of clips in this corpus. That is a total fall of 0.36.
It takes 3 clips to get there. Clip to clip the moves are -0.11, then -0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.18 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.17 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.18, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 25 s · en · podcast
hear it un-normalised (raw levels, max seam 1.8 dB)
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, some disfluency, average clarity, light breath
(confusion, embarrassment, sexual lust · brisk, normally alert, neutral tension, conversational)I gotta tell you, if if you walked into a into a studio or or into an office and you said, I want to be in this film, I think I would say yes. I is only because I'd be looking at you pointing at me and I said, I don't want to get them angry. So
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as confusion, embarrassment, sexual lust; style: conversational, casual; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 10.0s, EN.
265137_00039032 · in -20.6 dBFS · gain +0.6 dB · podcast-03335
(interest, teasing, affection· brisk, energised, slightly relaxed, casual)what about the do you get the same kick from everything that you that you did? Do do it does it ebb and flow for you sometimes? Do you say, you know what, last year I was really into doing some music, I want to take a break from that one and do some acting here.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, teasing, affection; style: casual, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 5.3/10; 8.6s, EN.
265137_00040120 · in -20.9 dBFS · gain +0.9 dB · podcast-03303
(fatigue exhaustion, embarrassment, sexual lust·measured, normally alert, slightly tense, casual)Absolutely. Like next year I'm gonna go out and and do talking shows all over. In twenty seventeen, I'll need to
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly tense, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as fatigue exhaustion, embarrassment, sexual lust; style: casual, storytelling; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 4.9/10; 6.4s, EN.
265137_00041104 · in -22.7 dBFS · gain +2.7 dB · podcast-03306
R_CHST — resonance: chest ↓c-podcast-VN1 · #6
This is a VoiceNet dimension, not an emotion: resonance: chest (R_CHST) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with resonance: chest (R_CHST) above average — 0.63, higher than 63 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.08, lower than 92 % of clips in this corpus. That is a total fall of 0.55.
It takes 5 clips to get there. Clip to clip the moves are -0.20, then -0.03, then -0.18, then -0.15 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 49 s · en · podcast
k 5d_a -0.551d_b -0.551step_a 0.205step_b 0.205min_cos_consec 0.6835min_cos_anchor 0.6969dataset podcastlang enspeaker 420057track 420057total 49.3slevel spread 2.7 dBmax seam 2.7 dBcos from recomputed from spkemb_traj
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, normally alert, average clarity, light breath
(contentment, thankfulness gratitude, relief · neutral tension, fairly steady, some disfluency, casual)I always will say, you know, I consider myself completely blessed because my kids are so relaxed. My kids have never been considered behavioral. You know,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, thankfulness gratitude, relief; style: casual, conversational; average recording, no background noise; genuineness 3.4/6; vocal-burst blend 2.5/10; 8.7s, EN.
420057_00149152 · in -21.8 dBFS · gain +1.8 dB · podcast-02454
(longing, contemplation, contentment · neutral tension, fairly steady, some disfluency, casual)I've had a harder time raising my other teen daughters than I did with my autistic kids, you know. (low mumble) Um, but they're they're just constantly teaching me new things about life, and they've taught me how to be even Emily taught, you know, teaches me to be more like
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as longing, contemplation, contentment; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 7.4/10; 16.2s, EN.
420057_00150032 · in -21.2 dBFS · gain +1.2 dB · podcast-02438
(relief, contemplation, embarrassment·slightly relaxed, fairly steady, some disfluency, casual)in tune with myself, if that makes sense. Like, I have
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as relief, contemplation, embarrassment; style: casual, conversational; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 5.3/10; 3.0s, EN.
420057_00151776 · in -20.6 dBFS · gain +0.6 dB · podcast-02442
(shame, disappointment, confusion·neutral tension, fairly steady, some disfluency, casual)to really like check myself sometimes too. And she takes things very literally, everything like you can't say any kind of jokes. We learn like you can't say things when you tell her goodnight. Her sister was like, Don't let the bed bugs bite. And she turned around, she's like, What?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is negative, neutral stance, neutral openness; reads as shame, disappointment, confusion; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.4/6; vocal-burst blend 7.0/10; 16.0s, EN.
420057_00152080 · in -21.9 dBFS · gain +1.9 dB · podcast-02435
(impatience and irritability, teasing, astonishment surprise·slightly relaxed, moderately variable, no disfluency, storytelling)What did you say? And I'm like, you cannot do that to her. So
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as impatience and irritability, teasing, astonishment surprise; style: storytelling, dramatic; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.6/10; 4.6s, EN.
420057_00153684 · in -19.2 dBFS · gain -0.8 dB · podcast-02431
TEMP — tempo ↓c-podcast-VN1 · #7
This is a VoiceNet dimension, not an emotion: tempo (TEMP) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with tempo (TEMP) above average — 0.59, higher than 59 % of clips in this corpus — and works its way down to low at 0.19, lower than 81 % of clips in this corpus. That is a total fall of 0.40.
It takes 5 clips to get there. Clip to clip the moves are -0.07, then -0.11, then -0.02, then -0.19 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.32 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.30 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.32, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, fairly steady, moderate pitch range, light breath
(normal-paced, subdued, slightly relaxed, casual)I got it. All right. So T 15 here in 20 or not here, but at the Open Championship T 15 2022, miss cut, miscut, T67, miscut before that. (low mumble) Um kind of similar for the PGA championship. So really he's got a top 20 in every major except for the Masters, which his best finish is a top 21.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 7.9/10; 22.1s, EN.
134603_00185496 · in -27.7 dBFS · gain +7.7 dB · podcast-03098
(triumph, contentment, relief· normal-paced, subdued, slightly relaxed, casual)mean, he hasn't really been this week, though, if you or this year. If you look at his results, we're talking, you know, like I said, just rattled off the last four, a bunch of top 40s, four top 40s in a row for him. So that's not really boom or bust, but (ahem) uh that's what he can be. I feel like his stat profiles profile is so good right now that we could see a big one coming. He could be sitting on a big week. And (low mumble) um I'm hoping it's this week because I'm gonna have some of him on DraftKings for sure.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as triumph, contentment, relief; style: casual, monologue; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.4/10; 25.0s, EN.
134603_00188120 · in -24.6 dBFS · gain +4.6 dB · podcast-06489
(interest, contentment, infatuation· normal-paced, subdued, slightly relaxed, casual)Yeah, no, I actually like Si Wu a lot this week. His name came up a lot when I was doing the first round leader research. (low mumble) Um and I think some of that is gonna obviously (ahem) uh transfer over to tournament success. But yeah, he definitely popped in a lot of like the approach numbers and the putting stuff and and things like that as far as
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as interest, contentment, infatuation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 10.0/10; 18.1s, EN.
134603_00190632 · in -29.2 dBFS · gain +9.2 dB · podcast-03091
(doubt, intoxication altered states of consciousness, sexual lust· normal-paced, normally alert, neutral tension, casual)the greens kind of sneaky good. He's not a good putter, but I think he's got there was a there's a couple stats like approach putt performance. I don't okay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, intoxication altered states of consciousness, sexual lust; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 5.4/6; vocal-burst blend 4.8/10; 8.0s, EN.
134603_00192848 · in -23.3 dBFS · gain +3.3 dB · podcast-03087
(measured, subdued, neutral tension, casual)(ahem) he was high up on that stat, but I know that a lot of guys that had (low mumble) uh similar T to green stuff with approach putt performance being like a high they had success (ahem) uh at the 2016. Tournament here. So (purr) that's kind of a weird stat, kind of a wonky stat, but it makes sense. I mean, like
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 3.9/10; 20.4s, EN.
134603_00193864 · in -26.3 dBFS · gain +6.3 dB · podcast-03121
DARC — darkness of timbre ↓c-podcast-VN1 · #8
This is a VoiceNet dimension, not an emotion: darkness of timbre (DARC) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with darkness of timbre (DARC) high — 0.79, higher than 79 % of clips in this corpus — and works its way down to around average at 0.45, lower than 56 % of clips in this corpus. That is a total fall of 0.34.
It takes 3 clips to get there. Clip to clip the moves are -0.16, then -0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, moderately variable, some disfluency, average clarity, wide pitch range
(hope enthusiasm optimism, elation, pride · brisk, energised, neutral tension, casual)of guys join over the last couple of weeks, which has been fantastic. (ahem) I'm always going to push this or encourage this, and that is we
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, pride; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.8/10; 7.3s, EN.
630156_00005544 · in -23.7 dBFS · gain +3.7 dB · podcast-02798
(hope enthusiasm optimism, affection, elation ·normal-paced, energised, neutral tension, casual)a ton in the group. It's maybe a bit of a different type of group that we have. Uh (ahem) hey, Gerald, good to see you. Um, (ahem)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, affection, elation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.5/10; 5.5s, EN.
630156_00006336 · in -24.0 dBFS · gain +4.0 dB · podcast-02786
(emotional numbness· normal-paced, normally alert, slightly relaxed, conversational)than other groups online where (low mumble) um so maybe I should say this first. Hero Collective is not
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as emotional numbness; style: conversational, casual; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.5/10; 7.0s, EN.
630156_00006896 · in -23.5 dBFS · gain +3.5 dB · podcast-02781
WARM — warmth ↑c-podcast-VN1 · #9
This is a VoiceNet dimension, not an emotion: warmth (WARM) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with warmth (WARM) low — 0.15, lower than 85 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.62.
It takes 5 clips to get there. Clip to clip the moves are +0.13, then +0.15, then +0.10, then +0.25 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.12 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 35 s · en · podcast
k 5d_a 0.623d_b 0.623step_a 0.246step_b 0.246min_cos_consec 0.1229min_cos_anchor 0.0816dataset podcastlang enspeaker 841117track 841117total 34.8slevel spread 5.6 dBmax seam 5.6 dBcos from recomputed from spkemb_traj
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · some disfluency
(disgust, hope enthusiasm optimism, sexual lust · brisk, energised, neutral tension, casual)gentlemen that's all that i have for y'all we're gonna get into these pitches that i
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disgust, hope enthusiasm optimism, sexual lust; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.2/6; vocal-burst blend 5.6/10; 4.3s, EN.
841117_00188584 · in -23.4 dBFS · gain +3.4 dB · podcast-02647
(impatience and irritability, jealousy and envy, intoxication altered states of consciousness· brisk, energised, neutral tension, casual)sean you ain't gonna ask this why we do it clean you ain't gonna ask why the name gentleman i mean dang i saw you talk longer than some of these other folks you know
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as impatience and irritability, jealousy and envy, intoxication altered states of consciousness; style: casual, dramatic; below-average recording, some background noise; mildly explicit content; genuineness 4.6/6; vocal-burst blend 6.6/10; 10.8s, EN.
841117_00189448 · in -18.0 dBFS · gain -2.0 dB · podcast-02681
(triumph, amusement, intoxication altered states of consciousness ·normal-paced, normally alert, neutral tension, casual)the reason why i decided to do it clean and bam can have his his answer but i decided to do it clean because i did do it (low mumble) uh
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as triumph, amusement, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.9/10; 8.8s, EN.
841117_00192312 · in -23.6 dBFS · gain +3.6 dB · podcast-02057
(sadness, malevolence malice, distress· normal-paced, normally alert, slightly relaxed, casual)that you've been making your family laugh mothers dad
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sadness, malevolence malice, distress; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.2/10; 4.9s, EN.
841117_00195248 · in -22.4 dBFS · gain +2.4 dB · podcast-02646
(normal-paced, normally alert, slightly relaxed, casual)aunties uncles you ain't cuss then so it just takes a little more
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.6/10; 5.5s, EN.
841117_00195736 · in -22.1 dBFS · gain +2.1 dB · podcast-02651
SMTH — smoothness ↓c-podcast-VN1 · #10
This is a VoiceNet dimension, not an emotion: smoothness (SMTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with smoothness (SMTH) high — 0.76, higher than 76 % of clips in this corpus — and works its way down to around average at 0.43, lower than 57 % of clips in this corpus. That is a total fall of 0.34.
It takes 3 clips to get there. Clip to clip the moves are -0.16, then -0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.03 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.03, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, balanced body, quiet background, normal-paced, some disfluency, light breath
(infatuation · normally alert, slightly relaxed, fairly steady, conversational)Okay, some woman? (surprised gasp) Okay. That's maybe the only thing I remember.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.0/10; 3.6s, EN.
333078_00072176 · in -16.9 dBFS · gain -3.1 dB · podcast-02705
(infatuation, teasing, amusement· normally alert, relaxed, moderately variable, conversational)Did he kill her? Do you try to kill her? Something like that. (low mumble) Well, she tried to kill him a few times. I think (low mumble) he's he's he's pretty like he's he's (ahem) a quite the character, but he's never really tried to kill anybody, I don't think. So yeah. All right.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as infatuation, teasing, amusement; style: conversational, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.1/10; 12.1s, EN.
333078_00072792 · in -15.7 dBFS · gain -4.3 dB · podcast-02659
(amusement, embarrassment, teasing ·very low-energy, relaxed, moderately variable, casual)think he was married five times twice to the same person. (ahem) You know, so it's yeah, it's a great what a life, man. (breathy giggle) Yeah. Ain't it the life as a foo fighters would say. You probably heard that song at the club,
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as amusement, embarrassment, teasing; style: casual, conversational; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.3/10; 12.6s, EN.
333078_00074112 · in -15.3 dBFS · gain -4.7 dB · podcast-00415
VULN — vulnerability ↑c-podcast-VN1 · #11
This is a VoiceNet dimension, not an emotion: vulnerability (VULN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with vulnerability (VULN) low — 0.13, lower than 87 % of clips in this corpus — and ends with it below average at 0.41, lower than 59 % of clips in this corpus. That is a total rise of 0.29.
It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.18 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, normal-paced, normally alert, fairly steady, some disfluency
(fatigue exhaustion, anger, thankfulness gratitude · slightly relaxed, wide pitch range, storytelling, monologue)Daniel, heute haben wir ja ganz, ganz mächtige Gäste, wie sie so heißt. Die geballte Radprominenz aus Linz mit einem so einem coolen Projekt, das wir uns heute vorstellen.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, anger, thankfulness gratitude; style: storytelling, monologue; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 4.0/10; 18.6s, DE.
575493_00010824 · in -18.7 dBFS · gain -1.3 dB · podcast-04517
(contempt, malevolence malice·neutral tension, wide pitch range, cartoonish, casual)Das heißt umstritten, das war eigentlich fast umkämpft in den letzten Jahren, aber jetzt ist alles andere. Jetzt gebe ich auch schon das Wort weiter. Mir gegenüber sitzt der Klaus, (low mumble) der Daniel und der Sebastian. Wie schaut es denn aus? Was ist da passiert?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, malevolence malice; style: cartoonish, casual; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.2/10; 20.8s, DE.
575493_00012680 · in -18.6 dBFS · gain -1.4 dB · podcast-04500
(shame, thankfulness gratitude, triumph·slightly relaxed, moderate pitch range, monologue, didactic)Weiß ich aus guter und einer eigener Erfahrung. Der Daniel ist ja ein langer Trail-Kollege von mir. Daniel, wir sind ja früher immer viel am Pfanningberg gefahren, aber da ist jetzt alles neu.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, thankfulness gratitude, triumph; style: monologue, didactic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.4/10; 10.8s, DE.
575493_00016856 · in -18.4 dBFS · gain -1.6 dB · podcast-04515
S_NEWS — style: newsreading ↓c-podcast-VN1 · #12
This is a VoiceNet dimension, not an emotion: style: newsreading (S_NEWS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: newsreading (S_NEWS) around average — 0.47, lower than 53 % of clips in this corpus — and works its way down to low at 0.24, lower than 76 % of clips in this corpus. That is a total fall of 0.23.
It takes 5 clips to get there. Clip to clip the moves are -0.00, then -0.16, then -0.08, then +0.02 — not a clean run: step 4 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.11 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, some disfluency, average clarity
(longing · normal-paced, normally alert, slightly relaxed, conversational)(low mumble) Um, well, yes, I think I'll start with last year. Last year I was still in an apartment building a little house, and
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing; style: conversational, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.3/10; 6.8s, EN.
618322_00181440 · in -22.8 dBFS · gain +2.8 dB · podcast-01562
(shame, embarrassment, pain· normal-paced, normally alert, neutral tension, casual)I decided to do Christmas with my ex. And that meant that I bought everything, wrapped everything, did the stockings, cooked all the food, got up at 6 a.m. on Christmas morning, drove to his house,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, embarrassment, pain; style: casual, conversational; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 4.2/10; 12.7s, EN.
618322_00182120 · in -26.8 dBFS · gain +6.8 dB · podcast-01567
(contentment, affection, pleasure ecstasy· normal-paced, normally alert, neutral tension, casual)set it all up before the kids woke up. And the kids came out to this magical Christmas morning that they always had, and I worked my guts out. And after that, (ahem) um,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, affection, pleasure ecstasy; style: casual, conversational; good recording, no background noise; genuineness 4.7/6; vocal-burst blend 6.3/10; 10.6s, EN.
618322_00183392 · in -25.0 dBFS · gain +5.0 dB · podcast-01585
(teasing, impatience and irritability, amusement·brisk, energised, neutral tension, conversational)I was like, girl, you are never doing that again. Never. If I have to tackle you. Right. (breathy giggle)
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, volatile; timbre is neutral-toned, slightly bright, very rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as teasing, impatience and irritability, amusement; style: conversational, casual; good recording, no background noise; mildly explicit content; genuineness 3.2/6; vocal-burst blend 2.0/10; 7.8s, EN.
618322_00185128 · in -19.7 dBFS · gain -0.3 dB · podcast-01564
(disappointment, sadness, relief· brisk, normally alert, neutral tension, casual)I think you you and every other really close friend through this was like, what are you doing? And that was the and then after that, the divorce just got so ugly. It's almost like the more I tried to give and the more I tried to keep the normal traditions going, the worse he behaved. And so I realized that that was the end.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as disappointment, sadness, relief; style: casual, conversational; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.8/10; 21.4s, EN.
618322_00185912 · in -24.8 dBFS · gain +4.8 dB · podcast-06179
METL — metallic quality ↓c-podcast-VN1 · #13
This is a VoiceNet dimension, not an emotion: metallic quality (METL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with metallic quality (METL) around average — 0.46, lower than 54 % of clips in this corpus — and works its way down to low at 0.22, lower than 78 % of clips in this corpus. That is a total fall of 0.24.
It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.97 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.97), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, fairly steady
(hope enthusiasm optimism, contentment, contemplation · subdued, slightly relaxed, frequent disfluency, casual)the idea of moving overseas is a good one, but geez, there's a lot of challenges that comes with it physically and (low mumble) (low mumble) (low mumble) (low mumble) mentally. (exhausted groan)
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as hope enthusiasm optimism, contentment, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 9.9/10; 29.8s, EN.
493871_00082464 · in -26.8 dBFS · gain +6.8 dB · podcast-02447
(affection, contentment, hope enthusiasm optimism ·normally alert, neutral tension, some disfluency, casual)So were you (low mumble) you (low mumble) were you (low mumble) born in New (ahem) (low mumble) South Wales?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, contentment, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 10.0/10; 26.4s, EN.
493871_00085436 · in -29.9 dBFS · gain +9.9 dB · podcast-02430
FOCS — vocal focus ↑c-podcast-VN1 · #14
This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with vocal focus (FOCS) low — 0.16, lower than 84 % of clips in this corpus — and ends with it around average at 0.56, higher than 56 % of clips in this corpus. That is a total rise of 0.41.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.29 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.29 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.29, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: an elderly somewhat feminine voice · neutral-toned, slow, very low-energy, relaxed, moderately variable, frequent disfluency
(contemplation, infatuation, doubt · average clarity, fairly narrow pitch, normal breath, ASMR)people will have something to think. (low mumble) (ahem) Um what do you look forward to the most besides your EP now? What do you maybe let me not ask you? Let me ask it like this. When you close your eyes and envision yourself, how do you how does the story go?
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as contemplation, infatuation, doubt; style: ASMR, whispered; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.3/10; 20.7s, EN.
931778_00051640 · in -31.9 dBFS · gain +11.9 dB · podcast-04675
(contentment, contemplation, elation·somewhat unclear, fairly narrow pitch, audible breath, casual)I see I see my hometown and I see it being a hub for art. Because I feel like people in small towns and people, I come from a rural area as well, and I feel like we don't dream in the entertainment space. And for me,
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is positive, submissive, slightly vulnerable; reads as contentment, contemplation, elation; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.1/10; 17.6s, EN.
931778_00053776 · in -28.7 dBFS · gain +8.7 dB · podcast-03258
(infatuation, amusement, embarrassment·average clarity, moderate pitch range, light breath, casual)It's too far. Uh like this lady (ahem) um I bumped into I think two weeks ago,
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation, amusement, embarrassment; style: casual, conversational; good recording, no background noise; genuineness 3.7/6; vocal-burst blend 1.4/10; 5.6s, EN.
931778_00055640 · in -28.7 dBFS · gain +8.7 dB · podcast-03213
This is a VoiceNet dimension, not an emotion: clarity / intelligibility (CLRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with clarity / intelligibility (CLRT) below average — 0.41, lower than 59 % of clips in this corpus — and ends with it above average at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.31.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.46 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, some disfluency, average clarity, light breath
(teasing, intoxication altered states of consciousness, amusement · brisk, energised, neutral tension, casual)ketchup his crust in the ketchup, he dips his whole pizza has marinara. My logic is you've got tomato base on it. So what's wrong with having a bit of tomato on the bottom as well? But with you, you do not have tea. You don't have tea as a base, do you? Okay, cool. Right.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as teasing, intoxication altered states of consciousness, amusement; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 5.4/6; vocal-burst blend 8.7/10; 16.2s, EN.
149469_00173591 · in -21.3 dBFS · gain +1.3 dB · podcast-06375
(embarrassment·normal-paced, normally alert, neutral tension, casual)we've been we've been saying like quite a lot that (ahem) uh yeah, with the four things on the YouTube. I think we just kinda decided Well, we got given the decision
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 7.6/10; 8.0s, EN.
149469_00175967 · in -23.7 dBFS · gain +3.7 dB · podcast-03533
(normal-paced, normally alert, slightly relaxed, casual)but everything's on Spotify for for you to listen to, so
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 3.6/10; 3.3s, EN.
149469_00177424 · in -24.3 dBFS · gain +4.3 dB · podcast-03530
TENS — tension ↓c-podcast-VN1 · #16
This is a VoiceNet dimension, not an emotion: tension (TENS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with tension (TENS) around average — 0.48, lower than 52 % of clips in this corpus — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.36.
It takes 3 clips to get there. Clip to clip the moves are -0.12, then -0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.51 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.51 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.51, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, some disfluency, average clarity, light breath
(jealousy and envy, doubt, contemplation · brisk, slightly relaxed, fairly steady, casual)cast somebody in a in an incorrect light, but I'm trying to find the their best side. And usually their best side is gonna be their like collaborative side or their open hearted side. And that that's what I'm trying to sort of find, I think.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, doubt, contemplation; style: casual, monologue; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 4.5/10; 10.8s, EN.
151471_00086664 · in -26.0 dBFS · gain +6.0 dB · podcast-01127
(contemplation, disappointment, fatigue exhaustion·normal-paced, neutral tension, moderately variable, casual)Right, totally. Like, yeah, all all the writing I've found from you so far, it's all been pretty positive press. You're not doing hit pieces on artists, you know. (low mumble) Um but even though you don't write net really negative press on people, I'm curious. Do you do you think there's value to negative press on artistic expression? (low mumble) (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, disappointment, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.4/10; 25.1s, EN.
151471_00087776 · in -26.6 dBFS · gain +6.6 dB · podcast-06312
(confusion, embarrassment, sourness· normal-paced, neutral tension, moderately variable, casual)sort of uh generous avenue, but uh you know, I don't or like ask you, you know, say why why certain parts of the conversation were awful. Like I don't know. What what's the point? (ahem) Um I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as confusion, embarrassment, sourness; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 9.1/10; 10.3s, EN.
151471_00096440 · in -24.8 dBFS · gain +4.8 dB · podcast-04890
STRU — structuredness of delivery ↑c-podcast-VN1 · #17
This is a VoiceNet dimension, not an emotion: structuredness of delivery (STRU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with structuredness of delivery (STRU) at the very bottom of the range — 0.03, lower than 97 % of clips in this corpus — and ends with it above average at 0.60, higher than 60 % of clips in this corpus. That is a total rise of 0.57.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.19, then -0.06, then +0.22 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.11 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.11, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body
(embarrassment, infatuation, doubt · normal-paced, very low-energy, relaxed, casual)I think that (contented sigh) I think that it (chuckle) uh (ahem) um I think that it exists apart from feeling, right? You've got
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, very dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, infatuation, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.4/10; 10.3s, EN.
519271_00139384 · in -33.0 dBFS · gain +13.0 dB · podcast-03165
(contemplation, thankfulness gratitude, contentment·measured, very low-energy, relaxed, casual)got what you've got within you, and we're all we're all up to the thing that is before us. And I think there's nothing that we can't handle. (low mumble) Um, given the right the right mindset, the right optimism and (low mumble) um readiness has very little to do with it. I think you can only find out
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as contemplation, thankfulness gratitude, contentment; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 10.0/10; 15.0s, EN.
519271_00140576 · in -34.5 dBFS · gain +14.5 dB · podcast-03157
(doubt, confusion·normal-paced, normally alert, neutral tension, casual)that question. I would say that Hillenberg's (soft hum) suggestion to this almost dark open-ended question is that you of course you're ready. You know, you
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as doubt, confusion; style: casual, conversational; good recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 5.6/10; 8.3s, EN.
519271_00142256 · in -33.5 dBFS · gain +13.5 dB · podcast-03174
(contemplation, awe· normal-paced, normally alert, slightly relaxed, casual)I (surprised gasp) it speaks a little bit to the nature of the human spirit. (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, awe; style: casual, playful; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.0/10; 3.9s, EN.
519271_00143744 · in -34.5 dBFS · gain +14.5 dB · podcast-03182
(doubt, embarrassment, contemplation · normal-paced, normally alert, fully relaxed, casual)(ahem) Like, oftentimes actually we're we're never ready. Like
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, embarrassment, contemplation; style: casual, conversational; average recording, no background noise; genuineness 3.5/6; vocal-burst blend 4.9/10; 4.8s, EN.
519271_00144132 · in -34.9 dBFS · gain +14.9 dB · podcast-03150
S_NEWS — style: newsreading ↑c-podcast-VN1 · #18
This is a VoiceNet dimension, not an emotion: style: newsreading (S_NEWS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.20.
The chain starts with style: newsreading (S_NEWS) at the very bottom of the range — 0.05, lower than 95 % of clips in this corpus — and ends with it below average at 0.28, lower than 72 % of clips in this corpus. That is a total rise of 0.23.
It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.19 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.19, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, average clarity, casual, conversational)ca that's what the bin bag bags are about. You put those up so you don't get the heat in.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 4.0/10; 5.2s, EN.
774376_00640536 · in -35.9 dBFS · gain +15.9 dB · podcast-01716
(astonishment surprise, emotional numbness, doubt·measured, somewhat unclear, conversational, casual)Oh, I thought they were just so you don't. Well,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, emotional numbness, doubt; style: conversational, casual; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 3.4/10; 3.0s, EN.
774376_00641072 · in -29.7 dBFS · gain +9.7 dB · podcast-01746
S_PLAY — style: playful ↑c-podcast-VN1 · #19
This is a VoiceNet dimension, not an emotion: style: playful (S_PLAY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with style: playful (S_PLAY) below average — 0.32, lower than 68 % of clips in this corpus — and ends with it high at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.53.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.21, then +0.11 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · en · podcast
k 4d_a 0.529d_b 0.529step_a 0.210step_b 0.210min_cos_consec 0.7793min_cos_anchor 0.7842dataset podcastlang enspeaker 459544track 459544total 33.0slevel spread 3.9 dBmax seam 3.6 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, normally alert, average clarity
(astonishment surprise, triumph, thankfulness gratitude · neutral tension, fairly steady, some disfluency, casual)and then once once those guys overdosed, we were on the every news station local news, national news. So people now that lived in Baltimore heard that we were still here cleaning and were like, you know, (ahem) uh restaurant owner heard about us, messaged me on Facebook, bought us dinner, breakfast, lunch every day for the rest of the time we were
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as astonishment surprise, triumph, thankfulness gratitude; style: casual; average recording, no background noise; mildly explicit content; genuineness 3.9/6; vocal-burst blend 5.4/10; 20.1s, EN.
459544_00079559 · in -20.6 dBFS · gain +0.6 dB · podcast-06150
(slightly relaxed, fairly steady, frequent disfluency, casual)He hired one of the guys that was help came out of the house and was helping (low mumble) um clean up with us,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 3.5/6; vocal-burst blend 5.2/10; 4.6s, EN.
459544_00081695 · in -20.9 dBFS · gain +0.9 dB · podcast-03247
(helplessness, relief, embarrassment· slightly relaxed, fairly steady, some disfluency, casual)Kid was like, I can't find work around here. (low mumble) Um it just turned out to be just this.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as helplessness, relief, embarrassment; style: casual, conversational; average recording, no background noise; genuineness 4.5/6; vocal-burst blend 5.8/10; 4.1s, EN.
459544_00082456 · in -24.5 dBFS · gain +4.5 dB · podcast-03252
(astonishment surprise, elation, fear·neutral tension, moderately variable, some disfluency, casual)on the news, they were like, Oh my god, this is amazing. How can I gotta help these guys?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, elation, fear; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 7.7/10; 3.7s, EN.
459544_00083216 · in -22.0 dBFS · gain +2.0 dB · podcast-03254
This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.50.
The chain starts with attack / onset sharpness (ATCK) below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the range at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.60.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.19, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.26 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.26, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 28 s · en · podcast
k 4d_a 0.599d_b 0.599step_a 0.211step_b 0.211min_cos_consec 0.1624min_cos_anchor 0.2577dataset podcastlang enspeaker 493671track 493671total 28.0slevel spread 4.0 dBmax seam 4.0 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, moderately variable
(astonishment surprise, awe, amusement · normally alert, slightly relaxed, some disfluency, casual)It's like the opening scene for some porn movie or something, Rich. That's like a dream
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise, awe, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.7/6; vocal-burst blend 0.9/10; 4.4s, EN.
493671_00101112 · in -28.1 dBFS · gain +8.1 dB · podcast-01632
(doubt, contempt, contemplation· normally alert, neutral tension, frequent disfluency, casual)That reminds me, by the way. (ahem) Uh, did anybody get any stimulus checks? I mean, is Trump's name really on there? I mean, if you cash that, doesn't that make you an adult film star? In a stormy Daniels way. Uh,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as doubt, contempt, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.8/6; vocal-burst blend 6.1/10; 14.5s, EN.
493671_00101944 · in -24.1 dBFS · gain +4.1 dB · podcast-01606
(astonishment surprise, confusion, doubt ·energised, neutral tension, frequent disfluency, casual)(surprised gasp) I don't think he signed that check though, right? That was that was uh (low mumble) Cohen.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, confusion, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.9/10; 4.4s, EN.
493671_00103408 · in -26.2 dBFS · gain +6.2 dB · podcast-01649
(teasing, astonishment surprise · energised, slightly tense, frequent disfluency, casual)For the for the porn star or the stimulus. Well, actually.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly tense, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as teasing, astonishment surprise; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.5/10; 4.3s, EN.
493671_00104648 · in -24.5 dBFS · gain +4.5 dB · podcast-01611