Manifest tier. emotion, rule B1, T=0.4, step cap 0.25. Population 1,711,911 chains (19,336 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 1,387,607.
Rule.B1 — one-sided: emotion B rises by >=T; the other axis is unconstrained Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.emotion__B1__T0.40__C0.25__INTERNAL — population 1,711,911 chains (19,336 h). SHAREABLE variant: 1,387,607. Filter.rule=='B1' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and abs(d_b)>=0.4 and step_b<=0.25 Sampled from 233,805 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Malevolence Malice barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it around average at 0.57, higher than 57 % of clips in this corpus. That is a total rise of 0.44.
Nothing was asked of the other axis, and in fact Confusion drifts down from 0.84 to 0.18 (-0.66), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.51 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 37 s · en · podcast
hear it un-normalised (raw levels, max seam 5.4 dB)
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, fairly steady, frequent disfluency, moderate pitch range
(measured, subdued, slightly relaxed, monologue)Also it's not a plafond vert, it's not our pain noir that we (low mumble) feel after other days. But
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.0/10; 11.0s, EN.
875194_00356864 · in -33.4 dBFS · gain +13.4 dB · podcast-02371
(sourness, jealousy and envy, contemplation· measured, very low-energy, relaxed, monologue)(low mumble) what's (ahem) what you're thinking concretely, Matt?
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, slightly guarded; reads as sourness, jealousy and envy, contemplation; style: monologue, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.1/10; 22.2s, EN.
875194_00357964 · in -28.1 dBFS · gain +8.1 dB · podcast-02370
(slow, very low-energy, fully relaxed, casual)hauteur on (low mumble) Ineos in the sport (low mumble)
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 3.6/6; vocal-burst blend 1.5/10; 3.3s, EN.
875194_00378072 · in -32.6 dBFS · gain +12.7 dB · podcast-05769
This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Interest around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.42.
Nothing was asked of the other axis, and in fact Doubt drifts down from 0.98 to 0.79 (-0.19), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.12, then +0.09, then -0.04 — not a clean run: step 4 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 45 s · en · emolia
hear it un-normalised (raw levels, max seam 2.4 dB)
k 5d_a -0.191d_b 0.416step_a 0.143step_b 0.245min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_ni6DB4uocOQtrack EN_ni6DB4uocOQtotal 44.9slevel spread 2.4 dBmax seam 2.4 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(doubt · steady, almost no disfluency, light breath, formal)Arise from realistic initial conditions, or is it possible to prove some version of the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: formal, newsreading; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.5/10; 6.3s, EN.
EN_ni6DB4uocOQ_W000068 · in -19.2 dBFS · gain -0.8 dB · emolia-00771
(concentration, doubt, confusion· steady, almost no disfluency, light breath, newsreading)Locality, are there non-local phenomena in quantum physics? If they exist, are non-local phenomena limited to the entanglement revealed in the violations of the Bell inequalities, or can information and conserved quantities also move in a non-local way?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, doubt, confusion; style: newsreading, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.5/10; 12.4s, EN.
EN_ni6DB4uocOQ_W000070 · in -16.8 dBFS · gain -3.2 dB · emolia-00771
(concentration, contemplation·fairly steady, almost no disfluency, light breath, authoritative)Under what circumstances a non-local phenomena observed? What does the existence or absence of non-local phenomena imply about the fundamental structure of spacetime
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation; style: authoritative, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.3/10; 8.4s, EN.
EN_ni6DB4uocOQ_W000071 · in -17.5 dBFS · gain -2.5 dB · emolia-00771
(interest, doubt, confusion· fairly steady, no disfluency, light breath, formal)How does this elucidate the proper interpretation
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, doubt, confusion; style: formal, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.2/10; 4.3s, EN.
EN_ni6DB4uocOQ_W000072 · in -17.5 dBFS · gain -2.5 dB · emolia-00771
(interest · fairly steady, no disfluency, minimal breath, authoritative)Hierarchy problem, why is gravity such a weak force? It becomes strong for particles only at the Planck scale, around 1019 GeV, much above the electroweak scale, 100 GeV, the energy scale dominating physics at low energies
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: authoritative, formal; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 12.9s, EN.
EN_ni6DB4uocOQ_W000074 · in -17.5 dBFS · gain -2.5 dB · emolia-00771
Pride ↑ (unconstrained axis: Disgust)emotion__B1__T0.40__C0.25__INTERNAL · #3
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Pride around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.44.
Nothing was asked of the other axis, and in fact Disgust drifts down from 1.00 to 0.89 (-0.10), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.21, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 61 s · no · eurospeech
hear it un-normalised (raw levels, max seam 1.3 dB)
k 4d_a -0.105d_b 0.437step_a 0.956step_b 0.223min_cos_consec —min_cos_anchor —dataset eurospeechlang nospeaker norway_9673-2track norway_9673-2total 60.5slevel spread 1.3 dBmax seam 1.3 dB
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · neutral-bright, average recording, neutral tension, moderately variable, some disfluency, wide pitch range
(disgust, concentration, shame · normal-paced, normally alert, average clarity, cartoonish)de ønsker seg aller mest, er kjærlighet, at noen er glade i dem, at noen har tid til dem, og at noen passer på dem og viser at de er verdt å være glad i.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disgust, concentration, shame; style: cartoonish, playful; average recording, some background noise; genuineness 3.2/6; vocal-burst blend 1.9/10; 13.0s, NO.
norway_9673-2_147680_160720 · in -23.6 dBFS · gain +3.6 dB · eurospeech-02367
(thankfulness gratitude, shame, relief· normal-paced, normally alert, average clarity, playful)fra veldig seriøse (low mumble) advokater som er inne i disse sakene, og som er bekymret for rettssikkerheten. Det er jeg også, men jeg er aller mest bekymret for at vi i debatten rundt det kaster både ungene og de viktige ansatte ut med badevannet.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as thankfulness gratitude, shame, relief; style: playful, cartoonish; average recording, some background noise; genuineness 3.9/6; vocal-burst blend 3.0/10; 18.4s, NO.
norway_9673-2_171137_189552 · in -24.0 dBFS · gain +4.0 dB · eurospeech-02367
(pride·brisk, normally alert, average clarity, cartoonish)Det er behov for å gå igjennom og styrke barnevernet, men det er også behov for å slå litt ring rundt dem som jobber i barnevernet, og som gjør noen av de aller vanskeligste jobbene
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as pride; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.8/10; 14.2s, NO.
norway_9673-2_189552_203776 · in -23.0 dBFS · gain +3.0 dB · eurospeech-02367
(pride, distress, helplessness·normal-paced, very low-energy, somewhat unclear, casual)som finnes i samfunnet. Og jeg har behov for å si det her nå, at vi som politikere må slå ring om dem som gjør disse veldig vanskelige (ahem) jobbene. Det synes jeg Stortinget skal si i dag.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, distress, helplessness; style: casual, playful; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 5.7/10; 14.4s, NO.
norway_9673-2_203776_218192 · in -24.3 dBFS · gain +4.3 dB · eurospeech-02367
This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Doubt below average — 0.37, lower than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.58.
Nothing was asked of the other axis, and in fact Teasing drifts down from 0.93 to 0.78 (-0.15), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.16, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.25 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.25, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 42 s · en · podcast
hear it un-normalised (raw levels, max seam 1.3 dB)
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average clarity, light breath
(teasing, infatuation, amusement · normal-paced, normally alert, slightly relaxed, casual)you mentioned them previously, but it's the perfect segue. Also Moona.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as teasing, infatuation, amusement; style: casual, conversational; good recording, no background noise; genuineness 3.7/6; vocal-burst blend 3.3/10; 3.6s, EN.
536184_00257048 · in -25.8 dBFS · gain +5.8 dB · podcast-02045
(elation, hope enthusiasm optimism, interest·brisk, energised, slightly relaxed, casual)So let's talk about Moon. How do you feel about their rise? You know, obviously their big self-titled album that came out in 2022 was a huge success, but it is not their first album. They've been around for a while, two albums before that. How do you feel about it? This big spotlight on them.
full caption & clip details
A child masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, interest; style: casual, didactic; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 4.5/10; 19.0s, EN.
536184_00257544 · in -25.0 dBFS · gain +5.0 dB · podcast-06164
(contentment, pleasure ecstasy, hope enthusiasm optimism ·normal-paced, normally alert, neutral tension, casual)I'm always of the mind of like whatever makes the musician happy, I feel like is cool. And they seem to be like doing well with the rise in fame. They're not at like, you know, like chapel roan level.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, pleasure ecstasy, hope enthusiasm optimism; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.8/6; vocal-burst blend 6.7/10; 12.2s, EN.
536184_00259479 · in -23.7 dBFS · gain +3.7 dB · podcast-03710
(doubt, amusement, infatuation· normal-paced, normally alert, neutral tension, casual)point. Yeah, I know. So yeah, I don't know. I think I'm a later fan. I kind of came in with Silk Chiffan.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, amusement, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.7/10; 7.0s, EN.
536184_00260808 · in -23.0 dBFS · gain +3.0 dB · podcast-02050
Sexual Lust ↑ (unconstrained axis: Interest)emotion__B1__T0.40__C0.25__INTERNAL · #5
This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Sexual Lust below average — 0.31, lower than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.62.
Nothing was asked of the other axis, and in fact Interest drifts down from 0.99 to 0.80 (-0.19), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.17, then +0.25, then +0.12 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 78 s · snippets
hear it un-normalised (raw levels, max seam 1.2 dB)
k 5d_a -0.192d_b 0.625step_a 0.869step_b 0.249min_cos_consec —min_cos_anchor —dataset snippetslang ?speaker batch66_part2_batch66_parttrack batch66_part2_batch66_parttotal 77.5slevel spread 1.6 dBmax seam 1.2 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, slightly relaxed, fairly steady, moderate pitch range, light breath
(interest · normal-paced, normally alert, some disfluency, casual)With these radical measures, the Soviet Union succeeded in becoming the largest cotton exporter in the world, but with a very dark side to it. The Aral Sea became one of mankind's worst environmental catastrophes. A gigantic environmental
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest; style: casual, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.7/10; 15.2s.
batch66_part2_batch66_part2_chunk_1595_1_2009723 · in -20.0 dBFS · gain +0.0 dB · snippets-01226
(disappointment, interest, sadness· normal-paced, normally alert, some disfluency, monologue)the inefficient water use, which is rampant in the regions cotton fields, led to insane amounts of water lost to waste. In some places, nearly half the water that was transported by irrigation canals simply seeped through the sand.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, interest, sadness; style: monologue, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.8/10; 14.6s.
batch66_part2_batch66_part2_chunk_1595_1_2009774 · in -19.1 dBFS · gain -0.9 dB · snippets-01226
(measured, normally alert, almost no disfluency, storytelling)started here. It dried up Lake in Eastern Austria.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; no dominant emotion; style: storytelling, casual; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.1/10; 3.6s.
batch66_part2_batch66_part2_chunk_1595_1_2009833 · in -19.4 dBFS · gain -0.6 dB · snippets-01226
(relief, disappointment, interest· measured, very low-energy, some disfluency, monologue)This was not just some pure fantasy dreamed up by some maniac Soviet scientist. No, it was a real plan discussed and planned by Soviet authorities for decades. Luckily, the enormous cost slowed it down. In 1985, the Soviet government began to use Gulag workers for the first conventional demolitions in the Taiga. It was only Mr. Glasnost, Mikhail Gorbachev, who stopped the project in 1986, fearing high costs and immeasurable environmental consequences. Even though only parts of Stalin's original plan were brought to life, large amounts of the damage had already been done.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as relief, disappointment, interest; style: monologue, storytelling; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 5.6/10; 34.8s.
batch66_part2_batch66_part2_chunk_1595_1_2009884 · in -19.5 dBFS · gain -0.5 dB · snippets-01226
(sexual lust, concentration·normal-paced, normally alert, some disfluency, whispered)the construction of a vast system of canals to irrigate large stretches of new farmland across the desert steps of Uzbekistan and Turkmenistan.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, concentration; style: whispered, narration; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.5/10; 8.8s.
batch66_part2_batch66_part2_chunk_1595_1_2010020 · in -20.7 dBFS · gain +0.7 dB · snippets-01226
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.
Nothing was asked of the other axis, and in fact Triumph drifts down from 0.95 to 0.82 (-0.12), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.74 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.74 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 19 s · en · emolia
k 3d_a -0.121d_b 0.417step_a 0.090step_b 0.242min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_ZDbRZ3T-Z5Atrack EN_ZDbRZ3T-Z5Atotal 19.0slevel spread 7.2 dBmax seam 7.2 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · almost no disfluency
(triumph · normal-paced, normally alert, slightly relaxed, authoritative)He once topped the charts with a cover of the Michael Jackson song, Beat It.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph; style: authoritative, monologue; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.5/10; 5.2s, EN.
EN_ZDbRZ3T-Z5A_W000031 · in -16.8 dBFS · gain -3.2 dB · emolia-02585
(contempt, teasing, malevolence malice·brisk, highly aroused, tense, ranting)You guessed it, Gavin is the doppelganger for Weird Al Yankovic!
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, volatile; timbre is cool, slightly bright, very rough, thin; very clear, almost no disfluency, very wide pitch range, normal breath; affect is deeply negative, very dominant, guarded; reads as contempt, teasing, malevolence malice; style: ranting, cartoonish; poor recording, some background noise; mildly explicit content; genuineness 1.7/6; vocal-burst blend 1.8/10; 4.2s, EN.
EN_ZDbRZ3T-Z5A_W000032 · in -9.5 dBFS · gain -10.5 dB · emolia-02585
(hope enthusiasm optimism, elation, thankfulness gratitude· brisk, highly aroused, slightly tense, cartoonish)Thanks for coming in, Gavin, and God bless you, buddy. We have some terrific parting gifts for you, including a full body waxing at Fantastic Sam's!
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; very clear, almost no disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as hope enthusiasm optimism, elation, thankfulness gratitude; style: cartoonish, casual; below-average recording, some background noise; genuineness 2.3/6; vocal-burst blend 1.3/10; 9.3s, EN.
EN_ZDbRZ3T-Z5A_W000033 · in -12.2 dBFS · gain -7.8 dB · emolia-02585
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Emotional Numbness around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.40.
Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.87 to 0.52 (-0.34), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.22, then +0.14 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · en · emolia
k 4d_a -0.345d_b 0.401step_a 0.510step_b 0.224min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00051_S09620track EN_B00051_S09620total 32.5slevel spread 1.9 dBmax seam 1.9 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, steady, almost no disfluency, narration)The first news reports that the garden was found 2,953 feet from Hitler's bunker.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.5/10; 7.1s, EN.
EN_B00051_S09620_W000004 · in -25.4 dBFS · gain +5.4 dB · emolia-01223
(awe·measured, fairly steady, some disfluency, didactic)An excavation uncovered the foundations of the garden's house, two greenhouses, and an underground boiler room that provided water and warm air to grow fresh product all year round.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: didactic, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.0/10; 13.1s, EN.
EN_B00051_S09620_W000005 · in -25.7 dBFS · gain +5.7 dB · emolia-01223
(measured, fairly steady, almost no disfluency, monologue)Wartime ceramics, porcelain, and glassware were also found.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.4/10; 4.6s, EN.
EN_B00051_S09620_W000006 · in -23.8 dBFS · gain +3.8 dB · emolia-01223
(emotional numbness· measured, steady, almost no disfluency, monologue)A vegetarian, tea toddler, and non-smoker, Hitler enjoyed plenty of green produce at the Wolf's Lair.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.6/10; 7.3s, EN.
EN_B00051_S09620_W000007 · in -24.7 dBFS · gain +4.7 dB · emolia-01223
Jealousy and Envy ↑ (unconstrained axis: Triumph)emotion__B1__T0.40__C0.25__INTERNAL · #8
This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Jealousy and Envy below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.58.
Nothing was asked of the other axis, and in fact Triumph drifts down from 0.90 to 0.70 (-0.19), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.08, then +0.14, then +0.15 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of -0.03 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.03 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 50 s · de · emolia
k 5d_a -0.195d_b 0.581step_a 0.492step_b 0.216min_cos_consec —min_cos_anchor —dataset emolialang despeaker DE_WOcDPspIfnotrack DE_WOcDPspIfnototal 49.7slevel spread 9.9 dBmax seam 7.1 dB
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(little disfluency, clear, didactic, formal)Wir müssen tatsächlich gerade hier in Berlin dem institutionellen Rassismus tatsächlich auch den Kampf ansagen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.1/10; 6.1s, DE.
DE_WOcDPspIfno_W000011 · in -15.4 dBFS · gain -4.6 dB · emolia-00001
(confusion, sourness, contemplation·frequent disfluency, somewhat unclear, monologue, casual)Die Erziehung fängt ja ziemlich früh an, selbstverständlich, ob in den Kindergarten oder in den Schulen, dass, (wistful sigh) äh, die Kinder aus den Migrantenfamilien nicht (ahem) isoliert (low mumble) werden, sondern, (low mumble) äh, auch nicht (wistful sigh) als Sonderlinge anzusehen, sondern als, (wistful sigh) äh,
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as confusion, sourness, contemplation; style: monologue, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 16.3s, DE.
DE_WOcDPspIfno_W000012 · in -18.7 dBFS · gain -1.3 dB · emolia-00001
(affection·little disfluency, average clarity, formal, monologue)Kinder, Berliner Familien anzusehen, dementsprechend gleich zu behandeln, gleiche Chancen zu bieten.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as affection; style: formal, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.0/10; 6.1s, DE.
DE_WOcDPspIfno_W000013 · in -16.4 dBFS · gain -3.6 dB · emolia-00001
(distress, longing, anger·some disfluency, average clarity, casual, monologue)Und das glaube ich schon auch die zentrale Forderung, einfach, (low mumble) man merkt es ja in der Stadt so, wem gehört die Stadt so, wer, wer hat einfach Anspruch auf den Raum, so geht's nur ums Wohnen, geht's ums Arbeiten, geht's ums Leben oder geht's um die Frage, wie funktioniert das, (low mumble) äh,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, longing, anger; style: casual, monologue; good recording, no background noise; genuineness 4.5/6; vocal-burst blend 2.5/10; 12.2s, DE.
DE_WOcDPspIfno_W000014 · in -23.5 dBFS · gain +3.5 dB · emolia-00001
(jealousy and envy· some disfluency, average clarity, monologue, casual)gemeinsam an, in dieser Stadt, in diesem (ahem) urbanen Kosmos, den wir haben. Und da ist es natürlich wichtig, dass wir an, an, an die Politik die Forderung stellen, diese Räume.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy; style: monologue, casual; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 0.9/10; 8.5s, DE.
DE_WOcDPspIfno_W000015 · in -25.4 dBFS · gain +5.3 dB · emolia-00001
This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Interest around average — 0.46, lower than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.54.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 1.00 to 0.39 (-0.61), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.22, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.52 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.57 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.52, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult somewhat feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, moderately variable, moderate pitch range, light breath
(infatuation, doubt, embarrassment · measured, relaxed, frequent disfluency, casual)projection map. (low mumble) Um, but I can find somebody in like Lithuania who does. Totally.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation, doubt, embarrassment; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.6/10; 7.5s, EN.
210630_00082904 · in -20.7 dBFS · gain +0.7 dB · podcast-00924
(sexual lust, infatuation, embarrassment ·normal-paced, slightly relaxed, some disfluency, casual)if they need some sort of like fashion styling sort of like thing, they need like, you know,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sexual lust, infatuation, embarrassment; style: casual, conversational; average recording, no background noise; genuineness 4.9/6; vocal-burst blend 8.8/10; 4.1s, EN.
210630_00083904 · in -19.1 dBFS · gain -0.9 dB · podcast-05307
(hope enthusiasm optimism·brisk, neutral tension, some disfluency, conversational)some art direction, I can totally do that. And it's like you could, you know, come up with a whole project, put (ahem) the entire thing together via
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism; style: conversational, casual; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 8.3/10; 8.0s, EN.
210630_00084312 · in -19.0 dBFS · gain -1.0 dB · podcast-05305
(interest, awe, astonishment surprise·normal-paced, slightly relaxed, some disfluency, conversational)(ahem) like literally 20 people can work on the same project right now without ever meeting each other. And that's crazy. That's so crazy because it takes video calls and and that concept of interacting with with technology to a whole nother level. The possibilities are endless.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, awe, astonishment surprise; style: conversational, casual; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 8.5/10; 13.8s, EN.
210630_00085112 · in -19.5 dBFS · gain -0.5 dB · podcast-05308
This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Astonishment Surprise around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.44.
Nothing was asked of the other axis, and in fact Contentment drifts down from 0.97 to 0.24 (-0.74), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.54 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.44 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.54, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, average recording, normal-paced, relaxed, fairly steady, average clarity
(contentment, relief, fatigue exhaustion · subdued, frequent disfluency, moderate pitch range, casual)Yeah. So yeah, so we took (low mumble) uh we drove to Sacramento and hung out there for a few days, and that is something where I r really wanted to try to do different things when we were traveling. Different things that each one of us would find entertaining, right?
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, relief, fatigue exhaustion; style: casual, ASMR; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.5/10; 21.6s, EN.
596320_00041896 · in -26.7 dBFS · gain +6.7 dB · podcast-00771
(longing, contemplation, helplessness·normally alert, frequent disfluency, moderate pitch range, casual)there were some things uh, (low mumble) you know, that I was looking forward to, some things that Ifa was
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, slightly vulnerable; reads as longing, contemplation, helplessness; style: casual, conversational; average recording, no background noise; genuineness 4.3/6; vocal-burst blend 3.2/10; 5.8s, EN.
596320_00044160 · in -27.2 dBFS · gain +7.2 dB · podcast-02027
(astonishment surprise, embarrassment· normally alert, some disfluency, wide pitch range, casual)I think is very weird. Well, as a librarian,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.4/10; 3.1s, EN.
596320_00045080 · in -28.4 dBFS · gain +8.4 dB · podcast-02021
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Infatuation around average — 0.48, lower than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.47.
Nothing was asked of the other axis, and in fact Disgust drifts down from 0.98 to 0.14 (-0.84), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.23, then +0.00, then +0.15 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 39 s · en · emolia
k 5d_a -0.836d_b 0.467step_a 0.526step_b 0.231min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_QKY9QXoCHF8track EN_QKY9QXoCHF8total 39.4slevel spread 0.4 dBmax seam 0.3 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(disgust, malevolence malice · fairly steady, no disfluency, formal, authoritative)The oracle suggested that, in payment for the bear's blood, no Athenian virgin should be allowed to marry until she had served Artemis in her temple, played the bear for the goddess
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, malevolence malice; style: formal, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.0s, EN.
EN_QKY9QXoCHF8_W000211 · in -14.9 dBFS · gain -5.1 dB · emolia-01705
(malevolence malice, awe· fairly steady, almost no disfluency, narration, formal)Boarth Boar is one of the favorite animals of the hunters, and also hard to tame. In honor of Artemis' skill, they sacrificed it to her. Oinius and Adonis were both killed by Artemis Boar
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, awe; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.0s, EN.
EN_QKY9QXoCHF8_W000212 · in -14.7 dBFS · gain -5.3 dB · emolia-01705
(malevolence malice, emotional numbness·steady, no disfluency, formal, narration)Guinea fowlerTemis felt pity for the Meliagrids as they mourned for their lost brother, Meliager, so she transformed them into Guinea fowl to be her favorite animals
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, emotional numbness; style: formal, narration; good recording, no background noise; mildly explicit content; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.4s, EN.
EN_QKY9QXoCHF8_W000213 · in -14.5 dBFS · gain -5.5 dB · emolia-01705
(awe·fairly steady, no disfluency, formal, authoritative)Buzzard hawkhawks were the favored birds of many of the gods, Artemis included
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 4.4s, EN.
EN_QKY9QXoCHF8_W000214 · in -14.6 dBFS · gain -5.4 dB · emolia-01705
(infatuation·steady, no disfluency, formal, authoritative)Palm and Cyprus were issued to be her birthplace. Other plants sacred to Artemis are Amaranth and Asphodel
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.1s, EN.
EN_QKY9QXoCHF8_W000215 · in -14.9 dBFS · gain -5.1 dB · emolia-01705
This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Anger around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.
Nothing was asked of the other axis, and in fact Contentment drifts down from 1.00 to 0.83 (-0.17), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 70 s · en · evasnippets
k 3d_a -0.165d_b 0.425step_a 0.354step_b 0.233min_cos_consec —min_cos_anchor —dataset evasnippetslang enspeaker cond_podcastt_00_03_0001_6track cond_podcastt_00_03_0001_6total 70.0slevel spread 6.2 dBmax seam 6.2 dB
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, slightly rough, balanced body, quiet background
(contentment, thankfulness gratitude, hope enthusiasm optimism · measured, very low-energy, slightly relaxed, didactic)Lord, we have saying about the immense love and forgiveness you offer us. We've saying about the unity that you bring to us, that your blood makes us one. Lord, we've saying about just how wonderful you are and how worthy you are of our praise. And so, Lord, let all those truths just show us that we need to remain pure before you.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, thankfulness gratitude, hope enthusiasm optimism; style: didactic, whispered; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.7/10; 28.0s.
cond_podcastt_00_03_0001_656886_00048288 · in -21.0 dBFS · gain +1.0 dB · evasnippets-00322
(concentration, pain· measured, energised, neutral tension, casual)All right, so that's kind of the the the where the way that we're going. We're unpacking through this passage what is the purity that the people of Israel were being called towards by God and how were they failing in that? Right? That's where we're going. So the Hebrew word for pure is to heart, and it means it's often linked to like flawless gold.
full caption & clip details
A young adult masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, pain; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.3/10; 26.3s.
cond_podcastt_00_03_0001_656886_00052400 · in -14.8 dBFS · gain -5.2 dB · evasnippets-00322
(anger, bitterness, disgust·normal-paced, normally alert, neutral tension, casual)I mean our our marriage would immediately crumble, right? This is what the Israelites were doing to God. God, I don't need you right now. I don't want you right now. I would rather go over here with these people, these gods.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as anger, bitterness, disgust; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.8/10; 15.4s.
cond_podcastt_00_03_0001_656886_00115624 · in -20.7 dBFS · gain +0.7 dB · evasnippets-00316
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Contemplation around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.43.
Nothing was asked of the other axis, and in fact Affection drifts down from 0.95 to 0.24 (-0.70), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then -0.12, then +0.25, then +0.20 — not a clean run: step 2 moves back the other way by 0.12 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.61 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.61 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 37 s · en · emolia
k 5d_a -0.702d_b 0.428step_a 0.353step_b 0.249min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_Ft1uZSgNd-0track EN_Ft1uZSgNd-0total 37.4slevel spread 3.0 dBmax seam 3.0 dB
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed, fairly steady, clear
(affection, infatuation · little disfluency, light breath, monologue, dramatic)For Lessing, the best examples of each type of art take advantage of their inherent connections to time or space.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, infatuation; style: monologue, dramatic; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.2/10; 5.3s, EN.
EN_Ft1uZSgNd-0_W000011 · in -19.5 dBFS · gain -0.5 dB · emolia-01343
(no disfluency, light breath, monologue, casual)And only poorer examples of art try to blur these boundaries.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.3/10; 3.3s, EN.
EN_Ft1uZSgNd-0_W000012 · in -16.5 dBFS · gain -3.5 dB · emolia-01343
(little disfluency, light breath, conversational, playful)As you can probably tell, I don't buy into the Sister Arts Theory very much. The moment it's written down, literature requires space.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.3/10; 6.6s, EN.
EN_Ft1uZSgNd-0_W000013 · in -17.4 dBFS · gain -2.6 dB · emolia-01343
(concentration·some disfluency, light breath, casual, monologue)Likewise, visual art necessarily requires time to understand and interpret it. And in comics, well, space is time.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: casual, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.1/10; 6.8s, EN.
EN_Ft1uZSgNd-0_W000016 · in -19.3 dBFS · gain -0.7 dB · emolia-01343
(contemplation, concentration, interest· some disfluency, normal breath, whispered, monologue)Each panel is a unit of the narrative, and time passes as we move within and between panels, as we move in space. We'll discuss panels another time, properly. This video is about the old stuff. I mean, the really old stuff.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, concentration, interest; style: whispered, monologue; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.8/10; 14.9s, EN.
EN_Ft1uZSgNd-0_W000017 · in -18.4 dBFS · gain -1.6 dB · emolia-01343
Impatience and Irritability ↑ (unconstrained axis: Emotional Numbness)emotion__B1__T0.40__C0.25__INTERNAL · #14
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Impatience and Irritability below average — 0.42, lower than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.58.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.97 to 0.44 (-0.53), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.18, then +0.17, then +0.00 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.20 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.23 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.20, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an adult masculine voice · neutral-toned
(emotional numbness · measured, normally alert, slightly relaxed, narration)standards for acceptable behavior in any number of areas dropped like a rock. Like what what was considered to be normal.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: narration, whispered; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.0/10; 9.0s, EN.
786523_00050688 · in -20.5 dBFS · gain +0.5 dB · podcast-01865
(contemplation, helplessness, affection·normal-paced, normally alert, slightly relaxed, casual)And look, like I'm someone who works from home. You know, I I get that, you know, so I I am not trying to speak from Ivy Ivory Tower here. But at the same time, it's like you look at you know, any number of of different areas, like we we've mentioned air travel, right, which is just an absolute rolling disaster. You know, to the point where it's like, I mean, when is the last time you talked to someone who flew and was not
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, helplessness, affection; style: casual, conversational; good recording, no background noise; genuineness 4.5/6; vocal-burst blend 10.0/10; 22.9s, EN.
786523_00051584 · in -19.9 dBFS · gain -0.1 dB · podcast-01958
(shame, fear, fatigue exhaustion· normal-paced, normally alert, slightly relaxed, casual)majorly delayed, didn't have something, you know, just on a basic level breakdown. Like I when I was flying out to you know an event for the OGC, I was my flight was delayed for several hours because the plane the door fell off the plane on the runway.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as shame, fear, fatigue exhaustion; style: casual, conversational; good recording, no background noise; genuineness 4.2/6; vocal-burst blend 10.0/10; 17.0s, EN.
786523_00053872 · in -20.8 dBFS · gain +0.8 dB · podcast-01862
(astonishment surprise, impatience and irritability, anger·brisk, normally alert, slightly relaxed, casual)Wait, what? You didn't tell me this part. The door fell off the fucking plane on
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as astonishment surprise, impatience and irritability, anger; style: casual, conversational; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.2/10; 3.9s, EN.
786523_00055600 · in -20.1 dBFS · gain +0.1 dB · podcast-01857
(impatience and irritability, anger, bitterness· brisk, energised, slightly tense, casual)the runway with you in the fucking plane. No, no, no, but before anyone got on.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, very rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, guarded; reads as impatience and irritability, anger, bitterness; style: casual, ranting; average recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 5.5/10; 3.6s, EN.
786523_00055984 · in -17.8 dBFS · gain -2.2 dB · podcast-01849
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Relief around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.41.
Nothing was asked of the other axis, and in fact Interest drifts down from 0.92 to 0.85 (-0.07), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.67 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.67 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.67, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(interest · measured, some disfluency, monologue, conversational)E a gente tem ali uma limitação de quantidade de integrantes, né? Então a gente tem essa questão do professor orientador, do capitão, que são peças fundamentais e tem que ter um representante de cada, mas a gente também é importante ressaltar que o limite máximo de integrantes é 20, né? Sim. A gente tem
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: monologue, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 6.0/10; 22.1s, PT.
492301_00034069 · in -19.6 dBFS · gain -0.4 dB · podcast-03730
(interest, hope enthusiasm optimism, elation·normal-paced, frequent disfluency, monologue, casual)máximo de 20 pessoas, né? Que é limitado por inscrição. E aí dentro dessas 20 pessoas, cada equipe se divide de uma forma os trabalhos, né?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, hope enthusiasm optimism, elation; style: monologue, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.5/10; 26.0s, PT.
492301_00036384 · in -17.6 dBFS · gain -2.4 dB · podcast-06160
(relief, triumph, pride·measured, frequent disfluency, monologue, authoritative)Legal de ressaltar também o fato de que não precisa de nenhum conhecimento prévio, né? Então não tem um período mínimo da faculdade para a pessoa poder participar do projeto Baja, né? Basta ela ter força de vontade,
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, triumph, pride; style: monologue, authoritative; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 3.3/10; 14.5s, PT.
492301_00047697 · in -18.4 dBFS · gain -1.6 dB · podcast-00023
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.40.
Nothing was asked of the other axis, and in fact Sourness drifts down from 1.00 to 0.79 (-0.21), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.17, then +0.00, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 5 clips: an elderly masculine voice · slightly rough, below-average recording, quiet background, normally alert, slurred, moderate pitch range
(sourness, disgust, contempt · fast, neutral tension, moderately variable, casual)Tu puedes creer lo que se te da tu regalada gana, pero eso no lo hace verdad. Tu puedes creer que hay unicornios en el lado oscuro de la luna, pero pues esa creencia es una creencia. Punto. No es conocimiento, no es este episteme, dijera Sócrates.
full caption & clip details
An elderly masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disgust, contempt; style: casual, conversational; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.5/10; 16.4s, ES.
126006_00095656 · in -31.2 dBFS · gain +11.2 dB · podcast-00365
(jealousy and envy, sourness, bitterness·measured, neutral tension, fairly steady, casual)Y hay una tercera opción, una tercera vía que vino a sintetizar o a ser popular o no sé cómo llamar (low mumble) San Agustín. No fue el único, ni fue el primero, pero fue el más popular y el que lo hizo mejor. O mejor precisamente porque lo hizo mejor, se volvió más popular, no. Es la mezcla de ambas.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, sourness, bitterness; style: casual, monologue; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.5/10; 21.0s, ES.
126006_00097288 · in -29.3 dBFS · gain +9.3 dB · podcast-00367
(affection, contentment, awe· measured, neutral tension, fairly steady, monologue)De esa medida, o in esa manera (low mumble) comenzó a surgir, por decir de alguna forma, la filosofía de la religión, o la filosofía cristiana, in this caso in particular. Es decir, la religión cristiana, al menos la católica, tiene fundamentos racionales muy fuertes. Santo Tomás, perdón, San Tomás, (surprised gasp) San Anselmo, Tu San Scotto, Guillermo de Oca,
126006_00101956 · in -31.4 dBFS · gain +11.4 dB · podcast-00378
(fast, slightly relaxed, fairly steady, casual)Incluso por si no lo sabían en el Vaticano existe el observatorio astronómico del Vaticano O sea, es uno de los observatorios astronómicos o un telescopio
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, thin; slurred, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.7/10; 11.3s, ES.
126006_00106820 · in -34.0 dBFS · gain +14.0 dB · podcast-00369
(pain, sadness·measured, neutral tension, moderately variable, casual)que desarrolla la iglesia católica va de la mano con la razón de tal manera que la iglesia misma busca poder explicar por ejemplo como es que Dios Jesucristo que creemos que es Dios
126006_00112628 · in -33.0 dBFS · gain +13.1 dB · podcast-00371
Pride ↑ (unconstrained axis: Hope Enthusiasm Optimism)emotion__B1__T0.40__C0.25__INTERNAL · #17
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Pride around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.44.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.90 to 0.40 (-0.49), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.25, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
Unchanged across all 4 clips: a young adult feminine voice · slightly cool, fairly smooth, quiet background
(normal-paced, normally alert, slightly relaxed, casual)either the wedding date or the cutting edge. Both rom coms. We're gonna stick with the rom com spirit this. Well,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.9/10; 7.8s, EN.
6714_00010192 · in -22.2 dBFS · gain +2.2 dB · podcast-00083
(jealousy and envy, impatience and irritability, embarrassment·brisk, energised, neutral tension, casual)And (low mumble) um, do you have any more announcements? Oh, also, I'd like to remind every single person who voted for this movie that's listening. I I don't like you right now. I'm I'm grateful that
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as jealousy and envy, impatience and irritability, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 7.1/10; 12.0s, EN.
6714_00011792 · in -19.5 dBFS · gain -0.5 dB · podcast-05175
(pride, thankfulness gratitude, fatigue exhaustion·normal-paced, normally alert, neutral tension, casual)We were gonna do Cherry Two Thousand. I really wanted to. We're doing it next week. I swear to God, it has to be better than this. We watched my boss's daughter. Why did I buy this on DVD?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is negative, slightly dominant, neutral openness; reads as pride, thankfulness gratitude, fatigue exhaustion; style: casual, dramatic; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.0/10; 12.1s, EN.
6714_00013608 · in -21.4 dBFS · gain +1.4 dB · podcast-05197
(pride ·brisk, energised, neutral tension, casual)boss's daughter, two thousand three. Four point seven out of ten stars on IMDB.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pride; style: casual, storytelling; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.5/10; 4.5s, EN.
6714_00015264 · in -21.4 dBFS · gain +1.4 dB · podcast-00082
This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Malevolence Malice below average — 0.38, lower than 62 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.48.
Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.83 to 0.52 (-0.30), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.18, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.80, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 20 s · en · emolia
k 4d_a -0.301d_b 0.483step_a 0.155step_b 0.179min_cos_consec 0.7589min_cos_anchor 0.7975dataset emolialang enspeaker EN_B00036_S00395track EN_B00036_S00395total 19.8slevel spread 2.2 dBmax seam 2.2 dBcos from recomputed from spkemb_traj
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(fairly steady, almost no disfluency, narration, formal)Skunks cover a surprising amount of ground, as much as three miles a night in search of food.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.8/10; 5.4s, EN.
EN_B00036_S00395_W000030 · in -19.2 dBFS · gain -0.8 dB · emolia-00944
(fear· fairly steady, no disfluency, formal, monologue)She also discovered that skunks can have nine or more different dens.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.0/10; 4.1s, EN.
EN_B00036_S00395_W000031 · in -17.7 dBFS · gain -2.3 dB · emolia-00944
(steady, almost no disfluency, formal, narration)Katie, for example, slept in this series of dens over the course of 16 days.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 5.4s, EN.
EN_B00036_S00395_W000032 · in -19.9 dBFS · gain -0.1 dB · emolia-00944
(fairly steady, no disfluency, narration, formal)Skunks sometimes share dens in the winter to conserve heat by snuggling.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 4.5s, EN.
EN_B00036_S00395_W000033 · in -19.1 dBFS · gain -0.9 dB · emolia-00944
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Concentration around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.44.
Nothing was asked of the other axis, and in fact Confusion drifts down from 0.79 to 0.18 (-0.61), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.10, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.75 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.75 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 38 s · en · emolia
k 4d_a -0.609d_b 0.445step_a 0.448step_b 0.209min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_fTcLGNxJnPUtrack EN_fTcLGNxJnPUtotal 38.2slevel spread 1.5 dBmax seam 1.5 dB
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(measured, frequent disfluency, average clarity, casual)This is the syngas. The only issue with syngas is you have a slurry. The feed is usually a slurry. Feed here is a slurry.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.8/10; 8.6s, EN.
EN_fTcLGNxJnPU_W000050 · in -17.1 dBFS · gain -2.9 dB · emolia-00432
(normal-paced, some disfluency, average clarity, didactic)Which is only having 60 to 70 percent of coal, rest remaining water. So, you are spending some amount in evaporating the liquid water. Then, (low mumble) you have the outlet temperature. If I want to compare the outlet temperature.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.8/10; 12.8s, EN.
EN_fTcLGNxJnPU_W000051 · in -18.6 dBFS · gain -1.4 dB · emolia-00432
(intoxication altered states of consciousness·measured, frequent disfluency, somewhat unclear, didactic)Outlet temperature is quite similar around 1500 to 1800 Kelvin and the pressure is pretty high as compared to shell.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: didactic, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.9/10; 8.4s, EN.
EN_fTcLGNxJnPU_W000052 · in -17.8 dBFS · gain -2.2 dB · emolia-00432
(concentration· measured, frequent disfluency, somewhat unclear, didactic)pressure is around 30 to 80 bar is varies. And the cold gas efficiency, this is important, the cold gas efficiency is.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.1/10; 8.0s, EN.
EN_fTcLGNxJnPU_W000053 · in -18.2 dBFS · gain -1.8 dB · emolia-00432
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Awe)emotion__B1__T0.40__C0.25__INTERNAL · #20
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.40. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.52.
Nothing was asked of the other axis, and in fact Awe drifts down from 0.99 to 0.47 (-0.51), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.20, then +0.10 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.76 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.76 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 23 s · en · emolia
k 4d_a -0.511d_b 0.523step_a 0.518step_b 0.223min_cos_consec —min_cos_anchor —dataset emolialang enspeaker EN_B00038_S09760track EN_B00038_S09760total 23.2slevel spread 1.0 dBmax seam 1.0 dB
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, measured, slightly relaxed, fairly steady
(awe, emotional numbness · energised, almost no disfluency, moderate pitch range, narration)They lift up onto the step and skim the water's surface on just a tiny patch of hull.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as awe, emotional numbness; style: narration, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 5.8s, EN.
EN_B00038_S09760_W000043 · in -21.7 dBFS · gain +1.7 dB · emolia-00972
(malevolence malice, emotional numbness ·normally alert, almost no disfluency, fairly narrow pitch, formal)The airfish borrows this design using its own step to push out of the water and to get to takeoff speed quicker.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, emotional numbness; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.3s, EN.
EN_B00038_S09760_W000044 · in -22.4 dBFS · gain +2.4 dB · emolia-00972
(awe, fear, astonishment surprise· normally alert, almost no disfluency, moderate pitch range, narration)When it hits 40 miles per hour, the wings take over and lift the ship into the air.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as awe, fear, astonishment surprise; style: narration, storytelling; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.5/10; 5.3s, EN.
EN_B00038_S09760_W000045 · in -21.4 dBFS · gain +1.4 dB · emolia-00972
(normally alert, no disfluency, moderate pitch range, formal)With clear water ahead, Kenneth increases the speed.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.8/10; 3.4s, EN.
EN_B00038_S09760_W000046 · in -22.2 dBFS · gain +2.2 dB · emolia-00972