rare-lang

The 10 scarcest languages, 2 chains each (supply 1,685-11,067).

Rule. mixed
Source. trajectories_v5.parquet  |  Family. the scarcest cells -- the ones the owner said matter most
Sampled from 9,704,889 matching rows, without replacement across the family, so no two tiers reuse a chain.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above, because a value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words are a different thing: a real non-speech sound, printed where it happens. Most clips have none; about a quarter do.

The full generated caption for any clip is still there, under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
S_RANT — style: rantingrare-lang · #1

The chain starts with style: ranting (S_RANT) around average — 0.47, lower than 53 % of clips in this corpus — and ends with it above average at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 40 s · et · podcast

hear it un-normalised (raw levels, max seam 0.3 dB)
k 2d_a 0.249d_b 0.249step_a 0.249step_b 0.249min_cos_consec 0.9185min_cos_anchor 0.9185dataset podcastlang etspeaker 17320track 17320total 40.4slevel spread 0.3 dBmax seam 0.3 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, balanced body, quiet background, normal-paced, moderately variable
(intoxication altered states of consciousness, pride, teasing · very low-energy, relaxed, frequent disfluency, casual) Planeerimine või kulgemine. (low mumble) Jõvad alati varem või alati (low mumble) (breathy giggle) hilaks. Aga no see selleks on. (low mumble) Merts. Sügis või kevad.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, pride, teasing; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.8/10; 21.9s, ET.
17320_00407424 · in -32.0 dBFS · gain +12.0 dB · podcast-03078
(pride, infatuation, contentment · normally alert, neutral tension, some disfluency, casual) Aga ei. Kulga ka. Ma ei tea küll, see on väga maata mõtlen, et mis järeldusi ja inimesed nendes küsimustes teevad, et kass või koer, aga ma loodan, et me siis kuuleme mingi hetke teiste tagasi, et mis ära selle kohta tehmis vastuseid tains,
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as pride, infatuation, contentment; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.8/10; 18.3s, ET.
17320_00410608 · in -32.2 dBFS · gain +12.2 dB · podcast-01667
TENS — tensionrare-lang · #2

The chain starts with tension (TENS) high — 0.85, higher than 85 % of clips in this corpus — and works its way down to above average at 0.63, higher than 63 % of clips in this corpus. That is a total fall of 0.23.

It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.44 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.44 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.44, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · et · podcast

hear it un-normalised (raw levels, max seam 1.3 dB)
k 2d_a -0.228d_b -0.228step_a 0.228step_b 0.228min_cos_consec 0.4389min_cos_anchor 0.4389dataset podcastlang etspeaker 427800track 427800total 34.0slevel spread 1.3 dBmax seam 1.3 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, fairly smooth, some disfluency, average clarity, light breath
(normal-paced, normally alert, slightly relaxed, didactic) Übülise seadusandlik mikrobaneerimise näida, aga Sandra, mõtleme positiivselt. US saas tahab president reguleeritud tursi otsikuid.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: didactic, whispered; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.4/10; 8.3s, ET.
427800_00249600 · in -20.8 dBFS · gain +0.8 dB · podcast-04794
(elation, amusement, sourness · brisk, energised, neutral tension, casual) Noh, täpselt, et sinna me ei peaks ise (low mumble) minema. Mul näiteks ei ole väga palju andeid, aga mul on erakordne talent ennast (ahem) kogemata, tahtmatult vigastada kogu aeg. Minu vaates tuleks ära keelata, teravad noadest iga korkuma süüa teen, siis (low mumble) ma lõikan näppu endale praktiselt iga kord siingi mõned haavad (ahem) (low mumble) kasva kinima, ma saanud näita, marju enda küünarnuki, mis on peaaene.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation, amusement, sourness; style: casual, dramatic; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 7.6/10; 25.6s, ET.
427800_00250431 · in -19.5 dBFS · gain -0.5 dB · podcast-06044
S_NEWS — style: newsreadingrare-lang · #3

The chain starts with style: newsreading (S_NEWS) below average — 0.36, lower than 64 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 9 s · ro · podcast

hear it un-normalised (raw levels, max seam 1.1 dB)
k 2d_a -0.246d_b -0.246step_a 0.246step_b 0.246min_cos_consec 0.7191min_cos_anchor 0.7191dataset podcastlang rospeaker 10997track 10997total 9.3slevel spread 1.1 dBmax seam 1.1 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
(fast, conversational) să crezi că mai livrează și știri serioase. Și le îmbra și într-o formare (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational; average recording, no background noise; genuineness 4.0/6; vocal-burst blend 1.3/10; 5.2s, RO.
10997_00066408 · in -32.8 dBFS · gain +12.8 dB · podcast-03843
(emotional numbness · measured, playful, conversational) (low mumble) de fapt îmbracă (low mumble) (low mumble) cel puțin parțial (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: playful, conversational; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.2/10; 4.0s, RO.
10997_00067620 · in -33.9 dBFS · gain +13.9 dB · podcast-03863
R_ORAL — resonance: oralrare-lang · #4

The chain starts with resonance: oral (R_ORAL) low — 0.17, lower than 83 % of clips in this corpus — and ends with it below average at 0.40, lower than 60 % of clips in this corpus. That is a total rise of 0.23.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 45 s · ro · podcast

hear it un-normalised (raw levels, max seam 0.0 dB)
k 2d_a 0.234d_b 0.234step_a 0.234step_b 0.234min_cos_consec 0.9250min_cos_anchor 0.9250dataset podcastlang rospeaker 434088track 434088total 45.2slevel spread 0.0 dBmax seam 0.0 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · slightly cool, neutral-bright, brisk, neutral tension, moderately variable, some disfluency, average clarity, wide pitch range
(contempt, disgust, elation · energised, normal breath, casual, dramatic) (ahem) ajutor. La statului să (ahem) pacea. Strategile naționale pe hi și pe droguri nu au buget Liviu. De ce vorbim? Deci au pus acolo acum, strategia națională pe Hiv, nu avem care trebuie să includă preveniere, care trebuie să includă consumatori de droguri la avenă. Deci nu avem din 2008, și acum au publicat o fără buget.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as contempt, disgust, elation; style: casual, dramatic; below-average recording, noisy background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 9.0/10; 20.6s, RO.
434088_00343256 · in -16.1 dBFS · gain -3.9 dB · podcast-00243
(astonishment surprise, elation, disgust · normally alert, light breath, casual, dramatic) Nu știu. Găsesc la primul festival, ca să facem (ahem) să încercăm (ahem) rope cei de la de lup din ei să vină să facem, știți? Să facem o testare. Apoi, poliția cu câini sau poliția fără câini la festivaluri, crește numărul de supradoze, pentru că și în ghi, două sau trei doze odată. Se sperie când văd poliția. Se întâmplă chestia asta. Sunt niște chestii de care trebuie să ținem cont. Dăm apă la festival.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as astonishment surprise, elation, disgust; style: casual, dramatic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 9.8/10; 24.5s, RO.
434088_00345792 · in -16.1 dBFS · gain -3.9 dB · podcast-04203
S_CASU — style: casualrare-lang · #5

The chain starts with style: casual (S_CASU) at the very bottom of the range — 0.06, lower than 94 % of clips in this corpus — and ends with it below average at 0.30, lower than 70 % of clips in this corpus. That is a total rise of 0.24.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · sk · podcast

hear it un-normalised (raw levels, max seam 0.8 dB)
k 2d_a 0.240d_b 0.240step_a 0.240step_b 0.240min_cos_consec 0.8762min_cos_anchor 0.8762dataset podcastlang skspeaker 130275track 130275total 24.2slevel spread 0.8 dBmax seam 0.8 dBcos from orange-id (speaker identity)
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady, light breath
(brisk, almost no disfluency, clear, dramatic) Vočíta si, že ak byt úspešne predá, zisk sa vyrovná piatim platom, ktoré zaráva na brigách zábovní.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; no dominant emotion; style: dramatic, storytelling; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 2.7/10; 7.9s, SK.
130275_00067928 · in -21.6 dBFS · gain +1.6 dB · podcast-05193
(sourness, malevolence malice, anger · measured, some disfluency, slurred, storytelling) Potešený výpočt vyhľadá na internete jeden z inzerátov na svoj byt. Niečo sa mu nestá. Fotky na webe od realitnej kancelárie sú tmavé, neostré, z kompozí a popisov nie je jasné, kde sa čo nachádza.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as sourness, malevolence malice, anger; style: storytelling, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 4.1/10; 16.2s, SK.
130275_00068716 · in -20.9 dBFS · gain +0.8 dB · podcast-05189
S_AUTH — style: authoritativerare-lang · #6

The chain starts with style: authoritative (S_AUTH) around average — 0.43, lower than 57 % of clips in this corpus — and works its way down to low at 0.22, lower than 78 % of clips in this corpus. That is a total fall of 0.21.

It takes 2 clips to get there. Clip to clip the moves are -0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · sk · podcast

k 2d_a -0.212d_b -0.212step_a 0.212step_b 0.212min_cos_consec 0.9306min_cos_anchor 0.9306dataset podcastlang skspeaker 265177track 265177total 33.4slevel spread 1.6 dBmax seam 1.6 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady
(thankfulness gratitude, contentment, affection · normally alert, audible breath, monologue, whispered) To je (low mumble) ten.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment, affection; style: monologue, whispered; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 7.0/10; 15.6s, SK.
265177_00102748 · in -29.9 dBFS · gain +9.8 dB · podcast-01316
(subdued, light breath, monologue, casual) A (low mumble) on dal ten obraz také církev jako skoro taká uzavretá záhrada, nebo trošku to prijá nejaká archa je No, kde my sa plávíme na potopé sveta a všetci ostatní sú zatopené, zahnuté, potrestaný.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 7.2/10; 17.6s, SK.
265177_00104308 · in -28.2 dBFS · gain +8.2 dB · podcast-01346
AGEV — perceived agerare-lang · #7

The chain starts with perceived age (AGEV) around average — 0.49, right about the corpus median — and ends with it above average at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.29 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.29 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.29, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 48 s · uk · podcast

k 2d_a 0.247d_b 0.247step_a 0.247step_b 0.247min_cos_consec 0.2936min_cos_anchor 0.2936dataset podcastlang ukspeaker 110375track 110375total 48.4slevel spread 1.8 dBmax seam 1.8 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, balanced body, quiet background, normally alert, neutral tension, some disfluency, light breath
(elation, hope enthusiasm optimism, pleasure ecstasy · brisk, moderately variable, average clarity, casual) кіпіш, не роздуває, ви головне приїжджаєте. Треба розуміти, що зараз актуально і про це комунікувати, тому що бізнес чомусь це робить, а (low mumble) місто якось воно таке велика, повільна машина, мені здається. Хоча воно змінюється, і я кажу, я бачу дуже позитивні якісь напрямки, я думаю, що буде розвиватися. І варто теж сказати, що ми з Андрієм не є якісь експерти міського
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.4/10; 24.7s, UK.
110375_00238975 · in -24.6 dBFS · gain +4.6 dB · podcast-02021
(triumph, awe, hope enthusiasm optimism · normal-paced, fairly steady, somewhat unclear, casual) В них кльова історія. В них та ж сама кльова архітектура, в них купа крутих людей. В них ідеальна штука. Я розумію, що це супер некомфортно для місцевих жителів, але тур по умовних двориках Одеси, в яких ти реально я стояв (low mumble) 20-30 хвилин, щоб хтось вийшов з дворика, щоб туди потрапити.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, awe, hope enthusiasm optimism; style: casual, conversational; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 7.3/10; 23.6s, UK.
110375_00243752 · in -22.9 dBFS · gain +2.9 dB · podcast-01017
AGEV — perceived agerare-lang · #8

The chain starts with perceived age (AGEV) below average — 0.33, lower than 67 % of clips in this corpus — and ends with it around average at 0.58, higher than 58 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 55 s · uk · podcast

k 2d_a 0.250d_b 0.250step_a 0.250step_b 0.250min_cos_consec 0.9307min_cos_anchor 0.9307dataset podcastlang ukspeaker 131417track 131417total 55.0slevel spread 0.1 dBmax seam 0.1 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, quiet background, normally alert, neutral tension, moderately variable, some disfluency
(teasing, pleasure ecstasy, affection · normal-paced, casual, conversational) Наскільки розумію, він початься, я сама не їла, його просто в мене багато колег з Галичини, я дуже багато чула про цей витвір.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, pleasure ecstasy, affection; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 7.0/10; 25.2s, UK.
131417_00294512 · in -18.2 dBFS · gain -1.8 dB · podcast-06167
(teasing, pleasure ecstasy, amusement · brisk, conversational, casual) Тепер (ahem) всі (breathy giggle) (wistful sigh) галичани, які слухають цей подкаст, мусяться спробувати.
full caption & clip details
A child feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, pleasure ecstasy, amusement; style: conversational, casual; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 4.5/10; 29.6s, UK.
131417_00297672 · in -18.4 dBFS · gain -1.6 dB · podcast-06174
EMPH — emphasis / stress strengthrare-lang · #9

The chain starts with emphasis / stress strength (EMPH) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to above average at 0.69, higher than 69 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.02 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 45 s · hu · podcast

k 2d_a -0.248d_b -0.248step_a 0.248step_b 0.248min_cos_consec -0.0161min_cos_anchor -0.0161dataset podcastlang huspeaker 118697track 118697total 45.0slevel spread 0.8 dBmax seam 0.8 dBcos from orange-id (speaker identity)
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-bright, fairly smooth, average recording, quiet background, normally alert, average clarity, moderate pitch range, light breath
(triumph, sourness, pain · measured, slightly relaxed, fairly steady, monologue) Hogy képzeltek el egy gyonít. Hát én abból az idők inspirálódni a játékból abból a társasból a similóból, a mitológiás similóból. Abban ilyen ilyen ördögszerű szarvai vannak, piros bőre, ilyen valami valami furcsa ilyen, majdnem úgy néz ki az öve, mint egy ilyen szumós öv. Nem mondasz rosszakat.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, sourness, pain; style: monologue, storytelling; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 7.2/10; 28.0s, HU.
118697_00039512 · in -23.0 dBFS · gain +3.0 dB · podcast-01017
(shame, embarrassment, longing · normal-paced, neutral tension, moderately variable, casual) Hát nekem bármok teljesen otthon a japán kultúrában, de valahogy úgy képzem el, hogy ez ilyen ördögmaszkos. Démonta, hogy valami maszk van rajta. Valamiért az van a fejem vagy a japánok erőszületette használnak különféle maszkokat.
full caption & clip details
A child somewhat masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, embarrassment, longing; style: casual, cartoonish; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.3/10; 16.8s, HU.
118697_00043064 · in -22.2 dBFS · gain +2.2 dB · podcast-01003
S_CART — style: cartoonishrare-lang · #10

The chain starts with style: cartoonish (S_CART) low — 0.16, lower than 84 % of clips in this corpus — and ends with it below average at 0.41, lower than 59 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 42 s · hu · podcast

k 2d_a 0.248d_b 0.248step_a 0.248step_b 0.248min_cos_consec 0.8284min_cos_anchor 0.8284dataset podcastlang huspeaker 158518track 158518total 41.7slevel spread 1.6 dBmax seam 1.6 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · neutral-toned, slightly rough, below-average recording, quiet background, measured, fairly steady, frequent disfluency, somewhat unclear
(fatigue exhaustion, intoxication altered states of consciousness, infatuation · very low-energy, relaxed, audible breath, casual) (ahem) egy beruházás szinten (low mumble) mondanám az első körben, ami azt jelenti, hogy (low mumble) most úgy tűnik, hogy az egyházi tanács megszavazta az őszi. (low mumble)
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, very dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, submissive, neutral openness; reads as fatigue exhaustion, intoxication altered states of consciousness, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 5.2/6; vocal-burst blend 7.5/10; 13.1s, HU.
158518_00231112 · in -23.1 dBFS · gain +3.1 dB · podcast-01542
(elation, pride, concentration · subdued, neutral tension, light breath, casual) A költségves egy 24 költségvetési tervben, hogy induljon (low mumble) (low mumble) elsen egy általos iskola, (low mumble) ha ennek meglesznek a feltételei, tehát, ha megtaláljuk az épületet, meg stb. stb. akkor úgy tűnik, hogy nagyon jó lenne, hogyha ez el tudna indulni, ez megint csak egy nagyon új terület lenne, megvan rá, az az emberi, (low mumble)
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as elation, pride, concentration; style: casual, monologue; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 10.0/10; 28.4s, HU.
158518_00232423 · in -24.6 dBFS · gain +4.6 dB · podcast-01523
METL — metallic qualityrare-lang · #11

The chain starts with metallic quality (METL) low — 0.21, lower than 79 % of clips in this corpus — and ends with it around average at 0.45, lower than 55 % of clips in this corpus. That is a total rise of 0.24.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 45 s · cs · podcast

k 2d_a 0.239d_b 0.239step_a 0.239step_b 0.239min_cos_consec 0.9554min_cos_anchor 0.9554dataset podcastlang csspeaker 124271track 124271total 44.7slevel spread 0.4 dBmax seam 0.4 dBcos from orange-id (speaker identity)
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, fairly steady
(contentment, elation, shame · normal-paced, normally alert, some disfluency, monologue) hodin jako diskuzem na jednotlivých tématech v rámci toho kurzu, protože to jsem vlastně možná jenom chtěl doplnit. Potom, co se týče role učitelů, tak to je velmi komplikovaná a jako dost hluboká otázka. Protože já si myslím, že ta role, že (low mumble) ta role učitele se bude hodně měnit.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contentment, elation, shame; style: monologue, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 10.0/10; 19.9s, CS.
124271_00125783 · in -21.2 dBFS · gain +1.2 dB · podcast-05242
(relief, elation, triumph · measured, subdued, frequent disfluency, monologue) A říkám, tady jako nejsem expert, ale co vidím, tak tyhle nástroje přinášejí systémy, přinášejí neuvěřitelnou možnost (low mumble) personalizace. A to platí pro tu výuku. Zkrátka dobře, že ten model má (low mumble) nekonečně mnoho času a nekonečně mnoho trpělivosti na to vysvětlit (low mumble) látku, což učitel nemá.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, elation, triumph; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 8.1/10; 24.6s, CS.
124271_00127768 · in -20.8 dBFS · gain +0.8 dB · podcast-06204
STNC — stance / assertivenessrare-lang · #12

The chain starts with stance / assertiveness (STNC) below average — 0.33, lower than 67 % of clips in this corpus — and ends with it above average at 0.58, higher than 58 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.53 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · cs · podcast

k 2d_a 0.248d_b 0.248step_a 0.248step_b 0.248min_cos_consec 0.5325min_cos_anchor 0.5325dataset podcastlang csspeaker 161379track 161379total 31.9slevel spread 2.7 dBmax seam 2.7 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, fairly steady, some disfluency, moderate pitch range, light breath
(triumph, relief, disappointment · normal-paced, very low-energy, neutral tension, casual) Tak ono je teď opravdu buzzlink. Jsou tam šuda lidi, je to takový těcký. (ahem) Ale dostáváte se do čtvrtě do míst, kde vidíte, že najednou tyhle baráčky, které jsou tak na sebe přilepovaný a do výšky stavěný, tak najednou se začínají zmenšovat, zvětšování. Takže opravdu to tady vypadá honosně mnohem všechno.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, relief, disappointment; style: casual, monologue; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 8.0/10; 24.3s, CS.
161379_00049092 · in -35.4 dBFS · gain +15.4 dB · podcast-00939
(sadness, shame, contemplation · fast, normally alert, slightly relaxed, casual) problémů, který jsem vám stal. Řekli bychom, nechme říct hodně, protože asi tady hodně užší můžeme vyjdocin.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, shame, contemplation; style: casual, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 3.2/10; 7.5s, CS.
161379_00061855 · in -32.7 dBFS · gain +12.7 dB · podcast-03338
ARSH — harshness of articulationrare-lang · #13

The chain starts with harshness of articulation (ARSH) above average — 0.58, higher than 58 % of clips in this corpus — and works its way down to below average at 0.33, lower than 67 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · fi · eurospeech

k 2d_a -0.249d_b -0.249step_a 0.249step_b 0.249min_cos_consec min_cos_anchor dataset eurospeechlang fispeaker finland_taysistunto-10-201track finland_taysistunto-10-201total 25.1slevel spread 1.1 dBmax seam 1.1 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-bright, fairly smooth, average recording, quiet background, normally alert, some disfluency, light breath
(pain, contempt · brisk, neutral tension, moderately variable, cartoonish) Asianajajaliiton mukaan Suomeen ei ole tarpeen säätää erillistä äitiyslakia, vaan tarvittavat muutokset olisi mahdollista tehdä jo olemassa oleviin asiaan liittyviin lakeihin.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pain, contempt; style: cartoonish, playful; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.1/10; 12.2s, FI.
finland_taysistunto-10-2018_1154336_1166496 · in -8.7 dBFS · gain -11.3 dB · eurospeech-00005
(sourness, disgust, contempt · normal-paced, slightly relaxed, fairly steady, monologue) Se kysyykin perustellusti, monimutkaistaisiko tämä lainsäädäntö tarpeettomasti oikeusjärjestystä. Myöskään kansainvälisesti ei tunneta erillisiä äitiyslakeja.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, disgust, contempt; style: monologue, didactic; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.4/10; 12.8s, FI.
finland_taysistunto-10-2018_1166496_1179280 · in -9.8 dBFS · gain -10.2 dB · eurospeech-00005
FOCS — vocal focusrare-lang · #14

The chain starts with vocal focus (FOCS) high — 0.80, higher than 80 % of clips in this corpus — and works its way down to around average at 0.55, higher than 55 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · fi · eurospeech

k 2d_a -0.245d_b -0.245step_a 0.245step_b 0.245min_cos_consec min_cos_anchor dataset eurospeechlang fispeaker finland_taysistunto-10-201track finland_taysistunto-10-201total 30.4slevel spread 1.9 dBmax seam 1.9 dB
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · slightly cool, neutral-bright, fairly smooth, thin, average recording, quiet background, brisk, normally alert
(contempt, bitterness, disgust · ranting, dramatic) Täällä on otettu esimerkiksi, (wistful sigh) kuinka vähän synnytyskuolemia on, ja tällä tavalla yritetty perustella lapsen oikeusturvattomuutta, mitä te edustatte muun muassa siellä vastustajien joukossa. On hyvä, että Suomessa
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, bitterness, disgust; style: ranting, dramatic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.2/10; 17.1s, FI.
finland_taysistunto-10-2018_3654895_3672015 · in -11.8 dBFS · gain -8.2 dB · eurospeech-00005
(pride, thankfulness gratitude, jealousy and envy · dramatic, casual) jopa pidempiä aikoja. Me emme tiedä, mitä sillä välillä voi tapahtua, ei ole tilastoja. Voi tulla autokolareita ja muita, ja silloin näillä lapsilla ei ole enää huoltajaa. (ahem) He jäävät vaille oikeusturvaa ja yhdenvertaista asemaa siinä.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pride, thankfulness gratitude, jealousy and envy; style: dramatic, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.6/10; 13.1s, FI.
finland_taysistunto-10-2018_3683312_3696400 · in -9.8 dBFS · gain -10.2 dB · eurospeech-00005
R_MIXD — resonance: mixedrare-lang · #15

The chain starts with resonance: mixed (R_MIXD) above average — 0.69, higher than 69 % of clips in this corpus — and works its way down to around average at 0.47, lower than 53 % of clips in this corpus. That is a total fall of 0.22.

It takes 2 clips to get there. Clip to clip the moves are -0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · sl · eurospeech

k 2d_a -0.224d_b -0.224step_a 0.224step_b 0.224min_cos_consec min_cos_anchor dataset eurospeechlang slspeaker slovenia_slovenia_10_Izredtrack slovenia_slovenia_10_Izredtotal 24.6slevel spread 0.9 dBmax seam 0.9 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly dark, slightly rough, balanced body, average recording, quiet background, slightly relaxed, somewhat unclear
(sadness, helplessness, thankfulness gratitude · measured, subdued, fairly steady, monologue) kar je po našem mnenju strokovno utemeljeno in ustrezno in je bilo usklajeno tudi z deležniki. Glede na navedeno tudi ne podpiramo amandmajev poslanske skupine SDS.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, helplessness, thankfulness gratitude; style: monologue, whispered; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.2/10; 13.0s, SL.
slovenia_slovenia_10_Izredna_0_14072022_1942735_1955712 · in -24.6 dBFS · gain +4.6 dB · eurospeech-00014
(thankfulness gratitude · normal-paced, normally alert, steady, monologue) Ostale določbe zakona so zelo tehnične oziroma specifične narave in v veliki meri povzemajo dosedanjo ureditev, dodajo le še natančen si zapis v angleškem jeziku.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 11.5s, SL.
slovenia_slovenia_10_Izredna_0_14072022_1955712_1967232 · in -23.8 dBFS · gain +3.8 dB · eurospeech-00014
TENS — tensionrare-lang · #16

The chain starts with tension (TENS) around average — 0.51, higher than 51 % of clips in this corpus — and ends with it high at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · sl · eurospeech

k 2d_a 0.246d_b 0.246step_a 0.246step_b 0.246min_cos_consec min_cos_anchor dataset eurospeechlang slspeaker slovenia_slovenia_10_Izredtrack slovenia_slovenia_10_Izredtotal 29.6slevel spread 0.3 dBmax seam 0.3 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-bright, balanced body, average recording, brisk, energised, neutral tension, moderately variable, some disfluency
(bitterness, disgust, thankfulness gratitude · casual, playful) poslanske skupine SDS to sedaj popravljamo, upamo, da boste tudi te amandmaje z dobršno mero racionalnosti tudi podprli. Zdaj, kaj pa je dejansko intenca tega vašega zakona?
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, disgust, thankfulness gratitude; style: casual, playful; average recording, some background noise; genuineness 4.5/6; vocal-burst blend 8.3/10; 11.7s, SL.
slovenia_slovenia_10_Izredna_0_14072022_4352559_4364304 · in -23.4 dBFS · gain +3.4 dB · eurospeech-00014
(contempt, anger, bitterness · cartoonish, authoritative) Želite si imeti nadzor nad vrtci in osnovnimi šolami preko vašega sindikata SVIZ, želite imeti policijsko ministrico, ki obvladuje slovensko policijo in verjetno tudi morebiti naroča politične pregone, želite si imeti strokovne komisije na Ministrstvu za kulturo, ki bodo sestavljene iz predstavnikov vam naklonjenih
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, anger, bitterness; style: cartoonish, authoritative; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 5.6/10; 17.7s, SL.
slovenia_slovenia_10_Izredna_0_14072022_4364304_4382016 · in -23.0 dBFS · gain +3.0 dB · eurospeech-00014
GEND — perceived genderrare-lang · #17

The chain starts with perceived gender (GEND) around average — 0.48, lower than 52 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.22.

It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 28 s · polish · mls

k 2d_a 0.219d_b 0.219step_a 0.219step_b 0.219min_cos_consec min_cos_anchor dataset mlslang polishspeaker 10900track 10900|nasrebrnymglobie_00_total 27.9slevel spread 0.3 dBmax seam 0.3 dB
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(pride · monologue, formal) twierdził że są to papiery zwęglone których treść on dopiero odczytuje za pomocą sztucznych zdjęć fotograficznych robionych z wielkim trudem i największą ostrożnością ta tajemniczość wzbudzała podejrzenia zwłaszcza że asystent (ahem) do tego czasu nie powiedział jak doszedł do posiadania kuli
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: monologue, formal; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.4/10; 17.3s, POLISH.
10900_6473_001038 · in -29.7 dBFS · gain +9.7 dB · mls-00109
(awe, pride, disappointment · monologue, authoritative) ale zaciekawienie wzrastało ciągle czekano z pewnem niedowierzaniem przyrzeczonych wyjaśnień a tymczasem zaczęto sobie przypominać ze współczesnych pism dzieje całej wyprawy
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pride, disappointment; style: monologue, authoritative; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 1.3/10; 10.4s, POLISH.
10900_6473_001235 · in -30.0 dBFS · gain +10.0 dB · mls-00109
CHNK — chunking / phrasing densityrare-lang · #18

The chain starts with chunking / phrasing density (CHNK) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to above average at 0.74, higher than 74 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · polish · mls

k 2d_a -0.241d_b -0.241step_a 0.241step_b 0.241min_cos_consec min_cos_anchor dataset mlslang polishspeaker 10900track 10900|nasrebrnymglobie_00_total 23.1slevel spread 0.0 dBmax seam 0.0 dB
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(pride, triumph, relief · normal-paced, some disfluency, monologue) kamienia nie znaleziono gdyż upadł w trzęsawisko i prawdopodobnie zagłębił się znacznie ale jeśli go chcę mieć on jest gotów dać mi kilku robotników dla jego wydobycia
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, triumph, relief; style: monologue; good recording, quiet background; mildly explicit content; genuineness 1.2/6; vocal-burst blend 0.0/10; 10.2s, POLISH.
10900_6473_002220 · in -29.7 dBFS · gain +9.7 dB · mls-00109
(shame, fear · measured, almost no disfluency, storytelling, narration) kamień naturalnie chciałem mieć i uwolniwszy się na parę dni z obserwatorjum pojechałem sam na miejsce w celu poszukiwań ale mimo niewątpliwych znaków i usilnej pracy nie mogliśmy nic znaleść
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, fear; style: storytelling, narration; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.4/10; 12.7s, POLISH.
10900_6473_001396 · in -29.7 dBFS · gain +9.7 dB · mls-00109
FOCS — vocal focusrare-lang · #19

The chain starts with vocal focus (FOCS) around average — 0.54, higher than 54 % of clips in this corpus — and works its way down to below average at 0.29, lower than 71 % of clips in this corpus. That is a total fall of 0.24.

It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · pl · podcast

k 2d_a -0.242d_b -0.242step_a 0.242step_b 0.242min_cos_consec 0.7168min_cos_anchor 0.7168dataset podcastlang plspeaker 143769track 143769total 32.5slevel spread 0.8 dBmax seam 0.8 dBcos from orange-id (speaker identity)
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · neutral-bright, fairly smooth, slightly thin, average recording, quiet background, energised, neutral tension, some disfluency
(embarrassment, distress, confusion · brisk, moderately variable, average clarity, cartoonish) mówiłaś. Nie no, spoko i to dobrze, to ja was zostawiam, ale zaraz jak wrócę i pokazujecie, co tam macie. Dobrze, dobrze. Ja pierdolę. Szybko wziąłyśmy po prostu jakiejś (ahem) igłę nic jakoś szmatę zaczęły się w niej wyszywać serduszko. Mama, coś tam i imię i w ogóle.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, distress, confusion; style: cartoonish, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.3/10; 18.4s, PL.
143769_00130464 · in -29.3 dBFS · gain +9.3 dB · podcast-05298
(embarrassment, shame, affection · fast, volatile, slurred, casual) Ona wraca tak pokoje. O, to dla pani. To chodziło. A my to byłyśmy takie obsrane.
full caption & clip details
A child feminine voice; delivery is energised, fast, neutral tension, volatile; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; slurred, some disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, vulnerable; reads as embarrassment, shame, affection; style: casual, storytelling; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.8/10; 13.9s, PL.
143769_00132300 · in -28.5 dBFS · gain +8.5 dB · podcast-05300
ROUG — roughnessrare-lang · #20

The chain starts with roughness (ROUG) high — 0.86, higher than 86 % of clips in this corpus — and works its way down to above average at 0.62, higher than 62 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · pl · podcast

k 2d_a -0.247d_b -0.247step_a 0.247step_b 0.247min_cos_consec 0.9020min_cos_anchor 0.9020dataset podcastlang plspeaker 17179track 17179total 22.8slevel spread 1.4 dBmax seam 1.4 dBcos from orange-id (speaker identity)
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(confusion, embarrassment, shame · slightly relaxed, casual, conversational) Wiadomo jest to, że przyjdzie druga fala epidemii, bo są wakacje, ludzie wyjechali, ludzie się mieszają. To jest normalne, że przyjdzie druga fala epidemia, ale no myślę, że tak nie warto pomagać tej drugiej fali taki. Gdzie się
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, embarrassment, shame; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 7.0/10; 10.9s, PL.
17179_00093960 · in -16.9 dBFS · gain -3.1 dB · podcast-01395
(embarrassment, pleasure ecstasy, shame · neutral tension, casual, conversational) bardzo wiele rok, to możecie sobie zrobić przed weselem, no co za problem. No właśnie, więc to jest strasznie. Rozumiem, to jest planowanie, jakieś zaliczki i tak dalej, no ale chyba od pieniędzy ważniejsze jest trochę zdrowie bliskich, nawet jeżeli.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, pleasure ecstasy, shame; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 5.3/10; 11.8s, PL.
17179_00096216 · in -18.4 dBFS · gain -1.6 dB · podcast-01398