Manifest tier. proxy_spearman, rule PXR, T=0.7, step cap 0.25. Population 2 chains (0 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 2.
Rule.PXR — proxy ramp-carrier: same two-sided test as AB2, but the per-step cap is applied on the PROXY axes for emotions that are not directly rampable Source. trajectories_v5.parquet | Family. the tier owner's manifest tiers -- the exact subsets used for training Manifest tier.proxy_spearman__PXR__T0.70__C0.25__INTERNAL — population 2 chains (0 h). SHAREABLE variant: 2. Filter.rule=='PXR' and T==0.2 and C==0.25 and dataset in ['mls', 'eurospeech', 'emolia', 'podcast', 'snippets', 'evasnippets'] and qmax>=0.7 and cmax<=0.25 This tier was resampled. The manifest's earlier published filter applied the one-sided test to every tier regardless of its actual rule. Tier populations were always counted under the correct rule, so the sizes quoted here were never wrong — but the earlier selection drew from a contaminated pool, of which 99.7 % were not members of this tier. These samples are drawn with the corrected filter, which tests qmax >= T and cmax <= C — per-row quantities computed under this tier's own rule. For a proxy tier no d_a/d_b predicate could be correct, because a proxy-rescued chain may legitimately exceed the cap on the named axis while its smoothness is certified on the proxy axis. This tier is tiny. It contains only 2 chain(s) in total, and all of them are shown below — this is the entire tier, not a sample.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above, because a
value that never moves says nothing about a trajectory.
(ahem) — brackets inside the words
are a different thing: a real non-speech sound, printed where it happens. Most clips have
none; about a quarter do.
The full generated caption for any clip is still there, under
“full caption & clip details”. Its perceived-gender and background-noise
clauses were re-rendered from the numeric buckets, because the versions stored in the
corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Fear ↓ / Astonishment Surprise ↑proxy_spearman__PXR__T0.70__C0.25__INTERNAL · #1
This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Astonishment Surprise barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.73.
At the same time Fear goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.25 (lower than 75 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.16, then +0.13, then +0.21 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 47 s · en · emolia
hear it un-normalised (raw levels, max seam 1.4 dB)
k 5d_a -0.713d_b 0.734step_a 0.238step_b 0.224min_cos_consec 0.9157min_cos_anchor 0.9265dataset emolialang enspeaker EN_c5BYOO0j3Fytrack EN_c5BYOO0j3Fytotal 47.2slevel spread 1.4 dBmax seam 1.4 dBqmax 0.713cmax 0.238cos from recomputed from spkemb_traj
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fear, concentration · steady, almost no disfluency, formal, newsreading)This means that the bright hemisphere is visible from Earth when Iapetus is on the western side of Saturn, and that the dark hemisphere is visible when Iapetus is on the eastern side.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, concentration; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.2s, EN.
EN_c5BYOO0j3Fy_W000008 · in -14.9 dBFS · gain -5.1 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue)The Dark Hemisphere was later named Cassini Regio in his honor
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 3.5s, EN.
EN_c5BYOO0j3Fy_W000009 · in -13.4 dBFS · gain -6.6 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue)Iopetus is named after the Titan Iopetus from Greek mythology
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.2/10; 3.6s, EN.
EN_c5BYOO0j3Fy_W000011 · in -14.4 dBFS · gain -5.6 dB · emolia-02588
(emotional numbness· fairly steady, no disfluency, newsreading, formal)The name was suggested by John Herschel – son of William Herschel, discoverer of Mimas and Enceladus – in his 1847 publication Results of Astronomical Observations Made at the Cape of Good Hope, in which he advocated naming the moons of Saturn after the Titans, brothers and sisters of the Titan Cronus – whom the Romans equated with their god Saturn.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.0s, EN.
EN_c5BYOO0j3Fy_W000012 · in -14.8 dBFS · gain -5.2 dB · emolia-02588
(fairly steady, no disfluency, authoritative, formal)When first discovered, Iapetus was among four Saturnian moons labeled the Sedera Lodoisia by their discoverer Giovanni Cassini after King Louis XIV. The other three were Tethys, Dione and Rhea
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.2s, EN.
EN_c5BYOO0j3Fy_W000013 · in -13.7 dBFS · gain -6.3 dB · emolia-02588
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration below average — 0.27, lower than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.71.
At the same time Confusion goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.02 (lower than 98 % of clips in this corpus), a change of -0.70. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.19, then +0.24, then +0.18 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.05 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.05, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 38 s · en · emolia
hear it un-normalised (raw levels, max seam 3.0 dB)
k 5d_a -0.705d_b 0.711step_a 0.239step_b 0.237min_cos_consec 0.0215min_cos_anchor -0.0465dataset emolialang enspeaker EN_8n0eK_7khaktrack EN_8n0eK_7khaktotal 38.3slevel spread 5.3 dBmax seam 3.0 dBqmax 0.705cmax 0.239cos from recomputed from spkemb_traj
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly steady
(measured, subdued, relaxed, casual)250, (low mumble) um, with the (ahem) technical writing and, (low mumble) um.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 1.9/10; 5.1s, EN.
EN_8n0eK_7khak_W000058 · in -16.4 dBFS · gain -3.6 dB · emolia-00318
(helplessness, sadness·normal-paced, normally alert, slightly relaxed, casual)Wordsmithing and so forth. We feel that for that level.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as helplessness, sadness; style: casual, monologue; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 2.1/10; 3.6s, EN.
EN_8n0eK_7khak_W000059 · in -13.9 dBFS · gain -6.1 dB · emolia-00318
(normal-paced, normally alert, slightly relaxed, monologue)And then we would be further directing them to prepare an RFP.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.8/10; 4.4s, EN.
EN_8n0eK_7khak_W000060 · in -16.4 dBFS · gain -3.6 dB · emolia-00318
(normal-paced, normally alert, slightly relaxed, monologue)that would outline the activities that the district needs (ahem) public outreach for.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.2/10; 5.7s, EN.
EN_8n0eK_7khak_W000061 · in -16.2 dBFS · gain -3.8 dB · emolia-00318
(concentration·slow, very low-energy, relaxed, monologue)Okay. So you're asking the board for two different things, I think. (low mumble) Um, one, do we increase the budget from 35,000 to 50,000 to, (low mumble) um, are we, (low mumble) uh, approving the district
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.1/10; 18.9s, EN.
EN_8n0eK_7khak_W000062 · in -19.2 dBFS · gain -0.8 dB · emolia-00318