Listening grids — mined emotional and VoiceNet trajectories
92 tiers · 1,764 trajectories · 6,477 clips · 20.5 h — each chain rendered as one concatenation you can play straight through, with per-clip playback and transcripts.
New — a voice-conversion-corrected version of the emotion tiers.
Every chain on this site is separate recordings stitched together, so the voice can be
heard shifting between segments. A conversion pass has now re-voiced
3,438 segments across 1,304 emotion chains onto their own chain's first segment,
with ChatterboxVC + SIDON, and published them as a
parallel set — the pages
here are unchanged and every corrected card carries both renders side by side.
It is not a clean win and the measurements say so.
The corrected tiers → ·
what was done and what it measured →
And a third build — the strict, no-voice-conversion path. For chains that
already clear a 0.80 speaker-identity cut, conversion
costs similarity rather
than adding it. This set is built by
filtering instead: similarity ≥ 0.80 on
both conditions, 150 ms equal-power crossfades, chain-level normalisation, no
conversion. It carries the counts of how many trajectories survive —
3,378,525 mined chains at T≥0.20 and 1,399,128
at T≥0.25, plus 949,212 voice-profile chains.
The strict subset →
What a trajectory is. An ordered list of speech
clips from one speaker/track in which a scored dimension moves monotonically. The rule
fixes what must move and by how much: a threshold T on the total
move end-to-end, and a cap C on each consecutive step, so the chain ramps
rather than jumping. k is the number of clips.
What you are being asked to judge. Play the
concatenation and ask whether the thing the rule claims is moving actually moves, and
whether it moves smoothly rather than lurching at one seam. Then compare tiers:
does a higher T sound like a bigger change? Does a longer k
sound smoother or just longer? Do the proxy (PXR) chains hold up as well as
the direct ones?
New: why Sadness produced nothing, and four
rules that fix it. Sadness, Awe, Distress and Disappointment return
zero
chains under the strict rule. The reason is measurable and it is not what it looks like.
There is a full plain-language walkthrough with the diagnosis, four alternative rules,
their yields and their costs, plus 180
listenable rescued examples.
Read it → Loudness has been normalised, deliberately.
The MOSS encoder normalises every clip to −20 dBFS independently and 59.5 %
of clips hit its ±3 dB clamp, so a raw concatenation carries level steps of
up to 8 dB that belong to the pipeline, not the trajectory. Each clip here is scaled
to −20 dBFS RMS exactly before joining, and the peak guard is applied
once to the finished chain so the level relationship between clips is preserved. The gain
applied is printed under every clip, and the first 5 items of each tier
(444 in total) also carry an un-normalised render so the raw seam can still
be heard. A 150 ms silence marks each boundary.
Manifest tiers — the subsets he will train on
the tier owner's manifest tiers -- the exact subsets used for training
| tier | rule | k | T | corpus | population | sampled from | shown | clips | audio | same speaker? | |
|---|
emotion__B1__T0.20__C0.25__INTERNAL | B1 | | 0.2 | | 21,543,847 | 2,666,590 | 20 | 70 | 10.4 min | loose · 0.72 6/20 | listen → |
emotion__B1__T0.25__C0.20__INTERNAL | B1 | | 0.25 | | 3,895,947 | 497,796 | 20 | 86 | 16.9 min | loose · 0.40 6/20 | listen → |
emotion__B1__T0.25__C0.25__INTERNAL | B1 | | 0.25 | | 9,393,332 | 1,196,403 | 20 | 70 | 12.6 min | loose · 0.74 6/20 | listen → |
emotion__B1__T0.40__C0.25__INTERNAL | B1 | | 0.4 | | 1,711,911 | 233,805 | 20 | 82 | 15.0 min | loose · 0.66 9/20 | listen → |
emotion__B1__T0.50__C0.25__INTERNAL | B1 | | 0.5 | | 383,880 | 59,271 | 20 | 93 | 15.8 min | loose · 0.43 5/20 | listen → |
emotion__B1__T0.60__C0.25__INTERNAL | B1 | | 0.6 | | 100,844 | 16,321 | 20 | 92 | 15.0 min | loose · 0.29 5/20 | listen → |
emotion__B1__T0.70__C0.25__INTERNAL | B1 | | 0.7 | | 16,409 | 2,970 | 20 | 100 | 14.9 min | loose · 0.66 3/20 | listen → |
emotion__B1__T0.80__C0.25__INTERNAL | B1 | | 0.8 | | 2,257 | 396 | 20 | 100 | 15.7 min | loose · 0.23 3/20 | listen → |
emotion_twosided__AB2__T0.20__C0.25__INTERNAL | AB2 | | 0.2 | | 1,224,717 | 324,658 | 20 | 73 | 14.7 min | tight · 0.82 20/20 | listen → |
emotion_twosided__AB2__T0.25__C0.25__INTERNAL | AB2 | | 0.25 | | 324,658 | 324,658 | 20 | 68 | 12.5 min | tight · 0.89 20/20 | listen → |
merged_emo_vn__T0.20__C0.25__INTERNAL | B1 UNION VN1 | | 0.2 | | — | 5,354,590 | 20 | 71 | 16.9 min | loose · 0.72 8/20 | listen → |
proxy_spearman__PXR__T0.20__C0.25__INTERNAL | PXR | | 0.2 | | 1,222,387 | 1,222,387 | 20 | 64 | 10.3 min | tight · 0.85 8/20 | listen → |
proxy_spearman__PXR__T0.25__C0.25__INTERNAL | PXR | | 0.25 | | 338,690 | 338,690 | 20 | 69 | 12.8 min | tight · 0.82 19/20 | listen → |
proxy_spearman__PXR__T0.40__C0.25__INTERNAL | PXR | | 0.4 | | 12,053 | 12,053 | 20 | 84 | 14.6 min | tight · 0.81 18/20 | listen → |
proxy_spearman__PXR__T0.50__C0.25__INTERNAL | PXR | | 0.5 | | 1,016 | 1,016 | 20 | 94 | 13.2 min | tight · 0.84 20/20 | listen → |
proxy_spearman__PXR__T0.60__C0.25__INTERNAL | PXR | | 0.6 | | 70 | 70 | 20 | 97 | 13.6 min | loose · 0.78 20/20 | listen → |
proxy_spearman__PXR__T0.70__C0.25__INTERNAL (whole tier) | PXR | | 0.7 | | 2 | 2 | 2 | 10 | 1.4 min | loose · 0.44 2/2 | listen → |
proxy_spearman__PXR__T0.80__C0.25__INTERNAL — EMPTY | PXR | | 0.8 | | — | — | 0 | 0 | 0.0 min | — | listen → |
proxy_taillift__PXR__T0.20__C0.25__INTERNAL | PXR | | 0.2 | | 1,233,330 | 1,233,330 | 20 | 58 | 11.0 min | tight · 0.83 7/20 | listen → |
proxy_taillift__PXR__T0.25__C0.25__INTERNAL | PXR | | 0.25 | | 346,174 | 346,174 | 20 | 74 | 13.4 min | loose · 0.80 18/20 | listen → |
proxy_taillift__PXR__T0.40__C0.25__INTERNAL | PXR | | 0.4 | | 12,397 | 12,397 | 20 | 81 | 11.4 min | tight · 0.84 19/20 | listen → |
proxy_taillift__PXR__T0.50__C0.25__INTERNAL | PXR | | 0.5 | | 1,035 | 1,035 | 20 | 90 | 14.0 min | loose · 0.77 19/20 | listen → |
proxy_taillift__PXR__T0.60__C0.25__INTERNAL | PXR | | 0.6 | | 71 | 71 | 20 | 96 | 14.0 min | loose · 0.78 20/20 | listen → |
proxy_taillift__PXR__T0.70__C0.25__INTERNAL (whole tier) | PXR | | 0.7 | | 2 | 2 | 2 | 10 | 1.4 min | loose · 0.44 2/2 | listen → |
proxy_taillift__PXR__T0.80__C0.25__INTERNAL — EMPTY | PXR | | 0.8 | | — | — | 0 | 0 | 0.0 min | — | listen → |
voicenet__VN1__T0.20__C0.25__INTERNAL | VN1 | | 0.2 | | 430,993,195 | 2,688,000 | 20 | 73 | 11.4 min | loose · 0.70 6/20 | listen → |
voicenet__VN1__T0.50__C0.25__INTERNAL | VN1 | | 0.5 | | 17,672,060 | 116,099 | 20 | 92 | 17.4 min | loose · 0.24 12/20 | listen → |
voicenet__VN1__T0.60__C0.25__INTERNAL | VN1 | | 0.6 | | 3,019,723 | 20,076 | 20 | 93 | 17.4 min | loose · 0.33 8/20 | listen → |
voicenet__VN1__T0.70__C0.25__INTERNAL | VN1 | | 0.7 | | 299,136 | 2,094 | 20 | 100 | 20.1 min | loose · 0.37 11/20 | listen → |
voicenet__VN1__T0.80__C0.25__INTERNAL | VN1 | | 0.8 | | 20,740 | 136 | 20 | 100 | 16.8 min | loose · 0.31 13/20 | listen → |
Rule × chain length
rule x chain length
| tier | rule | k | T | corpus | population | sampled from | shown | clips | audio | same speaker? | |
|---|
k-AB2-k3 | AB2 | 3 | | | — | 162,088 | 20 | 60 | 10.6 min | tight · 0.85 20/20 | listen → |
k-AB2-k4 | AB2 | 4 | | | — | 105,692 | 20 | 80 | 15.2 min | tight · 0.85 20/20 | listen → |
k-AB2-k5 | AB2 | 5 | | | — | 56,878 | 20 | 100 | 17.9 min | loose · 0.77 20/20 | listen → |
k-B1-k2 | B1 | 2 | | | — | 664,812 | 20 | 40 | 9.3 min | tight · 0.86 9/20 | listen → |
k-B1-k3 | B1 | 3 | | | — | 1,327,006 | 20 | 60 | 11.7 min | loose · 0.76 8/20 | listen → |
k-B1-k4 | B1 | 4 | | | — | 1,513,892 | 20 | 80 | 14.9 min | loose · 0.30 7/20 | listen → |
k-B1-k5 | B1 | 5 | | | — | 1,524,064 | 20 | 100 | 17.1 min | loose · 0.56 5/20 | listen → |
k-PXR-k3 | PXR | 3 | | | — | 169,361 | 20 | 60 | 11.8 min | tight · 0.91 19/20 | listen → |
k-PXR-k4 | PXR | 4 | | | — | 110,162 | 20 | 80 | 13.9 min | tight · 0.85 20/20 | listen → |
k-PXR-k5 | PXR | 5 | | | — | 59,167 | 20 | 100 | 21.3 min | loose · 0.79 20/20 | listen → |
k-VN1-k2 | VN1 | 2 | | | — | 672,000 | 20 | 40 | 8.6 min | tight · 0.89 7/20 | listen → |
k-VN1-k3 | VN1 | 3 | | | — | 672,000 | 20 | 60 | 10.1 min | loose · 0.53 5/20 | listen → |
k-VN1-k4 | VN1 | 4 | | | — | 1,333,380 | 20 | 80 | 15.7 min | loose · 0.55 6/20 | listen → |
k-VN1-k5 | VN1 | 5 | | | — | 1,334,387 | 20 | 100 | 16.5 min | loose · 0.14 3/20 | listen → |
One corpus at a time
one corpus in isolation
| tier | rule | k | T | corpus | population | sampled from | shown | clips | audio | same speaker? | |
|---|
c-emolia-AB2 | AB2 | | | emolia | — | 231,110 | 20 | 73 | 10.5 min | tight · 0.84 20/20 | listen → |
c-emolia-B1 | B1 | | | emolia | — | 2,982,304 | 20 | 82 | 11.4 min | not measured (9 timbre only) | listen → |
c-emolia-PXR | PXR | | | emolia | — | 240,455 | 20 | 71 | 12.0 min | tight · 0.86 20/20 | listen → |
c-emolia-VN1 | VN1 | | | emolia | — | 2,304,000 | 20 | 81 | 12.9 min | not measured (11 timbre only) | listen → |
c-eurospeech-AB2 | AB2 | | | eurospeech | — | 24,724 | 20 | 81 | 19.7 min | loose · 0.46 20/20 | listen → |
c-eurospeech-B1 | B1 | | | eurospeech | — | 350,044 | 20 | 77 | 19.4 min | not measured | listen → |
c-eurospeech-PXR | PXR | | | eurospeech | — | 26,258 | 20 | 76 | 20.0 min | loose · 0.29 18/20 | listen → |
c-eurospeech-VN1 | VN1 | | | eurospeech | — | 288,000 | 20 | 75 | 18.8 min | not measured | listen → |
c-mls-AB2 | AB2 | | | mls | — | 931 | 20 | 77 | 19.4 min | tight · 0.93 20/20 | listen → |
c-mls-B1 | B1 | | | mls | — | 78,932 | 20 | 64 | 16.7 min | tight · 0.94 1/20 | listen → |
c-mls-PXR | PXR | | | mls | — | 981 | 20 | 77 | 19.3 min | tight · 0.94 18/20 | listen → |
c-mls-VN1 | VN1 | | | mls | — | 72,000 | 20 | 78 | 20.2 min | not measured | listen → |
c-podcast-AB2 | AB2 | | | podcast | — | 63,806 | 20 | 70 | 17.9 min | loose · 0.74 20/20 | listen → |
c-podcast-B1 | B1 | | | podcast | — | 1,409,804 | 20 | 73 | 16.9 min | loose · 0.59 20/20 | listen → |
c-podcast-PXR | PXR | | | podcast | — | 66,556 | 20 | 69 | 16.1 min | loose · 0.32 20/20 | listen → |
c-podcast-VN1 | VN1 | | | podcast | — | 1,152,000 | 20 | 76 | 14.2 min | loose · 0.41 20/20 | listen → |
c-snippets-AB2 | AB2 | | | snippets | — | 3,890 | 20 | 71 | 8.8 min | loose · 0.14 20/20 | listen → |
c-snippets-B1 | B1 | | | snippets | — | 177,245 | 20 | 81 | 10.9 min | not measured | listen → |
c-snippets-PXR | PXR | | | snippets | — | 4,248 | 20 | 68 | 7.5 min | loose · 0.14 20/20 | listen → |
c-snippets-VN1 | VN1 | | | snippets | — | 144,000 | 20 | 79 | 8.2 min | not measured | listen → |
c-evasnippets-AB2 | AB2 | | | evasnippets | — | 197 | 20 | 84 | 27.9 min | loose · 0.45 20/20 | listen → |
c-evasnippets-B1 | B1 | | | evasnippets | — | 31,445 | 20 | 79 | 27.4 min | not measured | listen → |
c-evasnippets-PXR | PXR | | | evasnippets | — | 192 | 20 | 79 | 27.6 min | tight · 0.91 20/20 | listen → |
c-evasnippets-VN1 | VN1 | | | evasnippets | — | 51,767 | 20 | 68 | 19.6 min | not measured | listen → |
Speaker-cleaned set
speaker-cleaned two-sided set (WavLM -id >= 0.80, consecutive AND anchored)
Voice profiles (vprof_vc)
voice-profile grid (vprof_vc): one cloned voice, chains CONSTRUCTED not discovered
Low-resource cells
the scarcest cells -- the ones the owner said matter most
Rescue rules — kept separate on purpose
These come from
looser rules and are
not strict-rule trajectories. They exist so emotions the strict rule cannot reach
can still be listened to and judged. Every tier id starts with
sad- and every
example is labelled with the rule that produced it.
What was changed, and why →| tier | emotion | rule | k | rule yield | strict rule | shown | |
|---|
sad-Sadness-S3-k2 | Sadness | S3 | 2 | 7,826 | 0 | 20 | listen → |
sad-Sadness-S3-k3 | Sadness | S3 | 3 | 5,806 | 0 | 20 | listen → |
sad-Sadness-S1-k3 | Sadness | S1 | 3 | 3,162 | 0 | 20 | listen → |
sad-Sadness-S2-k2 | Sadness | S2 | 2 | 7,924 | 0 | 20 | listen → |
sad-Sadness-S4-k2 | Sadness | S4 | 2 | 19,060 | 0 | 20 | listen → |
sad-Awe-S3-k2 | Awe | S3 | 2 | 5,762 | 0 | 20 | listen → |
sad-Distress-S3-k2 | Distress | S3 | 2 | 7,875 | 0 | 20 | listen → |
sad-Disappointment-S3-k2 | Disappointment | S3 | 2 | 13,495 | 0 | 20 | listen → |
sad-Helplessness-BASE-k3 | Helplessness | BASE | 3 | 19,693 | 19,693 | 20 | listen → |
How to read the speaker numbers
Every tier table above carries a same speaker?
badge so you can tell at a glance what a tier's identity claim is worth without opening
it: tight · 0.9x means the measured identity cosine is
at or above the 0.80 threshold on median, loose · 0.6x
means it was measured and falls below it, and not measured
means no identity embedding covers those clips. The fraction beside the badge is how many
of the tier's chains carry a measurement at all.
min_cos_consec is the smallest WavLM
-id cosine between consecutive clips; min_cos_anchor the
smallest against the first clip. They are the check on whether a chain is really one
speaker. The speaker-cleaned tiers carry them from the mined parquet; elsewhere they are
recomputed here from the per-clip embedding store where it has coverage, and shown as
— where it does not. podcast and vprof_vc have no embedding store, so their
chains show no cosine rather than a guessed one.
What the speaker numbers say once you have
them. Most tiers here are candidate sets: the mining applied no speaker rule
to them at all. Only the sc- family was speaker-cleaned. The difference is
stark, and it is the single most useful thing to know before trusting a grid on
identity.
| family | chains with a score | median anchor cosine | at or above 0.80 |
|---|
| Manifest tiers — the subsets he will train on | 313 | 0.765 | 42 % |
| Rule × chain length | 169 | 0.821 | 54 % |
| One corpus at a time | 277 | 0.752 | 43 % |
| Speaker-cleaned set | 60 | 0.876 | 100 % |
| Low-resource cells | 45 | 0.690 | 36 % |
podcast in particular. Its embedding
store covers every sampled clip, and the picture is not flattering: median anchor cosine
0.56, with only
26 % of chains at or above
the 0.80 identity threshold. That matches what is known independently — only about
54.6 % of consecutive clips inside one nominal podcast “speaker” track
are actually the same person. Podcast chains outside the speaker-cleaned family should be
treated as candidates, not as verified single-speaker trajectories.
| corpus | clips wanted | with an -id embedding |
|---|
| emolia | 2,920 | 1,360 |
| eurospeech | 548 | 254 |
| evasnippets | 221 | 116 |
| mls | 367 | 137 |
| podcast | 1,209 | 1,209 |
| snippets | 387 | 167 |
| vprof_vc | 301 | 0 |
Provenance
Samples are drawn from
emolia, eurospeech, mls, podcast,
snippets, evasnippets and the vprof_vc voice
profiles. The annotations are CC-BY-4.0. This is a listening demo — a few
trajectories per tier — not a corpus release.