Listening grids — mined emotional and VoiceNet trajectories

92 tiers · 1,764 trajectories · 6,477 clips · 20.5 h — each chain rendered as one concatenation you can play straight through, with per-clip playback and transcripts.

New — a voice-conversion-corrected version of the emotion tiers. Every chain on this site is separate recordings stitched together, so the voice can be heard shifting between segments. A conversion pass has now re-voiced 3,438 segments across 1,304 emotion chains onto their own chain's first segment, with ChatterboxVC + SIDON, and published them as a parallel set — the pages here are unchanged and every corrected card carries both renders side by side. It is not a clean win and the measurements say so. The corrected tiers →  ·  what was done and what it measured →
And a third build — the strict, no-voice-conversion path. For chains that already clear a 0.80 speaker-identity cut, conversion costs similarity rather than adding it. This set is built by filtering instead: similarity ≥ 0.80 on both conditions, 150 ms equal-power crossfades, chain-level normalisation, no conversion. It carries the counts of how many trajectories survive — 3,378,525 mined chains at T≥0.20 and 1,399,128 at T≥0.25, plus 949,212 voice-profile chains. The strict subset →
What a trajectory is. An ordered list of speech clips from one speaker/track in which a scored dimension moves monotonically. The rule fixes what must move and by how much: a threshold T on the total move end-to-end, and a cap C on each consecutive step, so the chain ramps rather than jumping. k is the number of clips.
What you are being asked to judge. Play the concatenation and ask whether the thing the rule claims is moving actually moves, and whether it moves smoothly rather than lurching at one seam. Then compare tiers: does a higher T sound like a bigger change? Does a longer k sound smoother or just longer? Do the proxy (PXR) chains hold up as well as the direct ones?
New: why Sadness produced nothing, and four rules that fix it. Sadness, Awe, Distress and Disappointment return zero chains under the strict rule. The reason is measurable and it is not what it looks like. There is a full plain-language walkthrough with the diagnosis, four alternative rules, their yields and their costs, plus 180 listenable rescued examples. Read it →
Loudness has been normalised, deliberately. The MOSS encoder normalises every clip to −20 dBFS independently and 59.5 % of clips hit its ±3 dB clamp, so a raw concatenation carries level steps of up to 8 dB that belong to the pipeline, not the trajectory. Each clip here is scaled to −20 dBFS RMS exactly before joining, and the peak guard is applied once to the finished chain so the level relationship between clips is preserved. The gain applied is printed under every clip, and the first 5 items of each tier (444 in total) also carry an un-normalised render so the raw seam can still be heard. A 150 ms silence marks each boundary.

Manifest tiers — the subsets he will train on

the tier owner's manifest tiers -- the exact subsets used for training

tierrulekTcorpuspopulationsampled fromshownclipsaudiosame speaker?
emotion__B1__T0.20__C0.25__INTERNALB10.221,543,8472,666,590207010.4 minloose · 0.72 6/20listen →
emotion__B1__T0.25__C0.20__INTERNALB10.253,895,947497,796208616.9 minloose · 0.40 6/20listen →
emotion__B1__T0.25__C0.25__INTERNALB10.259,393,3321,196,403207012.6 minloose · 0.74 6/20listen →
emotion__B1__T0.40__C0.25__INTERNALB10.41,711,911233,805208215.0 minloose · 0.66 9/20listen →
emotion__B1__T0.50__C0.25__INTERNALB10.5383,88059,271209315.8 minloose · 0.43 5/20listen →
emotion__B1__T0.60__C0.25__INTERNALB10.6100,84416,321209215.0 minloose · 0.29 5/20listen →
emotion__B1__T0.70__C0.25__INTERNALB10.716,4092,9702010014.9 minloose · 0.66 3/20listen →
emotion__B1__T0.80__C0.25__INTERNALB10.82,2573962010015.7 minloose · 0.23 3/20listen →
emotion_twosided__AB2__T0.20__C0.25__INTERNALAB20.21,224,717324,658207314.7 mintight · 0.82 20/20listen →
emotion_twosided__AB2__T0.25__C0.25__INTERNALAB20.25324,658324,658206812.5 mintight · 0.89 20/20listen →
merged_emo_vn__T0.20__C0.25__INTERNALB1 UNION VN10.25,354,590207116.9 minloose · 0.72 8/20listen →
proxy_spearman__PXR__T0.20__C0.25__INTERNALPXR0.21,222,3871,222,387206410.3 mintight · 0.85 8/20listen →
proxy_spearman__PXR__T0.25__C0.25__INTERNALPXR0.25338,690338,690206912.8 mintight · 0.82 19/20listen →
proxy_spearman__PXR__T0.40__C0.25__INTERNALPXR0.412,05312,053208414.6 mintight · 0.81 18/20listen →
proxy_spearman__PXR__T0.50__C0.25__INTERNALPXR0.51,0161,016209413.2 mintight · 0.84 20/20listen →
proxy_spearman__PXR__T0.60__C0.25__INTERNALPXR0.67070209713.6 minloose · 0.78 20/20listen →
proxy_spearman__PXR__T0.70__C0.25__INTERNAL (whole tier)PXR0.7222101.4 minloose · 0.44 2/2listen →
proxy_spearman__PXR__T0.80__C0.25__INTERNAL — EMPTYPXR0.8000.0 minlisten →
proxy_taillift__PXR__T0.20__C0.25__INTERNALPXR0.21,233,3301,233,330205811.0 mintight · 0.83 7/20listen →
proxy_taillift__PXR__T0.25__C0.25__INTERNALPXR0.25346,174346,174207413.4 minloose · 0.80 18/20listen →
proxy_taillift__PXR__T0.40__C0.25__INTERNALPXR0.412,39712,397208111.4 mintight · 0.84 19/20listen →
proxy_taillift__PXR__T0.50__C0.25__INTERNALPXR0.51,0351,035209014.0 minloose · 0.77 19/20listen →
proxy_taillift__PXR__T0.60__C0.25__INTERNALPXR0.67171209614.0 minloose · 0.78 20/20listen →
proxy_taillift__PXR__T0.70__C0.25__INTERNAL (whole tier)PXR0.7222101.4 minloose · 0.44 2/2listen →
proxy_taillift__PXR__T0.80__C0.25__INTERNAL — EMPTYPXR0.8000.0 minlisten →
voicenet__VN1__T0.20__C0.25__INTERNALVN10.2430,993,1952,688,000207311.4 minloose · 0.70 6/20listen →
voicenet__VN1__T0.50__C0.25__INTERNALVN10.517,672,060116,099209217.4 minloose · 0.24 12/20listen →
voicenet__VN1__T0.60__C0.25__INTERNALVN10.63,019,72320,076209317.4 minloose · 0.33 8/20listen →
voicenet__VN1__T0.70__C0.25__INTERNALVN10.7299,1362,0942010020.1 minloose · 0.37 11/20listen →
voicenet__VN1__T0.80__C0.25__INTERNALVN10.820,7401362010016.8 minloose · 0.31 13/20listen →

Rule × chain length

rule x chain length

tierrulekTcorpuspopulationsampled fromshownclipsaudiosame speaker?
k-AB2-k3AB23162,088206010.6 mintight · 0.85 20/20listen →
k-AB2-k4AB24105,692208015.2 mintight · 0.85 20/20listen →
k-AB2-k5AB2556,8782010017.9 minloose · 0.77 20/20listen →
k-B1-k2B12664,81220409.3 mintight · 0.86 9/20listen →
k-B1-k3B131,327,006206011.7 minloose · 0.76 8/20listen →
k-B1-k4B141,513,892208014.9 minloose · 0.30 7/20listen →
k-B1-k5B151,524,0642010017.1 minloose · 0.56 5/20listen →
k-PXR-k3PXR3169,361206011.8 mintight · 0.91 19/20listen →
k-PXR-k4PXR4110,162208013.9 mintight · 0.85 20/20listen →
k-PXR-k5PXR559,1672010021.3 minloose · 0.79 20/20listen →
k-VN1-k2VN12672,00020408.6 mintight · 0.89 7/20listen →
k-VN1-k3VN13672,000206010.1 minloose · 0.53 5/20listen →
k-VN1-k4VN141,333,380208015.7 minloose · 0.55 6/20listen →
k-VN1-k5VN151,334,3872010016.5 minloose · 0.14 3/20listen →

One corpus at a time

one corpus in isolation

tierrulekTcorpuspopulationsampled fromshownclipsaudiosame speaker?
c-emolia-AB2AB2emolia231,110207310.5 mintight · 0.84 20/20listen →
c-emolia-B1B1emolia2,982,304208211.4 minnot measured (9 timbre only)listen →
c-emolia-PXRPXRemolia240,455207112.0 mintight · 0.86 20/20listen →
c-emolia-VN1VN1emolia2,304,000208112.9 minnot measured (11 timbre only)listen →
c-eurospeech-AB2AB2eurospeech24,724208119.7 minloose · 0.46 20/20listen →
c-eurospeech-B1B1eurospeech350,044207719.4 minnot measuredlisten →
c-eurospeech-PXRPXReurospeech26,258207620.0 minloose · 0.29 18/20listen →
c-eurospeech-VN1VN1eurospeech288,000207518.8 minnot measuredlisten →
c-mls-AB2AB2mls931207719.4 mintight · 0.93 20/20listen →
c-mls-B1B1mls78,932206416.7 mintight · 0.94 1/20listen →
c-mls-PXRPXRmls981207719.3 mintight · 0.94 18/20listen →
c-mls-VN1VN1mls72,000207820.2 minnot measuredlisten →
c-podcast-AB2AB2podcast63,806207017.9 minloose · 0.74 20/20listen →
c-podcast-B1B1podcast1,409,804207316.9 minloose · 0.59 20/20listen →
c-podcast-PXRPXRpodcast66,556206916.1 minloose · 0.32 20/20listen →
c-podcast-VN1VN1podcast1,152,000207614.2 minloose · 0.41 20/20listen →
c-snippets-AB2AB2snippets3,89020718.8 minloose · 0.14 20/20listen →
c-snippets-B1B1snippets177,245208110.9 minnot measuredlisten →
c-snippets-PXRPXRsnippets4,24820687.5 minloose · 0.14 20/20listen →
c-snippets-VN1VN1snippets144,00020798.2 minnot measuredlisten →
c-evasnippets-AB2AB2evasnippets197208427.9 minloose · 0.45 20/20listen →
c-evasnippets-B1B1evasnippets31,445207927.4 minnot measuredlisten →
c-evasnippets-PXRPXRevasnippets192207927.6 mintight · 0.91 20/20listen →
c-evasnippets-VN1VN1evasnippets51,767206819.6 minnot measuredlisten →

Speaker-cleaned set

speaker-cleaned two-sided set (WavLM -id >= 0.80, consecutive AND anchored)

tierrulekTcorpuspopulationsampled fromshownclipsaudiosame speaker?
sc-AB2-k3AB2390,36220609.5 mintight · 0.87 20/20listen →
sc-AB2-k4AB2451,769208011.8 mintight · 0.87 20/20listen →
sc-AB2-k5AB2525,6762010018.0 mintight · 0.89 20/20listen →

Voice profiles (vprof_vc)

voice-profile grid (vprof_vc): one cloned voice, chains CONSTRUCTED not discovered

tierrulekTcorpuspopulationsampled fromshownclipsaudiosame speaker?
vp-AB2-k3AB230.2520609.1 minnot measuredlisten →
vp-AB2-k4AB240.25208011.4 minnot measuredlisten →
vp-AB2-k5AB250.252010016.0 minnot measuredlisten →
vp-B1-k3B130.2520609.7 minnot measuredlisten →
vp-PXR-k3PXR30.25206010.2 minnot measuredlisten →
vp-VN1-k3VN130.2520609.1 minnot measuredlisten →
vp-VN1-k4VN140.25208012.5 minnot measuredlisten →

Low-resource cells

the scarcest cells -- the ones the owner said matter most

tierrulekTcorpuspopulationsampled fromshownclipsaudiosame speaker?
rare-AB2-pairsAB2324,658206515.0 minloose · 0.72 20/20listen →
rare-PXR-pairsPXR338,690207514.0 minloose · 0.31 9/20listen →
rare-B1-pairsB15,029,774205815.7 minloose · 0.64 2/20listen →
rare-VN1-dimsVN14,011,767204010.0 minnot measuredlisten →
rare-langmixed9,704,889204011.2 mintight · 0.85 14/20listen →

Rescue rules — kept separate on purpose

These come from looser rules and are not strict-rule trajectories. They exist so emotions the strict rule cannot reach can still be listened to and judged. Every tier id starts with sad- and every example is labelled with the rule that produced it. What was changed, and why →
tieremotionrulekrule yieldstrict ruleshown
sad-Sadness-S3-k2SadnessS327,826020listen →
sad-Sadness-S3-k3SadnessS335,806020listen →
sad-Sadness-S1-k3SadnessS133,162020listen →
sad-Sadness-S2-k2SadnessS227,924020listen →
sad-Sadness-S4-k2SadnessS4219,060020listen →
sad-Awe-S3-k2AweS325,762020listen →
sad-Distress-S3-k2DistressS327,875020listen →
sad-Disappointment-S3-k2DisappointmentS3213,495020listen →
sad-Helplessness-BASE-k3HelplessnessBASE319,69319,69320listen →

How to read the speaker numbers

Every tier table above carries a same speaker? badge so you can tell at a glance what a tier's identity claim is worth without opening it: tight · 0.9x means the measured identity cosine is at or above the 0.80 threshold on median, loose · 0.6x means it was measured and falls below it, and not measured means no identity embedding covers those clips. The fraction beside the badge is how many of the tier's chains carry a measurement at all.
min_cos_consec is the smallest WavLM -id cosine between consecutive clips; min_cos_anchor the smallest against the first clip. They are the check on whether a chain is really one speaker. The speaker-cleaned tiers carry them from the mined parquet; elsewhere they are recomputed here from the per-clip embedding store where it has coverage, and shown as — where it does not. podcast and vprof_vc have no embedding store, so their chains show no cosine rather than a guessed one.
What the speaker numbers say once you have them. Most tiers here are candidate sets: the mining applied no speaker rule to them at all. Only the sc- family was speaker-cleaned. The difference is stark, and it is the single most useful thing to know before trusting a grid on identity.
familychains with a scoremedian anchor cosineat or above 0.80
Manifest tiers — the subsets he will train on3130.76542 %
Rule × chain length1690.82154 %
One corpus at a time2770.75243 %
Speaker-cleaned set600.876100 %
Low-resource cells450.69036 %
podcast in particular. Its embedding store covers every sampled clip, and the picture is not flattering: median anchor cosine 0.56, with only 26 % of chains at or above the 0.80 identity threshold. That matches what is known independently — only about 54.6 % of consecutive clips inside one nominal podcast “speaker” track are actually the same person. Podcast chains outside the speaker-cleaned family should be treated as candidates, not as verified single-speaker trajectories.
corpusclips wantedwith an -id embedding
emolia2,9201,360
eurospeech548254
evasnippets221116
mls367137
podcast1,2091,209
snippets387167
vprof_vc3010

Provenance

Samples are drawn from emolia, eurospeech, mls, podcast, snippets, evasnippets and the vprof_vc voice profiles. The annotations are CC-BY-4.0. This is a listening demo — a few trajectories per tier — not a corpus release.