Two labellers, three arms
One fits a model and always answers. The other declines. A third finds the pieces before naming them.
keypoint-MoSeq
MoSeq fits a model that assumes behaviour is a sequence of discrete chunks, and slices every frame into one. It always returns an answer, including when it is given noise.
nonparametric v1
The nonparametric labeller finds stretches that recur across animals and labels only frames that closely match one, leaving the rest unlabelled. It declines to label a smooth continuum entirely.
nonparametric v2
The same labeller, with each seed's threshold calibrated against a null instead of set to one arbitrary distance. Five times the coverage, at contamination below MoSeq's.

| MoSeq | nonparam v1 | nonparam v2 | |
|---|---|---|---|
| coverage | 100% by construction | 14.0% | 76.0% |
| abstention | none — no abstain bin exists | 86.01% | 24.02% |
| artifact mass | 2.8% | 0.00% | 2.16% |
| surrogate assignment | not calibrated against one | 0.00% | α = 1e-4 per seed |
| grammar | 1,491provisional | INCONCLUSIVE, 0 of 0 | 8 of 8 surviving |
| noise control | white noise survives it too | no n-gram to test | not yet run |
| human validation | none | none | none |
The 2.8% figure is the conservative half of two: a further 9 labels (18.9% of frames) are enriched for tracking flags but not in every stratum, which means their failure tracks posture — real behaviour that degrades tracking, not the tracker alone. Counting both as one number would overstate the contamination.
The v1 arm's 0.00% surrogate assignment has a 24× margin and does not show its acceptance regions are tight. Phase, AR and OU surrogates all sit about 24× further from the corpus than the corpus sits from itself, so any threshold calibrated against them lands past full coverage.
Finding the pieces without naming them first
Every arm above starts by assuming what a piece of behaviour is. This one does not: it finds where the movement stops being smooth, takes what lies between two breaks as the unit, and asks whether those units turn up in other animals.
| channel group | segments | length-matched windows | difference | do segments recur? |
|---|---|---|---|---|
| shape | +5.930%[+4.608%, +7.184%] | +0.737%[+0.469%, +1.005%] | +5.193%[+3.959%, +6.449%] | yes |
| both | +4.048%[+3.071%, +4.998%] | +0.364%[+0.146%, +0.568%] | +3.684%[+2.723%, +4.639%] | yes |
| twist | +0.117%[-0.185%, +0.427%] | -0.535%[-0.794%, -0.308%] | +0.652%[+0.336%, +0.985%] | no — interval spans zero |
The comparison is the result, not the left-hand column. A bank of any units cut from a real animal beats a surrogate, so the segment number alone says nothing about the boundaries. What says something is the same statistic recomputed with random windows drawn to the same lengths, in the same space, against the same dwell-matched nulls.
A continuum with islands
Those groups are not reproduced by surrogates that match the corpus's dwell times, so the structure is real. It covers under two per cent of the segments, which is why no vocabulary is claimed here.
The widest group is the one thing in this programme that has been validated against the experimental design rather than against its appearance. What it is, what it is not, and its clips against a blind control are on The island.