What the pose data is

Every other number here is computed on a cleaned array. This is what the cleaning does, and which method is frozen as the default.

raw · viterbi · disposition
carried forward
frozen 2026-09-14
median 0.50 s
reported, not carried
best pooled, harmful on fast frames
wiener · butterworth
cannot be benchmarked
read off disk, never reimplemented

The incumbent Wiener filter is retired as the default. It is not rejected on evidence — it cannot be applied to a corrupted array, so its recovery performance is unknown and this page does not claim otherwise.

123456789101112131415% of a corruption’s error removed, by its length in frames
A 15-frame median is unchanged unless more than 7 of its 15 samples are corrupted. So the rolling median removes almost all of a 7-frame error and almost none of an 8-frame one. This is arithmetic, not tuning, and it explains every other number measured about this arm: teleports are 1–3 frames and it repairs 95.4% of them. Parked landmarks have a median length of 2 frames, so most fall under the cliff too and it repairs 58% of those — but the ones that survive are the long ones, which is the error class nothing here reaches.

Why the median is not carried

armslowestfastest
median 0.50 s−0.0026−0.0013−0.0003+0.0005+0.0053
Anipose Viterbi−0.0016−0.0018−0.0018−0.0023−0.0030
disposition−0.0012−0.0013−0.0015−0.0017−0.0022

Error against injected corruption with known truth, by speed of the animal. Negative is better; positive means the arm left the data further from the truth than it found it. The median wins on still frames and does damage on moving ones; Viterbi is steady and best where the animal moves. On a fear-conditioning corpus the moving frames are the signal, which is why the median is reported and not carried. Its park lengths were drawn from a biased distribution and correcting that did not change this.

What each method does

methodmechanismconsequence
median 0.50 s15-frame centred rolling median, x and y filtered separatelyUnchanged unless 8 of 15 samples are corrupted — a cliff at 7 frames
Wiener (incumbent)Frequency shrinkage, gain = SNR/(1+SNR), no cutoffAttenuates every frequency by how much of it is thought to be signal
ButterworthOrder-4 zero-phase low-pass at 4.833 HzHard cutoff; everything faster is removed, signal or not
Savitzky–GolaySliding quadratic fit, 11-frame windowPreserves peak height, but one outlier pulls the whole window
refineDLC outlierGate frame-to-frame jumps, drop, interpolateThe only arm that leaves ordinary frames exactly as measured
Anipose ViterbiPath selection over 4 candidates per frame, motion priorAccepts or rejects a jump; never averages, so 99.7% passes untouched
disposition (ours)Find the suspect keypoint, correct short runs, abstain on long onesReplaces one keypoint with a prediction; refuses where the run is sustained

All of them receive the same input: DeepLabCut’s 7 keypoints at 30 fps, after gaps of 3 frames or fewer are interpolated (1.34% of keypoint-frames) and longer ones abandoned (3.39%).

What is wrong underneath

3.80%
frames with a continuity spike
a keypoint that disagrees with both neighbours
1.12%
frames the bone check flags
at the same ε, on the same frames
8.2%
of spiking frames it also flags
two nearly disjoint nets

A bone check can only see an error that changes a length. A keypoint sliding along a bone, or a whole skull triangle drifting, changes none — so 232,792 report frames carry a visible discontinuity no bone check at any threshold could have found. Anipose’s Viterbi filter removes 67% of them while moving 0.19 px on average, which is 4.3× the efficiency of either smoother.

The Wiener filter is where the published data moves: 56% of keypoint-frames by more than half a pixel, removing 41% of the median instantaneous speed and 24% of the median turning. That had never been measured, and every number on the other six tabs is computed on its output. The rolling median goes further and deletes 86–90% of median movement; Anipose’s Viterbi filter keeps 98–100% of it.

What the flagged frames actually look like

0
of 100 frames occluded
not one animal genuinely hidden
23
visible but wrong
the landmark is there; the network missed it
32% vs 14%
wrong, flagged vs not
the check works, and is noisy

Twenty-five frames from each of four cells — worst and passing recordings, flagged and unflagged — rendered with no flag marker and scored without knowing which was which. Zero occluded overturns the reading that tracking fails because the animal is hidden. Six frames show the same specific fault: a keypoint locked onto a bright object above the arena, across three boxes and four days. That is not occlusion, and it is the kind of error better detections would fix.

Should we retrain? The open question

350×
more concentrated than chance
half the failure in a tenth of recordings
3.7×
higher at the arena wall
monotone across ten deciles
0
labelled frames on disk
no model, no project, no environment

Whether to train a five-network ensemble is open, and what it would buy is a real uncertainty estimate, which nothing here has. Tracking failure is concentrated, rises toward the wall, and is flat across the three apparatus units; I read that as a footage limit, and scoring 100 frames by eye found zero occluded and overturned it. The wall effect is real but it is high-contrast bars distracting the detector rather than the animal being hidden — a network problem, which reopens this decision and raises the expected return.

Before and after

128 clips across eight sets, every pane the same frames and the same crop, with the same red outline on flagged frames — the only difference between panes is the skeleton. The first three compare the pose before and after cleaning, on the array every downstream number is computed from. The fourth is what the best filter still gets wrong: a landmark parked off the animal and held there, which no temporal filter can reach because it is temporally smooth. The last is the answer to that — one suspect keypoint carried back from the nearest clean frame, on the 0.50% of frames whose run is short enough to trust, with the runs the corrector refused to touch outlined beside them, Anipose's Viterbi de-glitcher on the 0.31% of frames it reassigned, and the two candidate arms against the floor on the frames where they disagree.

No method repairs a landmark parked on the wrong body part — 10.5% is the ceiling, and two of the three arms manage none. That is what a detector produces when it locks onto the wrong thing, and no temporal or geometric filter raises it. A pose-conditional model would be neither, and is a candidate counterexample — not used here, because the one available is also a contestant in the model comparison.