What the pose data is
Every other number here is computed on a cleaned array. This is what the cleaning does, and which method is frozen as the default.
The incumbent Wiener filter is retired as the default. It is not rejected on evidence — it cannot be applied to a corrupted array, so its recovery performance is unknown and this page does not claim otherwise.
Why the median is not carried
| arm | slowest | fastest | |||
|---|---|---|---|---|---|
| median 0.50 s | −0.0026 | −0.0013 | −0.0003 | +0.0005 | +0.0053 |
| Anipose Viterbi | −0.0016 | −0.0018 | −0.0018 | −0.0023 | −0.0030 |
| disposition | −0.0012 | −0.0013 | −0.0015 | −0.0017 | −0.0022 |
Error against injected corruption with known truth, by speed of the animal. Negative is better; positive means the arm left the data further from the truth than it found it. The median wins on still frames and does damage on moving ones; Viterbi is steady and best where the animal moves. On a fear-conditioning corpus the moving frames are the signal, which is why the median is reported and not carried. Its park lengths were drawn from a biased distribution and correcting that did not change this.
What each method does
| method | mechanism | consequence |
|---|---|---|
| median 0.50 s | 15-frame centred rolling median, x and y filtered separately | Unchanged unless 8 of 15 samples are corrupted — a cliff at 7 frames |
| Wiener (incumbent) | Frequency shrinkage, gain = SNR/(1+SNR), no cutoff | Attenuates every frequency by how much of it is thought to be signal |
| Butterworth | Order-4 zero-phase low-pass at 4.833 Hz | Hard cutoff; everything faster is removed, signal or not |
| Savitzky–Golay | Sliding quadratic fit, 11-frame window | Preserves peak height, but one outlier pulls the whole window |
| refineDLC outlier | Gate frame-to-frame jumps, drop, interpolate | The only arm that leaves ordinary frames exactly as measured |
| Anipose Viterbi | Path selection over 4 candidates per frame, motion prior | Accepts or rejects a jump; never averages, so 99.7% passes untouched |
| disposition (ours) | Find the suspect keypoint, correct short runs, abstain on long ones | Replaces one keypoint with a prediction; refuses where the run is sustained |
All of them receive the same input: DeepLabCut’s 7 keypoints at 30 fps, after gaps of 3 frames or fewer are interpolated (1.34% of keypoint-frames) and longer ones abandoned (3.39%).
What is wrong underneath
A bone check can only see an error that changes a length. A keypoint sliding along a bone, or a whole skull triangle drifting, changes none — so 232,792 report frames carry a visible discontinuity no bone check at any threshold could have found. Anipose’s Viterbi filter removes 67% of them while moving 0.19 px on average, which is 4.3× the efficiency of either smoother.
The Wiener filter is where the published data moves: 56% of keypoint-frames by more than half a pixel, removing 41% of the median instantaneous speed and 24% of the median turning. That had never been measured, and every number on the other six tabs is computed on its output. The rolling median goes further and deletes 86–90% of median movement; Anipose’s Viterbi filter keeps 98–100% of it.
What the flagged frames actually look like
Twenty-five frames from each of four cells — worst and passing recordings, flagged and unflagged — rendered with no flag marker and scored without knowing which was which. Zero occluded overturns the reading that tracking fails because the animal is hidden. Six frames show the same specific fault: a keypoint locked onto a bright object above the arena, across three boxes and four days. That is not occlusion, and it is the kind of error better detections would fix.
Should we retrain? The open question
Whether to train a five-network ensemble is open, and what it would buy is a real uncertainty estimate, which nothing here has. Tracking failure is concentrated, rises toward the wall, and is flat across the three apparatus units; I read that as a footage limit, and scoring 100 frames by eye found zero occluded and overturned it. The wall effect is real but it is high-contrast bars distracting the detector rather than the animal being hidden — a network problem, which reopens this decision and raises the expected return.
Before and after
128 clips across eight sets, every pane the same frames and the same crop, with the same red outline on flagged frames — the only difference between panes is the skeleton. The first three compare the pose before and after cleaning, on the array every downstream number is computed from. The fourth is what the best filter still gets wrong: a landmark parked off the animal and held there, which no temporal filter can reach because it is temporally smooth. The last is the answer to that — one suspect keypoint carried back from the nearest clean frame, on the 0.50% of frames whose run is short enough to trust, with the runs the corrector refused to touch outlined beside them, Anipose's Viterbi de-glitcher on the 0.31% of frames it reassigned, and the two candidate arms against the floor on the frames where they disagree.
No method repairs a landmark parked on the wrong body part — 10.5% is the ceiling, and two of the three arms manage none. That is what a detector produces when it locks onto the wrong thing, and no temporal or geometric filter raises it. A pose-conditional model would be neither, and is a candidate counterexample — not used here, because the one available is also a contestant in the model comparison.