Can the crop be lined up without keypoints?

On a still animal, yes. On a moving one, not well enough.

Stage 3. SAM prompted by its own last mask (red) starts on the mouse, then moves to the bright arena wall at predicted IoU 0.94–0.96 and stays there. Keypoints in green.
Stage 3. SAM prompted by its own last mask (red) starts on the mouse, then moves to the bright arena wall at predicted IoU 0.94–0.96 and stays there. Keypoints in green.
Stage 4, same recording. Prompted from the padded keypoint box (yellow) at every keyframe, the mask stays on the mouse. The blue cross is the head-disc centre.
Stage 4, same recording. Prompted from the padded keypoint box (yellow) at every keyframe, the mask stays on the mouse. The blue cross is the head-disc centre.
[0.999, 1.000]
stage 4 energy under planted jitter
2 px of noise in every keypoint input; 1 means untouched
[8.931, 15.137]
the incumbent under the same jitter
its still-window head energy, multiplied
alignment0 on animal1 duplicates2 jitter3 still frames4 moving plantarm
1 · median-background mask—PASSNOT_A_RESULTNOT_A_RESULTFAIL [1134.851, 2474.794]NOT_A_RESULT
1 · phase correlation on that mask—PASSNOT_A_RESULTNOT_A_RESULTFAIL [619.103, 1559.352]NOT_A_RESULT
2 · keypoint-masked background—PASSNOT_A_RESULTNOT_A_RESULTNOT_A_RESULTNOT_A_RESULT
2 · ECC on that mask—PASSNOT_A_RESULTNOT_A_RESULTNOT_A_RESULTNOT_A_RESULT
3 · SAM, moment warpFAIL [0.585, 0.750]FAILFAIL [0.039, 0.148]NOT_A_RESULTFAIL [14.994, 19.947]NOT_A_RESULT
3 · SAM, ECCFAIL [0.585, 0.750]PASSFAIL [0.208, 0.299]NOT_A_RESULTFAIL [13.551, 19.453]NOT_A_RESULT
4 · SAM from the located animal, ECCPASS [0.923, 0.950]PASSPASS [0.999, 1.000]PASS [1.016, 1.063]FAIL [15.343, 18.736]FAIL
5 · ECC on the animal onlyPASS [0.923, 0.950]FAILPASS [1.000, 1.002]FAIL [1.854, 2.779]FAIL [1.266, 1.344]FAIL
6 · same, from the better of two startsPASS [0.923, 0.950]PASSPASS [1.000, 1.002]FAIL [1.844, 2.773]FAIL [1.266, 1.344]FAIL
7 · stage 6, scored on the animal's own pixelsPASS [0.923, 0.950]PASSPASS [0.999, 1.000]PASS [1.003, 1.060]FAIL [1.934, 2.346]FAIL
incumbent, animal pixelsPASS [0.905, 0.933]FAILNOT_A_RESULT [7.709, 12.266]FAIL [1.373, 2.171]FAIL [6.435, 10.323]FAIL
real fast frames, 10-frame chains6 · body motion vs keypoints
4’s unmasked ECC, re-runFAIL [0.520, 0.611]
5 · ECC on the animal onlyPASS [1.073, 1.095]
6 · from the better of two startsPASS [1.056, 1.078]

Ten full recordings, both detectors on. Red fill is SAM's mask as the pipeline used it; amber is the DLC skeleton. The mask's outline shows whether the two agree (green), disagree (amber, magenta, orange) or SAM was refused, and a red dot marks any keypoint outside the mask.

Watch it

Box 3 · day 3 · context A · animal 975 — SAM drifted onto the wall here in stage 3
210 s · refused 0.8% of frames · head disc on the mask 78% of keyframes
Box 2 · day 6 · context B · animal 9025 — the recording stage 7 refused most
180 s · refused 94.8% of frames · head disc on the mask 92% of keyframes
Box 2 · day 3 · context A · animal 104 — the smoke-test recording
210 s · refused 0.1% of frames · head disc on the mask 92% of keyframes
Box 3 · day 7 · context B · animal 2941
180 s · refused 2.1% of frames · head disc on the mask 64% of keyframes
Box 2 · day 3 · context A · animal 9037
210 s · refused 0.0% of frames · head disc on the mask 99% of keyframes
Box 1 · day 4 · context B · animal 9005
180 s · refused 12.6% of frames · head disc on the mask 90% of keyframes
Box 3 · day 6 · context A · animal 716
210 s · refused 0.2% of frames · head disc on the mask 97% of keyframes
Box 2 · day 3 · context B · animal 134
180 s · refused 0.1% of frames · head disc on the mask 96% of keyframes
Box 3 · day 4 · context A · animal HM03
210 s · refused 0.7% of frames · head disc on the mask 81% of keyframes
Box 1 · day 4 · context B · animal 572
180 s · refused 1.1% of frames · head disc on the mask 99% of keyframes

Two detectors, two checks

keypointoutside SAM's mask, all keyframeson keyframes where they disagree
right ear16.0%64.0%
nose13.8%61.0%
left ear13.0%54.3%
tail base5.8%19.9%
right hip3.3%17.6%
left hip2.4%13.0%
body centre0.5%3.4%

Mask and keypoints agree on 81% of frames. When they do not, the usual cause is SAM cutting off the head: the nose and ears fall outside the mask on about one frame in seven, the body centre almost never. The head is what a grooming measure needs, so the next step is to prompt SAM with the nose and ears as points.

ears outside the maskcontext Acontext B
box prompt (current)10.0%20.1%
box + keypoints as points5.4%9.0%, mask balloons onto the wall
guarded: points only near the mask, area capped5.9%15.8%

Giving SAM the nose, centre and tail as points brings the head back, measured on the ears it was never given. But where DLC puts a keypoint on the bright wall, SAM follows it there. Guarded, each detector checks the other first: the head returns in context A with no ballooning, much less so in B.

Head coverage depends on the arena. Even the plain box mask clips the head twice as often in context B. Any A-versus-B comparison of a head-region measure has to control for how much head each frame's mask contains.

Top two rows: point-prompted masks that grew most, all in context B, pulled onto the bright wall by a keypoint that sits there (yellow dots are the prompts, blue the ears). Bottom row: typical frames.
Top two rows: point-prompted masks that grew most, all in context B, pulled onto the bright wall by a keypoint that sits there (yellow dots are the prompts, blue the ears). Bottom row: typical frames.
Examples of each verdict. The “DLC suspect” frames (top row) are really the mask missing the head: the automatic blame split was too crude and is not trusted.
Examples of each verdict. The “DLC suspect” frames (top row) are really the mask missing the head: the automatic blame split was too crude and is not trusted.
Gate 4, one planted pair. Middle row: the 0.2 body-length head disc, zoomed, with its mean |difference|. Stage 4 removes about a third of the motion; the true transform removes almost all of it. The full-frame panels saturate and can hide this.
Gate 4, one planted pair. Middle row: the 0.2 body-length head disc, zoomed, with its mean |difference|. Stage 4 removes about a third of the motion; the true transform removes almost all of it. The full-frame panels saturate and can hide this.
The most-refused recording in stage 4. The mask takes in a white object beside the head, and sometimes the tail.
The most-refused recording in stage 4. The mask takes in a white object beside the head, and sometimes the tail.

Stage 4 passes every still-frame gate. It is exact on byte-identical frames, within a few percent of perfect on truly still ones, and planted keypoint jitter does not reach it. The incumbent reports motion on every duplicate frame.

Stage 5 follows moving animals. Restricting the alignment to the mouse's own mask brings the moving plant within about 30% of a perfect alignment, where stage 4 was 17 times off. On real video its body motion agrees with the keypoints (column 6). But it fails the still-frame gates stage 4 passed, so each arm solves the case the other fails.

Stage 7 ends the line. Scored on the mouse's own pixels, the stage-6 alignment passes every still-frame test, within 3% of perfect, and follows the body on real video. It fails only the moving-head test: about twice a perfect alignment, where perfect keypoints manage 1.4. A stop rule set before the run ends the search there.

Stage 6 starts the masked fit from no motion and keeps whichever start fits better. That made it exact on identical frames. It still fails the still-frame gate, and the likely reason is the gate: its disc reaches the bar floor, which the fit's sub-pixel animal motion moves. The next test scores the animal's own pixels.

Stage 4 fails on moving animals. Inside SAM's box the static bar floor pulls the alignment toward zero while the mouse pulls it toward its motion, and the fit settles in between at every speed. An earlier write-up said it locked onto the floor; the pictures above corrected that.

Column 6 is the keypoints used as a referee, on the fastest quarter of real frames: they are too noisy for fine head motion but sound for where the body went. Chained over ten frames, the masked fits agree with them to within about 3–8% either way; the unmasked fit recovers at most about half of the body's motion.

Some FAILs above carry no information. Stage 1's gate 4 could not be passed by any alignment, even a perfect one (DEVIATIONS D23), and stage 2's gates 3 and 4 were ill-posed (D24). The jitter gate of stages 1–3 rested on a doubtful premise (D25), which is why stage 4 plants the jitter instead.

Each defect is recorded rather than fixed after the fact. Every cell opens its record. Nothing on this page is read as grooming.