XMem under stress
Probing a video object segmentation model with synthetic benchmarks
XMem is a video object segmentation model built around an explicit memory system. Benchmark numbers tell you it works; they don’t tell you when it stops working.
So I built synthetic sequences with controlled failure modes — occlusion, appearance change, fast motion — and watched the memory behave under each one. Synthetic data means the ground truth is exact and the difficulty is a dial, not a property of whatever footage happened to be in the dataset.
An occlusion sequence, frame by frame.