Tracking Any Point
任意点跟踪TAPAdvancedGiven any point in a video, outputting its position in every later frame, and whether it becomes occluded.
Tracking Any Point (TAP) is a vision task formally introduced by Google DeepMind in 2022 alongside the TAP-Vid benchmark: a user marks an arbitrary point on some frame of a video — on an object’s surface, on a piece of cloth, or in the background — and the model must output that point’s pixel position in every other frame, and whether it is occluded. It differs from optical flow, which only computes motion between two adjacent frames and drifts when accumulated over a long sequence, and can’t handle a point disappearing behind an occluder and reappearing; it differs from object tracking in that it tracks a point rather than a whole object’s bounding box. Representative models include TAPIR and CoTracker. In robotics, point trajectories serve as an embodiment-agnostic intermediate representation: the motion of points on an object can be extracted from human video and used to guide policy learning (as in ATM) or to drive visual servoing.
ExampleClicking a few points on a cup’s handle in a video of a person pouring water, and tracking their trajectories over time, gives a policy-learning target for “how the cup should move.”
- Also called
- TAP, Point Tracking, Long-Term Point Tracking
- Related
- TAPIR · CoTracker · Optical Flow · Scene Flow · ATM · 3D Point Tracking
- Sources
- TAP-Vid: A Benchmark for Tracking Any Point in a Video (arXiv 2211.03726)
CoTracker: It is Better to Track Together (arXiv 2307.07635)