KLT Revisited: Spatio-Temporal Affine Regression for Feature Tracking
DOI:
https://doi.org/10.22456/2175-2745.150903Keywords:
feature, tracking, spatio-temporal, deep neural networksAbstract
Feature association is an essential step for vision-based localization methods. These methods rely on feature matching to estimate relative motion between consecutive frames using projective geometry. Regardless of the advances in feature association, most existing methods still rely on pairwise feature matching approaches and ignore the rich temporal context in image sequences. In this paper, we revisit the well-known Kanade-Lucas-Tomasi (KLT) feature tracker algorithm and propose a differentiable tracker model. The Proposed method is a fully convolutional neural network that learns spatio-temporal features to track keypoints across videos. Experimental results show that the proposed method outperforms the KLT method, especially in challenging environments.
Downloads
References
[1] HARTLEY, R.; ZISSERMAN, A. Multiple View Geometry in Computer Vision. 2. ed. [S.l.]: Cambridge University Press, 2004.
[2] FRAUNDORFER, F.; SCARAMUZZA, D. Visual odometry: Part ii: Matching, robustness, optimization, and applications. IEEE Robotics and Automation Magazine, v. 19, n. 2, p. 78–90, 2012.
[3] SARLIN, P.-E. et al. Superglue: Learning feature matching with graph neural networks. In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). [S.l.: s.n.], 2020. p. 4937–4946.
[4] POTJE, G. et al. Xfeat: Accelerated features for lightweight image matching. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). [S.l.: s.n.], 2024. p. 2682–2691.
[5] LUCAS, B. D.; KANADE, T. An iterative image registration technique with an application to stereo vision. In: Proceedings of the 7th International Joint Conference on Artificial Intelligence - Volume 2. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1981. (IJCAI’81), p. 674–679.
[6] SAND, P.; TELLER, S. Particle video: Long-range motion estimation using point trajectories. In: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06). [S.l.: s.n.], 2006. v. 2, p. 2195–2202.
[7] HARLEY, A. W.; FANG, Z.; FRAGKIADAKI, K. Particle video revisited: Tracking through occlusions using point trajectories. In: ECCV. [S.l.: s.n.], 2022.
[8] DOERSCH, C. et al. TAPIR: Tracking any point with per-frame initialization and temporal refinement. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. [S.l.: s.n.], 2023. p. 10061–10072.
[9] SHI, J.; TOMASI, C. Good features to track. In: 1994 Proceedings of IEEE Conference on Computer Vision and Pattern Recognition. [S.l.: s.n.], 1994. p. 593–600.
[10] CARREIRA, J.; ZISSERMAN, A. Quo vadis, action recognition? a new model and the kinetics dataset. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [S.l.: s.n.], 2017. p. 4724–4733.
[11] JADERBERG, M. et al. Spatial transformer networks. In: Proceedings of the 29th International Conference on Neural Information Processing Systems - Volume 2. Cambridge, MA, USA: MIT Press, 2015. (NIPS’15), p. 2017–2025.
[12] TOMASI, C. Detection and tracking of point features. In: [s.n.], 1991. Disponível em: https://api.semanticscholar.org/CorpusID:238434334.
[13] BOUGUET, J.-Y. Pyramidal implementation of the lucas kanade feature tracker. In: [s.n.], 1999. Disponível em: https://api.semanticscholar.org/CorpusID:9350588.
[14] DENG, J. et al. Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition. [S.l.: s.n.], 2009. p. 248–255.
[15] WANG, X.; JABRI, A.; EFROS, A. A. Learning Correspondence from the Cycle-Consistency of Time. 2019. Disponível em: https://arxiv.org/abs/1903.07593.
[16] JADERBERG, M. et al. Spatial Transformer Networks. 2016. Disponível em: https://arxiv.org/abs/1506.02025.
[17] WANG, Z.; SIMONCELLI, E.; BOVIK, A. Multiscale structural similarity for image quality assessment. In: The Thirty-Seventh Asilomar Conference on Signals, Systems and Computers, 2003. [S.l.: s.n.], 2003. v. 2, p. 1398–1402 Vol.2.
[18] LI, Z.; SNAVELY, N. Megadepth: Learning single-view depth prediction from internet photos. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.: s.n.], 2018. p. 2041–2050.
[19] DOERSCH, C. et al. TAP-Vid: A Benchmark for Tracking Any Point in a Video. 2023. Disponível em: https://arxiv.org/abs/2211.03726.
[20] GREFF, K. et al. Kubric: a scalable dataset generator. 2022.
[21] STURM, J. et al. A benchmark for the evaluation of rgb-d slam systems. In: Proc. of the International Conference on Intelligent Robot Systems (IROS). [S.l.: s.n.], 2012.
[22] DETONE, D.; MALISIEWICZ, T.; RABINOVICH, A. SuperPoint: Self-Supervised Interest Point Detection and Description. 2018. Disponível em: https://arxiv.org/abs/1712.07629.
[23] HE, K. et al. Deep Residual Learning for Image Recognition. 2015. Disponível em: https://arxiv.org/abs/1512.03385.
[24] DEWANCKER, I.; MCCOURT, M.; CLARK, S. Bayesian optimization primer. 2015. Disponível em: https://app.sigopt.com/static/pdf/SigOpt_Bayesian_Optimization_Primer.pdf.
[25] BRADSKI, G. The OpenCV Library. Dr. Dobb’s Journal of Software Tools, 2000.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Nigel Joseph Bandeira Dias, Gustavo Teodoro Laureano, Ronaldo Martins Da Costa

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.













