Evaluating Monocular Depth Estimation on Embedded Platforms for Autonomous Navigation

Authors

DOI:

https://doi.org/10.22456/2175-2745.150996

Keywords:

depth estimation, autonomous navigation, visual perception, self-driving car

Abstract

Deep learning-based monocular depth estimation has achieved significant advancements on urban benchmarks, but its embedded application remains limited by efficiency constraints. Vision Transformers (ViTs) and Foundation Models (FMs) show promising zero-shot generalization capabilities, yet their adaptation to resource-constrained hardware requires careful study. In this work, we investigate the development of the DepthAnything model on an NVIDIA Jetson Orin, analyzing the trade-off between accuracy and inference speed for different backbones (ViT-S, ViT-B, and ViT-L). We report quantitative metrics including AbsRel, δ1, RMSE, and FPS on the KITTI dataset, along with qualitative results. Our experiments show that the ViT-S backbone offers the best balance of accuracy and real-time performance (44 FPS), whereas ViT-B suffers from degradation and ViT-L exhibits significant instability due to optimization artifacts. These findings highlight the viability of compact backbones for embedded visual perception and suggest future optimizations, such as quantization-aware training and pruning, in larger architectures.

Downloads

Download data is not yet available.

References

[1] HORN, B. K. Shape from shading: A method for obtaining the shape of a smooth opaque object from one view. 1970.

[2] BRUNO, D. R. et al. Carina project: Visual perception systems applied for autonomous vehicles and advanced driver assistance systems (adas). IEEE Access, IEEE, 2023. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/10155133⟩.

[3] ZHANG, Z. et al. Review of monocular depth estimation methods. Journal of Electronic Imaging, SPIE, v. 34, n. 2, p. 020901, 2025. Disponível em: ⟨https://doi.org/10.1117/1.JEI.34.2.020901⟩.

[4] KRIZHEVSKY, A.; SUTSKEVER, I.; HINTON, G. E. Imagenet classification with deep convolutional neural networks. In: PEREIRA, F. et al. (Ed.). Advances in Neural Information Processing Systems. Curran Associates, Inc., 2012. v. 25. Disponível em: ⟨https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf⟩.

[5] HE, K. et al. Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. [S.l.: s.n.], 2016. p. 770–778.

[6] WIGGINS, W. F.; TEJANI, A. S. On the opportunities and risks of foundation models for natural language processing in radiology. Radiology: Artificial Intelligence, Radiological Society of North America, v. 4, n. 4, p. e220119, 2022. Disponível em: ⟨https://pubs.rsna.org/doi/full/10.1148/ryai.220119⟩.

[7] RADFORD, A. et al. Learning transferable visual models from natural language supervision. In: PMLR. International conference on machine learning. 2021. p. 8748–8763. Disponível em: ⟨https://proceedings.mlr.press/v139/radford21a.html⟩.

[8] VASWANI, A. et al. Attention is all you need. In: GUYON, I. et al. (Ed.). Advances in Neural Information Processing Systems. Curran Associates, Inc., 2017. v. 30. Disponível em: ⟨https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf⟩.

[9] DOSOVITSKIY, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations. [s.n.], 2021. Disponível em: ⟨https://arxiv.org/abs/2010.11929⟩.

[10] RADFORD, A. et al. Learning transferable visual models from natural language supervision. In: MEILA, M.; ZHANG, T. (Ed.). Proceedings of the 38th International Conference on Machine Learning. PMLR, 2021. (Proceedings of Machine Learning Research, v. 139), p. 8748–8763. Disponível em: ⟨https://proceedings.mlr.press/v139/radford21a.html⟩.

[11] ZHOU, X. et al. Vision language models in autonomous driving: A survey and outlook. IEEE Transactions on Intelligent Vehicles, p. 1–20, 2024. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/10531702⟩.

[12] RANFTL, R. et al. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis and machine intelligence, IEEE, v. 44, n. 3, p. 1623–1637, 2020. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/9178977⟩.

[13] BHAT, S. F. et al. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023. Disponível em: ⟨https://arxiv.org/abs/2302.12288⟩.

[14] YIN, W. et al. Metric3d: Towards zero-shot metric 3d prediction from a single image. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). [s.n.], 2023. p. 9043–9053. Disponível em: ⟨https://arxiv.org/abs/2307.10984⟩.

[15] JI, G.; ZHU, Z. Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher. In: LAROCHELLE, H. et al. (Ed.). Advances in Neural Information Processing Systems. Curran Associates, Inc., 2020. v. 33, p. 20823–20833. Disponível em: ⟨https://proceedings.neurips.cc/paper_files/paper/2020/file/ef0d3930a7b6c95bd2b32ed45989c61f-Paper.pdf⟩.

[16] BOCHKOVSKII, A. et al. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073, 2024. Disponível em: ⟨https://arxiv.org/abs/2410.02073⟩.

[17] YANG, L. et al. Depth anything: Unleashing the power of large-scale unlabeled data. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [s.n.], 2024. p. 10371–10381. Disponível em: ⟨https://arxiv.org/abs/2401.10891⟩.

[18] GUO, Y. et al. Depth any camera: Zero-shot metric depth estimation from any camera. In: Proceedings of the Computer Vision and Pattern Recognition Conference. [s.n.], 2025. p. 26996–27006. Disponível em: ⟨https://arxiv.org/abs/2501.02464⟩.

[19] OQUAB, M. et al. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. Disponível em: ⟨https://arxiv.org/abs/2304.07193⟩.

[20] MAN, A.; ABUAJWA, O. M. S.; AZIZ, N. H. A. A review of techniques for lightweight monocular depth estimation. In: 2024 IEEE Asia-Pacific Conference on Applied Electromagnetics (APACE). [s.n.], 2024. p. 351–354. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/10877399⟩.

[21] DAO, T.-T.; PHAM, Q.-V.; HWANG, W.-J. Fastmde: A fast cnn architecture for monocular depth estimation at high resolution. IEEE Access, v. 10, p. 16111–16122, 2022. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/9690863⟩.

[22] LIU, X. et al. Real-time monocular depth estimation merging vision transformers on edge devices for aiot. IEEE Transactions on Instrumentation and Measurement, v. 72, p. 1–9, 2023. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/10091191⟩.

[23] ŁUCKI, J. et al. Visual perception engine: Fast and flexible multi-head inference for robotic vision tasks. arXiv preprint arXiv:2508.11584, 2025. Disponível em: ⟨https://arxiv.org/abs/2508.11584⟩.

[24] UHRIG, J. et al. Sparsity invariant cnns. In: 2017 International Conference on 3D Vision (3DV). [s.n.], 2017. p. 11–20. Disponível em: ⟨https://ieeexplore.ieee.org/abstract/document/8374553⟩.

Downloads

Published

2026-03-23

How to Cite

Chiuyari Veramendi, W. N., & Santos Osório, F. (2026). Evaluating Monocular Depth Estimation on Embedded Platforms for Autonomous Navigation. Revista De Informática Teórica E Aplicada, 33(2), 364–369. https://doi.org/10.22456/2175-2745.150996

Issue

Section

WVC2025

Most read articles by the same author(s)

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.