Skip to main navigation Skip to search Skip to main content

Resource-efficient multiview perception: integrating semantic masking with masked autoencoders

  • The University of Sydney

Research output: Chapter in Book / Conference PaperConference Paperpeer-review

Abstract

Multiview systems have become a key technology in modern computer vision, offering advanced capabilities in scene understanding and analysis. However, these systems face critical challenges in bandwidth limitations and computational constraints, particularly for resource-limited camera nodes. This paper presents a novel approach for communication-efficient distributed multiview detection and tracking using masked autoencoders (MAEs). We introduce a semantic-guided masking strategy that leverages pre-trained segmentation models and a tunable power function to prioritize informative image regions. This approach, combined with an MAE, reduces communication overhead while preserving essential visual information. We evaluate our method on both virtual and real-world multiview datasets, demonstrating comparable performance in terms of detection and tracking performance metrics compared to state-of-the-art techniques, even at high masking ratios. Our selective masking algorithm outperforms random masking, maintaining higher accuracy and precision as the masking ratio increases. Furthermore, our approach achieves a significant reduction in transmission data volume compared to baseline methods, thereby balancing multiview tracking performance with communication efficiency.

Original languageEnglish
Title of host publicationProceedings of the 2025 IEEE International Conference on Pervasive Computing and Communications (PerCom 2025), Washington D.C., USA, 17-21 March 2025
Place of PublicationU.S.
PublisherIEEE
Pages145-151
Number of pages7
ISBN (Electronic)9798331535513
DOIs
Publication statusPublished - 2025
Event23rd IEEE International Conference on Pervasive Computing and Communications, PerCom 2025 - Washington, United States
Duration: 17 Mar 202521 Mar 2025

Conference

Conference23rd IEEE International Conference on Pervasive Computing and Communications, PerCom 2025
Country/TerritoryUnited States
CityWashington
Period17/03/2521/03/25

Keywords

  • Communication-Efficient
  • Distributed Vision
  • Masked Autoencoders
  • Multiview
  • Semantic Masking

Fingerprint

Dive into the research topics of 'Resource-efficient multiview perception: integrating semantic masking with masked autoencoders'. Together they form a unique fingerprint.

Cite this