Skip to main navigation Skip to search Skip to main content

Visual and text prompt segmentation: a novel multi-model framework for remote sensing

  • Xing Zi
  • , Kairui Jin
  • , Xian Tao
  • , Jun Li
  • , Ali Braytee
  • , Rajiv Ratn Shah
  • , Karthick Thiyagarajan
  • , Mukesh Prasad
  • University of Technology Sydney
  • CAS - Institute of Automation
  • Indraprastha Institute of Information Technology Delhi

Research output: Contribution to journalArticlepeer-review

1 Downloads (Pure)

Abstract

Pixel-level segmentation is critical for remote sensing applications, yet traditional supervised methods suffer from high annotation costs. While foundational vision models like CLIP and the Segment Anything Model (SAM) offer zero-shot capabilities, they struggle with domain-specific challenges in aerial imagery, specifically: (1) scale variation causing attention drift, (2) lack of semantic discrimination leading to mask redundancy, and (3) poor adaptation to overhead perspectives. To bridge this gap without task-specific fine-tuning, VTPSeg is presented as a coarse-to-fine multi-model framework designed for high-precision off-line mapping. Unlike generic integrations, VTPSeg introduces a cohesive semantic-geometric synergy. Specifically, the Grounding DINO+ (GD+) module employs a novel synonym-based prompt strategy to maximize recall for overhead objects. The CLIP Filter++ module then utilizes a dual-prompt mechanism (visual attention circles and negative text constraints) to eliminate false positives caused by background clutter. Finally, these refined priors serve as precise point prompts for FastSAM, ensuring instance-level granularity. Validated on five diverse datasets (WHU, LoveDA, Inria, xBD, and iSAID), VTPSeg achieves state-of-the-art or highly competitive performance across five diverse datasets, demonstrating that strategic prompt engineering can effectively adapt frozen foundational models to complex remote sensing tasks.

Original languageEnglish
Pages (from-to)63151-63161
Number of pages11
JournalIEEE Access
Volume14
DOIs
Publication statusPublished - 2026

Keywords

  • Multi-model
  • Remote Sensing
  • Visual Prompt
  • Visual-Language Model
  • Zero-Shot Segmentation

Fingerprint

Dive into the research topics of 'Visual and text prompt segmentation: a novel multi-model framework for remote sensing'. Together they form a unique fingerprint.

Cite this