Abstract
Pixel-level segmentation is critical for remote sensing applications, yet traditional supervised methods suffer from high annotation costs. While foundational vision models like CLIP and the Segment Anything Model (SAM) offer zero-shot capabilities, they struggle with domain-specific challenges in aerial imagery, specifically: (1) scale variation causing attention drift, (2) lack of semantic discrimination leading to mask redundancy, and (3) poor adaptation to overhead perspectives. To bridge this gap without task-specific fine-tuning, VTPSeg is presented as a coarse-to-fine multi-model framework designed for high-precision off-line mapping. Unlike generic integrations, VTPSeg introduces a cohesive semantic-geometric synergy. Specifically, the Grounding DINO+ (GD+) module employs a novel synonym-based prompt strategy to maximize recall for overhead objects. The CLIP Filter++ module then utilizes a dual-prompt mechanism (visual attention circles and negative text constraints) to eliminate false positives caused by background clutter. Finally, these refined priors serve as precise point prompts for FastSAM, ensuring instance-level granularity. Validated on five diverse datasets (WHU, LoveDA, Inria, xBD, and iSAID), VTPSeg achieves state-of-the-art or highly competitive performance across five diverse datasets, demonstrating that strategic prompt engineering can effectively adapt frozen foundational models to complex remote sensing tasks.
| Original language | English |
|---|---|
| Pages (from-to) | 63151-63161 |
| Number of pages | 11 |
| Journal | IEEE Access |
| Volume | 14 |
| DOIs | |
| Publication status | Published - 2026 |
Keywords
- Multi-model
- Remote Sensing
- Visual Prompt
- Visual-Language Model
- Zero-Shot Segmentation
Fingerprint
Dive into the research topics of 'Visual and text prompt segmentation: a novel multi-model framework for remote sensing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver