Publication:
A vision-language framework for multispectral scene representation using language-grounded features

dc.conference.dateAUG 3–8, 2025
dc.conference.locationBrisbane, Australia
dc.contributor.coauthorKaranfil, E.
dc.contributor.coauthorImamoglu, N.
dc.contributor.coauthorErdem, E.
dc.contributor.departmentDepartment of Computer Engineering
dc.contributor.kuauthorErdem, Aykut
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.date.accessioned2026-08-14T11:20:07Z
dc.date.issued2025
dc.description.abstractScene understanding in remote sensing often faces challenges in generating accurate representations for complex environments such as various land use areas or coastal regions, which may also include snow, clouds or haze. To address this, we present a vision-language framework named Spectral-LLaVA, which integrates multispectral data with vision-language alignment techniques to enhance scene representation and description. Using the BigEarthNet-v2 dataset from Sentinel-2, we establish a baseline with RGB-based scene descriptions and further demonstrate substantial improvements through the incorporation of multispectral information. Our framework optimizes a lightweight linear projection layer for alignment while keeping the vision backbone of SpectralGPT frozen. Our experiments encompass (1) scene classification using linear probing, and (2) language modeling for jointly performing scene classification and description generation. Our results highlight Spectral-LLaVA’s ability to produce detailed and accurate descriptions, particularly for scenarios where RGB data alone proves inadequate, while also enhancing classification performance by refining SpectralGPT features into semantically meaningful representations. The code and dataset for this project are available here.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuN/A
dc.description.versionPublished Version
dc.identifier.ScopusPercentile36
dc.identifier.ScopusQuartileQ3
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.1109/igarss55030.2025.11242427
dc.identifier.embargoN/A
dc.identifier.endpage6259
dc.identifier.isbn9798331508111
dc.identifier.issn2153-6996
dc.identifier.scopus2-s2.0-105034029333
dc.identifier.startpage6255
dc.identifier.urihttp://doi.org/10.1109/igarss55030.2025.11242427
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34284
dc.identifier.wos001704609200650
dc.keywordsMultispectral image
dc.keywordsRepresentation (politics)
dc.keywordsRGB color model
dc.keywordsProjection (relational algebra)
dc.keywordsCode (set theory)
dc.keywordsBaseline (sea)
dc.keywordsPattern recognition (psychology)
dc.languageeng
dc.publisherIEEE
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofIgarss 2025 - 2025 IEEE International Geoscience and Remote Sensing Symposium
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectGeography
dc.subjectGeosciences
dc.subjectInstruments and instrumentation
dc.subjectImaging science
dc.subjectPhotographic technology
dc.titleA vision-language framework for multispectral scene representation using language-grounded features
dc.typeConference Proceeding
dspace.entity.typePublication
relation.isOrgUnitOfPublication89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isOrgUnitOfPublication.latestForDiscovery89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication.latestForDiscovery8e756b23-2d4a-4ce8-b1b3-62c794a8c164

Files