Publication: A vision-language framework for multispectral scene representation using language-grounded features
| dc.conference.date | AUG 3–8, 2025 | |
| dc.conference.location | Brisbane, Australia | |
| dc.contributor.coauthor | Karanfil, E. | |
| dc.contributor.coauthor | Imamoglu, N. | |
| dc.contributor.coauthor | Erdem, E. | |
| dc.contributor.department | Department of Computer Engineering | |
| dc.contributor.kuauthor | Erdem, Aykut | |
| dc.contributor.schoolcollegeinstitute | College of Engineering | |
| dc.date.accessioned | 2026-08-14T11:20:07Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | Scene understanding in remote sensing often faces challenges in generating accurate representations for complex environments such as various land use areas or coastal regions, which may also include snow, clouds or haze. To address this, we present a vision-language framework named Spectral-LLaVA, which integrates multispectral data with vision-language alignment techniques to enhance scene representation and description. Using the BigEarthNet-v2 dataset from Sentinel-2, we establish a baseline with RGB-based scene descriptions and further demonstrate substantial improvements through the incorporation of multispectral information. Our framework optimizes a lightweight linear projection layer for alignment while keeping the vision backbone of SpectralGPT frozen. Our experiments encompass (1) scene classification using linear probing, and (2) language modeling for jointly performing scene classification and description generation. Our results highlight Spectral-LLaVA’s ability to produce detailed and accurate descriptions, particularly for scenarios where RGB data alone proves inadequate, while also enhancing classification performance by refining SpectralGPT features into semantically meaningful representations. The code and dataset for this project are available here. | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | WOS | |
| dc.description.indexedby | Scopus | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | N/A | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | 36 | |
| dc.identifier.ScopusQuartile | Q3 | |
| dc.identifier.WoSPercentile | N/A | |
| dc.identifier.WoSQuartile | N/A | |
| dc.identifier.doi | 10.1109/igarss55030.2025.11242427 | |
| dc.identifier.embargo | N/A | |
| dc.identifier.endpage | 6259 | |
| dc.identifier.isbn | 9798331508111 | |
| dc.identifier.issn | 2153-6996 | |
| dc.identifier.scopus | 2-s2.0-105034029333 | |
| dc.identifier.startpage | 6255 | |
| dc.identifier.uri | http://doi.org/10.1109/igarss55030.2025.11242427 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34284 | |
| dc.identifier.wos | 001704609200650 | |
| dc.keywords | Multispectral image | |
| dc.keywords | Representation (politics) | |
| dc.keywords | RGB color model | |
| dc.keywords | Projection (relational algebra) | |
| dc.keywords | Code (set theory) | |
| dc.keywords | Baseline (sea) | |
| dc.keywords | Pattern recognition (psychology) | |
| dc.language | eng | |
| dc.publisher | IEEE | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | Igarss 2025 - 2025 IEEE International Geoscience and Remote Sensing Symposium | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Geography | |
| dc.subject | Geosciences | |
| dc.subject | Instruments and instrumentation | |
| dc.subject | Imaging science | |
| dc.subject | Photographic technology | |
| dc.title | A vision-language framework for multispectral scene representation using language-grounded features | |
| dc.type | Conference Proceeding | |
| dspace.entity.type | Publication | |
| relation.isOrgUnitOfPublication | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isOrgUnitOfPublication.latestForDiscovery | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isParentOrgUnitOfPublication | 8e756b23-2d4a-4ce8-b1b3-62c794a8c164 | |
| relation.isParentOrgUnitOfPublication.latestForDiscovery | 8e756b23-2d4a-4ce8-b1b3-62c794a8c164 |
