Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models

Publication:
Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models

dc.contributor.department	KUIS AI (Koç University & İş Bank Artificial Intelligence Center)
dc.contributor.department	Department of Electrical and Electronics Engineering
dc.contributor.kuauthor	Korkmaz, Cansu
dc.contributor.kuauthor	Tekalp, Ahmet Murat
dc.contributor.kuauthor	Doğan, Zafer
dc.contributor.schoolcollegeinstitute	Research Center
dc.contributor.schoolcollegeinstitute	College of Engineering
dc.date.accessioned	2026-07-02T07:03:29Z
dc.date.available	2026-03-27
dc.date.issued	2026
dc.description.abstract	Super-resolution (SR) is an ill-posed inverse problem with many feasible solutions that are consistent with a given low-resolution image. On one hand, regressive SR models aim to balance fidelity and perceptual quality to yield a single solution; but this trade-off often leads to artifacts that introduce ambiguity in information-critical applications such as identifying digits or letters. On the other hand, diffusion models generate a diverse set of SR images; but now selecting the most trustworthy solution out of this set becomes a challenge. This paper introduces a robust, automated framework for identifying the most trustworthy SR sample from a diffusion-generated set by leveraging the semantic reasoning capabilities of vision-language models (VLMs). Specifically, VLMs such as BLIP-2, GPT-4o, and their variants are prompted with structured queries to evaluate semantic correctness, visual quality, and the presence of artifacts. The top-ranked SR candidates are then ensembled to yield a single trustworthy output in a cost-effective manner. To rigorously assess the validity of VLM-selected samples, we propose a novel Trustworthiness Score (TWS)-a hybrid metric that quantifies SR reliability based on three complementary components: semantic similarity using CLIP embeddings, structural integrity via SSIM on edge maps, and artifact sensitivity measured through a multi-level wavelet decomposition. We empirically demonstrate that TWS correlates strongly with human preference in both ambiguous and natural images, and that VLM-guided selections consistently yield high TWS values. Compared to conventional metrics like PSNR, LPIPS, and DISTS-which fail to reflect information fidelity-our approach offers a principled, scalable, and generalizable solution for navigating the uncertainty of the diffusion SR space. By aligning model outputs with human expectations and semantic correctness, this work sets a new benchmark for trustworthiness in generative SR tasks.
dc.description.fulltext	No
dc.description.harvestedfrom	Manual
dc.description.indexedby	WOS
dc.description.indexedby	Scopus
dc.description.openaccess	Green Submitted
dc.description.publisherscope	International
dc.description.readpublish	N/A
dc.description.sponsoredbyTubitakEu	TÜBİTAK
dc.description.sponsorship	The work of Cansu Korkmaz was supported by the Koc University Artificial Intelligence (KUIS AI) Center Fellowship. The work of A. Murat Tekalp was supported in part by TUBITAK 2247-A under Award 120C156 and in part by Turkish Academy of Sciences (TUBA). The work of Zafer Dogan was supported by the TUBITAK 2232 International Fellowship for Outstanding Researchers under Award 118C337.
dc.description.version	Published Version
dc.identifier.WoSQuartile	Q1
dc.identifier.doi	10.1109/TCSVT.2025.3585092
dc.identifier.eissn	1558-2205
dc.identifier.embargo	No
dc.identifier.endpage	1432
dc.identifier.grantno	118C337
dc.identifier.grantno	120C156
dc.identifier.issn	1051-8215
dc.identifier.issue	2
dc.identifier.scopus	2-s2.0-105010047435
dc.identifier.startpage	1419
dc.identifier.uri	https://doi.org10.1016/j.nsa.2026.106990
dc.identifier.uri	https://hdl.handle.net/20.500.14288/32850
dc.identifier.volume	36
dc.identifier.wos	001687411500018
dc.keywords	Training
dc.keywords	Semantics
dc.keywords	Diffusion models
dc.keywords	Measurement
dc.keywords	Accuracy
dc.keywords	Visualization
dc.keywords	Image reconstruction
dc.keywords	Circuits and systems
dc.keywords	Image edge detection
dc.keywords	Generative adversarial networks
dc.keywords	Super-resolution
dc.keywords	Trustworthy SR
dc.keywords	Vision-language models
dc.keywords	Human evaluation
dc.language	eng
dc.publisher	IEEE
dc.relation.affiliation	Koç University
dc.relation.collection	Koç University Institutional Repository
dc.relation.ispartof	IEEE Transactions on Circuits and Systems for Video Technology
dc.relation.openaccess	N/A
dc.rights	N/A
dc.rights.uri	N/A
dc.subject	Engineering
dc.title	Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models
dc.type	Journal Article
dspace.entity.type	Publication
relation.isOrgUnitOfPublication	77d67233-829b-4c3a-a28f-bd97ab5c12c7
relation.isOrgUnitOfPublication	21598063-a7c5-420d-91ba-0cc9b2db0ea0
relation.isOrgUnitOfPublication.latestForDiscovery	77d67233-829b-4c3a-a28f-bd97ab5c12c7
relation.isParentOrgUnitOfPublication	d437580f-9309-4ecb-864a-4af58309d287
relation.isParentOrgUnitOfPublication	8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication.latestForDiscovery	d437580f-9309-4ecb-864a-4af58309d287

Collections

Publications without Fulltext

Publication: Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models

Files

Collections

Publication:
Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models