Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models

Publication:
Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models

Files

Primary IR06906.pdf (2.81 MB)

Departments

Organizational Unit

Department of Electrical and Electronics Engineering

Organizational Unit

KUIS AI (Koç University & İş Bank Artificial Intelligence Center)

School / College / Institute

Organizational Unit

College of Engineering

Organizational Unit

Research Center

KU-Authors

Korkmaz, Cansu

Tekalp, Ahmet Murat

Doğan, Zafer

Date

2026

Type

Journal Article

Embargo Status

No

Abstract

Super-resolution (SR) is an ill-posed inverse problem with many feasible solutions consistent with a given low-resolution image. On one hand, regressive SR models aim to balance fidelity and perceptual quality to yield a single solution, but this trade-off often introduces artifacts that create ambiguity in information-critical applications such as recognizing digits or letters. On the other hand, diffusion models generate a diverse set of SR images, but selecting the most trustworthy solution from this set remains a challenge. This paper introduces a robust, automated framework for identifying the most trustworthy SR sample from a diffusion-generated set by leveraging the semantic reasoning capabilities of vision-language models (VLMs). Specifically, VLMs such as BLIP-2, GPT-4o, and their variants are prompted with structured queries to assess semantic correctness, visual quality, and artifact presence. The top-ranked SR candidates are then ensembled to yield a single trustworthy output in a cost-effective manner. To rigorously assess the validity of VLM-selected samples, we propose a novel Trustworthiness Score (TWS) a hybrid metric that quantifies SR reliability based on three complementary components: semantic similarity via CLIP embeddings, structural integrity using SSIM on edge maps, and artifact sensitivity through multi-level wavelet decomposition. We empirically show that TWS correlates strongly with human preference in both ambiguous and natural images, and that VLM-guided selections consistently yield high TWS values. Compared to conventional metrics like PSNR, LPIPS, which fail to reflect information fidelity, our approach offers a principled, scalable, and generalizable solution for navigating the uncertainty of the diffusion SR space. By aligning outputs with human expectations and semantic correctness, this work sets a new benchmark for trustworthiness in generative SR.

Publisher

IEEE

Subject

Computer vision

Source

IEEE Transactions on Circuits and Systems for Video Technology

DOI

10.1109/tcsvt.2025.3585092

URI

https://hdl.handle.net/20.500.14288/32598
https://doi.org/10.1109/tcsvt.2025.3585092

Rights

CC BY (Attribution)

Creative Commons license

Except where otherwised noted, this item's license is described as CC BY (Attribution)

Publication: Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models

Files

Departments

School / College / Institute

Program

KU-Authors

KU Authors

Co-Authors

Editor & Affiliation

Compiler & Affiliation

Translator

Other Contributor

Date

Language

Type

Embargo Status

Journal Title

Journal ISSN

Volume Title

Alternative Title

Abstract

Source

Publisher

Subject

Citation

Has Part

Source

Book Series Title

Edition

DOI

URI

item.page.datauri

Link

Rights

Copyrights Note

Creative Commons license

Collections

Endorsement

Review

Supplemented By

Referenced By

Related Goal

2

Views

3

Downloads

Publication:
Leveraging vision-language models to select trustworthy super-resolution samples generated by diffusion models