Publication:
Content-adaptive inference for state-of-the-art learned video compression

dc.contributor.coauthorYilmaz, M. A.
dc.contributor.departmentDepartment of Electrical and Electronics Engineering
dc.contributor.departmentGraduate School of Sciences and Engineering
dc.contributor.kuauthorTekalp, Ahmet Murat
dc.contributor.kuauthorBilican, Ahmet
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.contributor.schoolcollegeinstituteGRADUATE SCHOOL OF SCIENCES AND ENGINEERING
dc.date.accessioned2026-08-14T11:21:22Z
dc.date.issued2025
dc.description.abstractWhile the BD-rate performance of recent learned video codec models in both low-delay and random-access modes exceed that of respective modes of traditional codecs on average over common benchmarks, the performance improvements for individual videos with complex/large motions is much smaller compared to scenes with simple motion. This is related to the inability of a learned encoder model to generalize to motion vector ranges that have not been seen in the training set, which causes loss of performance in both coding of flow fields as well as frame prediction and coding. As a remedy, we propose a generic (model-agnostic) framework to control the scale of motion vectors in a scene during inference (encoding) to approximately match the range of motion vectors in the test and training videos by adaptively downsampling frames. This results in down-scaled motion vectors enabling: i) better flow estimation; hence, frame prediction and ii) more efficient flow compression. We show that the proposed framework for content-adaptive inference improves the BD-rate performance of already state-of-the-art low-delay video codec DCVC-FM by up to 41% on individual videos without any model fine tuning. We present ablation studies to show measures of motion and scene complexity can be used to predict the effectiveness of the proposed framework. The code is available at https://github.com/KUIS-AI-Tekalp-Research-Group/video-compression/tree/master/OJSP2025.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuN/A
dc.description.versionPublished Version
dc.identifier.ScopusPercentile66
dc.identifier.ScopusQuartileQ2
dc.identifier.WoSPercentile47,6
dc.identifier.WoSQuartileQ3
dc.identifier.doi10.1109/ojsp.2025.3564817
dc.identifier.eissn2644-1322
dc.identifier.embargoN/A
dc.identifier.endpage506
dc.identifier.issn2644-1322
dc.identifier.scopus2-s2.0-105003938158
dc.identifier.startpage498
dc.identifier.urihttp://doi.org/10.1109/ojsp.2025.3564817
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34361
dc.identifier.volume6
dc.identifier.wos001487987700002
dc.keywordsEncoding
dc.keywordsAdaptation models
dc.keywordsVectors
dc.keywordsTraining
dc.keywordsVideo compression
dc.keywordsVideo codecs
dc.keywordsBidirectional control
dc.keywordsEstimation
dc.keywordsRate-distortion
dc.keywordsContext modeling
dc.keywordsAdaptive inference
dc.keywordsFrame downsampling
dc.keywordsMotion modeling
dc.keywordsMotion domain shift
dc.languageeng
dc.publisherIEEE
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofIEEE Open Journal of Signal Processing
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectEngineering
dc.subjectElectrical and electronic engineering
dc.titleContent-adaptive inference for state-of-the-art learned video compression
dc.typeJournal Article
dspace.entity.typePublication
relation.isOrgUnitOfPublication21598063-a7c5-420d-91ba-0cc9b2db0ea0
relation.isOrgUnitOfPublication3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isOrgUnitOfPublication.latestForDiscovery21598063-a7c5-420d-91ba-0cc9b2db0ea0
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication434c9663-2b11-4e66-9399-c863e2ebae43
relation.isParentOrgUnitOfPublication.latestForDiscovery8e756b23-2d4a-4ce8-b1b3-62c794a8c164

Files