Publication:
Class-agnostic visio-temporal scene sketch semantic segmentation

dc.conference.dateFEB 28-MAR 04, 2025
dc.conference.locationTucson, AZ, USA
dc.contributor.departmentKUIS AI (Koç University & İş Bank Artificial Intelligence Center)
dc.contributor.kuauthorKütük, Aleyna
dc.contributor.kuauthorSezgin, Tevfik Metin
dc.contributor.schoolcollegeinstituteResearch Center
dc.date.accessioned2026-08-14T11:20:53Z
dc.date.issued2025
dc.description.abstractScene sketch semantic segmentation is a crucial task for various applications including sketch-to-image retrieval and scene understanding. Existing sketch segmentation methods treat sketches as bitmap images, leading to the loss of temporal order among strokes due to the shift from vector to image format. Moreover, these methods struggle to segment objects from categories absent in the training data. In this paper, we propose a Class-Agnostic Visio-Temporal Network (CAVT) for scene sketch semantic segmentation. CAVT employs a class-agnostic object detector to detect individual objects in a scene and groups the strokes of instances through its post-processing module. This is the first approach that performs segmentation at both the instance and stroke levels within scene sketches. Furthermore, there is a lack of free-hand scene sketch datasets with both instance and stroke-level class annotations. To fill this gap, we collected the largest Free-hand Instance-and Stroke-level Scene Sketch Dataset (FrISS) that contains 1K scene sketches and covers 403 object classes with dense annotations. Extensive experiments on FrISS and other datasets demonstrate the superior performance of our method over state-of-the-art scene sketch segmenttion models. Our code and dataset can be accessed from https://github.com/aleynakutuk6/CAVT.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuTÜBİTAK
dc.description.sponsorshipThis study is written as a part of a Research Project supported by grants from the Scientific and Technological Research Council of Turkey (TUBITAK), Turkey (Project No. 120E489). We are also thankful for the support of the council and KUIS AI Center.
dc.description.versionPublished Version
dc.identifier.ScopusPercentileN/A
dc.identifier.ScopusQuartileN/A
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.1109/wacv61041.2025.00818
dc.identifier.embargoN/A
dc.identifier.endpage8453
dc.identifier.grantno120E489
dc.identifier.isbn9798331510848
dc.identifier.issn2472-6737
dc.identifier.scopus2-s2.0-105003622388
dc.identifier.startpage8444
dc.identifier.urihttp://doi.org/10.1109/wacv61041.2025.00818
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34340
dc.identifier.wos001521272600328
dc.keywordsSketch
dc.keywordsComputer science
dc.keywordsClass (philosophy)
dc.keywordsArtificial intelligence
dc.keywordsSegmentation
dc.keywordsNatural language processing
dc.keywordsComputer vision
dc.keywordsAlgorithm
dc.languageeng
dc.publisherIEEE
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofIEEE/CVF Winter Conference on Applications of Computer Vision
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectComputer science
dc.subjectComputer vision and pattern recognition
dc.titleClass-agnostic visio-temporal scene sketch semantic segmentation
dc.title.alternativeSınıftan bağımsız görsel-zamansal sahne taslağı anlamsal bölümlemesi
dc.typeConference Proceeding
dspace.entity.typePublication
relation.isOrgUnitOfPublication77d67233-829b-4c3a-a28f-bd97ab5c12c7
relation.isOrgUnitOfPublication.latestForDiscovery77d67233-829b-4c3a-a28f-bd97ab5c12c7
relation.isParentOrgUnitOfPublicationd437580f-9309-4ecb-864a-4af58309d287
relation.isParentOrgUnitOfPublication.latestForDiscoveryd437580f-9309-4ecb-864a-4af58309d287

Files