Publication:
A cross-validation study of Turkish sentiment analysis datasets and tools

dc.contributor.coauthorCakici, S.
dc.contributor.coauthorHurriyetoglu, A.
dc.contributor.departmentDepartment of Economics
dc.contributor.departmentKUASIA (Center for Asian Studies)
dc.contributor.kuauthorÇırlan, Mehmet Akif
dc.contributor.kuauthorKaraduman, Dilara
dc.contributor.schoolcollegeinstituteResearch Center
dc.contributor.schoolcollegeinstituteCollege of Administrative Sciences and Economics
dc.date.accessioned2026-08-14T11:20:00Z
dc.date.issued2025
dc.description.abstractIn recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability and reuse of Turkish datasets across studies has yielded highly diverse outcomes. To address this, we conducted a systematic review of sentiment analysis studies on Turkish text. Our search identified 78 relevant studies, from which we extracted over 80 datasets. These studies were labeled using a comprehensive sentiment analysis taxonomy, and the dataset details were compiled into a structured repository. Furthermore, we evaluated the performance of four state-of-the-art models-XLM-T, BERTurk (fine-tuned with the BounTi dataset), TSAM, and TurkishBERTweet-on four widely-used Turkish datasets. Among the models, XLM-T achieved the highest performance with an accuracy of 0.92 and F1 score of 0.95 on the Twt dataset, while TSAM reached 0.97 accuracy and F1 score on the Humir dataset. Our empirical results demonstrate that model performance varies significantly based on dataset characteristics such as domain, balance, and linguistic structure. Our review revealed key research gaps, including the limited application of emotion-based and concept-based sentiment analysis techniques and the lack of domain diversity in Turkish sentiment datasets. By highlighting such gaps and compiling a centralized repository, this study provides a comprehensive and publicly accessible resource to guide future research in Turkish sentiment analysis.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuN/A
dc.description.versionPublished Version
dc.identifier.ScopusPercentile95
dc.identifier.ScopusQuartileQ1
dc.identifier.WoSPercentile28,4
dc.identifier.WoSQuartileQ3
dc.identifier.doi10.1007/s10579-025-09869-6
dc.identifier.eissn1574-0218
dc.identifier.embargoN/A
dc.identifier.endpage4041
dc.identifier.issn1574-020X
dc.identifier.issue4
dc.identifier.scopus2-s2.0-105014161670
dc.identifier.startpage4003
dc.identifier.urihttp://doi.org/10.1007/s10579-025-09869-6
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34267
dc.identifier.volume59
dc.identifier.wos001559398900001
dc.keywordsTurkish
dc.keywordsSentiment analysis
dc.keywordsComputer science
dc.keywordsCross-validation
dc.keywordsData science
dc.keywordsNatural language processing
dc.keywordsTurkish dataset
dc.keywordsTaxonomy
dc.keywordsDeep learning models
dc.languageeng
dc.publisherSpringer
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofLanguage Resources and Evaluation
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectPhysical sciences
dc.subjectComputer science
dc.subjectArtificial intelligence
dc.titleA cross-validation study of Turkish sentiment analysis datasets and tools
dc.typeJournal Article
dspace.entity.typePublication
relation.isOrgUnitOfPublication7ad2a3bb-d8d9-4cbd-a6a3-3ca4b30b40c3
relation.isOrgUnitOfPublicationba8da4aa-eff0-4ed3-9791-9fc3161993ad
relation.isOrgUnitOfPublication.latestForDiscovery7ad2a3bb-d8d9-4cbd-a6a3-3ca4b30b40c3
relation.isParentOrgUnitOfPublicationd437580f-9309-4ecb-864a-4af58309d287
relation.isParentOrgUnitOfPublication972aa199-81e2-499f-908e-6fa3deca434a
relation.isParentOrgUnitOfPublication.latestForDiscoveryd437580f-9309-4ecb-864a-4af58309d287

Files