Publication:
CETVEL: a unified benchmark for evaluating language understanding, generation and cultural capacity of LLMs for Turkish

dc.conference.dateMAR 24-29, 2026
dc.conference.locationRabat, Morocco
dc.contributor.coauthorKesen, I.
dc.contributor.departmentKUIS AI (Koç University & İş Bank Artificial Intelligence Center)
dc.contributor.departmentDepartment of Computer Engineering
dc.contributor.kuauthorErdem, Aykut
dc.contributor.kuauthorŞahin, Gözde Gül
dc.contributor.kuauthorEr, Yakup Abrek
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.contributor.schoolcollegeinstituteResearch Center
dc.date.accessioned2026-08-14T11:20:40Z
dc.date.issued2026
dc.description.abstractWe introduce CETVEL, a comprehensive benchmark designed to evaluate large language models (LLMs) in Turkish.Existing Turkish benchmarks often lack either task diversity or culturally relevant content, or both.CETVEL addresses these gaps by combining a broad range of both discriminative and generative tasks ensuring content that reflects the linguistic and cultural richness of Turkish language.CETVEL covers 23 tasks grouped into seven categories, including tasks such as grammatical error correction, machine translation, and question answering rooted in Turkish history and idiomatic language.We evaluate 33 open-weight LLMs (up to 70B parameters) covering different model families and instruction paradigms.Our experiments reveal that Turkish-centric instruction-tuned models generally underperform relative to multilingual or general-purpose models (e.g.Llama 3 and Mistral), despite being tailored for the language.Moreover, we show that tasks such as grammatical error correction and extractive question answering are particularly discriminative in differentiating model capabilities.CETVEL offers a comprehensive and culturally grounded evaluation suite for advancing the development and assessment of LLMs in Turkish.
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuEU - TÜBİTAK
dc.description.versionPublished Version
dc.identifier.ScopusPercentileN/A
dc.identifier.ScopusQuartileN/A
dc.identifier.WoSPercentileN/A
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.18653/v1/2026.eacl-long.46
dc.identifier.embargoN/A
dc.identifier.endpage1085
dc.identifier.grantno101135671
dc.identifier.grantno121C132
dc.identifier.isbn9798891763807
dc.identifier.scopus2-s2.0-105040538108
dc.identifier.startpage1052
dc.identifier.urihttp://doi.org/10.18653/v1/2026.eacl-long.46
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34328
dc.identifier.volume1
dc.keywordsTurkish
dc.keywordsBenchmark (surveying)
dc.keywordsCultural diversity
dc.languageeng
dc.publisherAssociation for Computational Linguistics
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofProceedings of the 19Th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectPhysical sciences
dc.subjectComputer science
dc.subjectArtificial intelligence
dc.subjectSocial sciences
dc.subjectArts and humanities
dc.subjectLanguage and linguistics
dc.titleCETVEL: a unified benchmark for evaluating language understanding, generation and cultural capacity of LLMs for Turkish
dc.typeConference Proceeding
dspace.entity.typePublication
relation.isOrgUnitOfPublication77d67233-829b-4c3a-a28f-bd97ab5c12c7
relation.isOrgUnitOfPublication89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isOrgUnitOfPublication.latestForDiscovery77d67233-829b-4c3a-a28f-bd97ab5c12c7
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublicationd437580f-9309-4ecb-864a-4af58309d287
relation.isParentOrgUnitOfPublication.latestForDiscovery8e756b23-2d4a-4ce8-b1b3-62c794a8c164

Files