Publication: Overview of the SIGTURK 2026 shared Task: Terminology-aware machine translation for English-Turkish scientific texts
| dc.conference.date | MAR 29, 2026 | |
| dc.conference.location | Rabat, Morocco | |
| dc.contributor.department | Department of Computer Engineering | |
| dc.contributor.department | Graduate School of Sciences and Engineering | |
| dc.contributor.department | KUIS AI (Koç University & İş Bank Artificial Intelligence Center) | |
| dc.contributor.kuauthor | Gebeşçe, Ali | |
| dc.contributor.kuauthor | Şahin, Gözde Gül | |
| dc.contributor.kuauthor | Amasya, Ege Uğur | |
| dc.contributor.kuauthor | Safa, Abdalfatah Rashid | |
| dc.contributor.schoolcollegeinstitute | GRADUATE SCHOOL OF SCIENCES AND ENGINEERING | |
| dc.contributor.schoolcollegeinstitute | College of Engineering | |
| dc.contributor.schoolcollegeinstitute | Research Center | |
| dc.date.accessioned | 2026-08-14T11:25:50Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | This paper presents an overview of the SIG-TURK 2026 Shared Task on Terminology-Aware Machine Translation for English-Turkish Scientific Texts.We address the critical challenge of terminological accuracy in low-resource settings by constructing the first terminology-rich English-Turkish parallel corpus, comprising 3,300 sentence pairs from STEM domains with 10,157 expert-validated term pairs.The shared task consists of three subtasks: term detection, expert-guided correction, and end-to-end post-editing.We evaluate state-of-the-art baselines (including GPT-5.2 and Claude Sonnet 4.5) alongside participant systems employing diverse strategies from fine-tuning to Retrieval-Augmented Generation (RAG).Our results highlight that while massive generalist models dominate zero-shot detection, smaller, domain-adapted models using Supervised Fine-Tuning and Reinforcement Learning can significantly outperform them in end-toend post-editing.Furthermore, we find that rigid retrieval pipelines often disrupt fluency, whereas Chain-of-Thought prompting allows models to integrate terminology more naturally.Despite these advances, a significant gap remains between automated systems and human expert performance in strict terminology correction. | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | Scopus | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | N/A | |
| dc.description.sponsorship | Wikimedia Foundation | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | N/A | |
| dc.identifier.ScopusQuartile | N/A | |
| dc.identifier.WoSPercentile | N/A | |
| dc.identifier.WoSQuartile | N/A | |
| dc.identifier.doi | 10.18653/v1/2026.sigturk-1.20 | |
| dc.identifier.embargo | N/A | |
| dc.identifier.endpage | 247 | |
| dc.identifier.isbn | 9798891763708 | |
| dc.identifier.scopus | 2-s2.0-105042197352 | |
| dc.identifier.startpage | 236 | |
| dc.identifier.uri | http://doi.org/10.18653/v1/2026.sigturk-1.20 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34557 | |
| dc.keywords | Machine translation | |
| dc.keywords | Translation (biology) | |
| dc.keywords | Computational linguistics | |
| dc.keywords | Natural language | |
| dc.keywords | Conjunction (astronomy) | |
| dc.language | eng | |
| dc.publisher | Association for Computational Linguistics | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | Proceedings of the Second Workshop Natural Language Processing for Turkic Languages (Sigturk 2026) | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Language and linguistics | |
| dc.subject | Physical sciences | |
| dc.subject | Computer science | |
| dc.subject | Artificial intelligence | |
| dc.subject | Computer engineering | |
| dc.title | Overview of the SIGTURK 2026 shared Task: Terminology-aware machine translation for English-Turkish scientific texts | |
| dc.type | Conference Proceeding | |
| dspace.entity.type | Publication | |
| relation.isOrgUnitOfPublication | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isOrgUnitOfPublication | 3fc31c89-e803-4eb1-af6b-6258bc42c3d8 | |
| relation.isOrgUnitOfPublication | 77d67233-829b-4c3a-a28f-bd97ab5c12c7 | |
| relation.isOrgUnitOfPublication.latestForDiscovery | 89352e43-bf09-4ef4-82f6-6f9d0174ebae | |
| relation.isParentOrgUnitOfPublication | 434c9663-2b11-4e66-9399-c863e2ebae43 | |
| relation.isParentOrgUnitOfPublication | 8e756b23-2d4a-4ce8-b1b3-62c794a8c164 | |
| relation.isParentOrgUnitOfPublication | d437580f-9309-4ecb-864a-4af58309d287 | |
| relation.isParentOrgUnitOfPublication.latestForDiscovery | 434c9663-2b11-4e66-9399-c863e2ebae43 |
