Publication: Unsupervised learning of Turkish morphology with multiple codebook VQ-VAE
dc.contributor.department | KUIS AI (Koç University & İş Bank Artificial Intelligence Center) | |
dc.contributor.kuauthor | Yüret, Deniz | |
dc.contributor.kuauthor | Kural, Müge | |
dc.contributor.schoolcollegeinstitute | Research Center | |
dc.date.accessioned | 2025-03-06T21:00:31Z | |
dc.date.issued | 2024 | |
dc.description.abstract | This paper presents an interpretable unsupervised morphological learning model, showing comparable performance to supervised models in learning complex morphological rules of Turkish as evidenced by its application to the problem of morphological inflection within the SIGMORPHON Shared Tasks. The significance of our unsupervised approach lies in its alignment with how humans naturally acquire rules from raw data without supervision. To achieve this, we construct a model with multiple codebooks of VQ-VAE employing continuous and discrete latent variables during word generation. We evaluate the model’s performance under high and low-resource scenarios, and use probing techniques to examine encoded information in latent representations. We also evaluate its generalization capabilities by testing unseen suffixation scenarios within the SIGMORPHON-UniMorph 2022 Shared Task 0. Our results demonstrate our model’s ability to distinguish word structures into lemmas and suffixes, with each codebook specialized for different morphological features, contributing to the interpretability of our model and effectively performing morphological inflection on both seen and unseen morphological features. | |
dc.description.indexedby | Scopus | |
dc.description.publisherscope | International | |
dc.description.sponsoredbyTubitakEu | N/A | |
dc.description.sponsorship | We gratefully acknowledge the support of the KUIS AI Center at Koc University, Istanbul, for this work. | |
dc.identifier.isbn | 9798891761407 | |
dc.identifier.quartile | N/A | |
dc.identifier.scopus | 2-s2.0-85204733595 | |
dc.identifier.uri | https://hdl.handle.net/20.500.14288/27907 | |
dc.keywords | Unsupervised learning | |
dc.keywords | Turkish morphology | |
dc.keywords | Multiple codebook | |
dc.keywords | VQ-VAE | |
dc.keywords | Vector quantization | |
dc.keywords | Natural language processing | |
dc.keywords | Morphological analysis | |
dc.keywords | Deep learning | |
dc.keywords | Neural networks | |
dc.keywords | Computational linguistics | |
dc.language.iso | eng | |
dc.publisher | Association for Computational Linguistics (ACL) | |
dc.relation.ispartof | SIGTURK 2024 - 1st Workshop on Natural Language Processing for Turkic Languages, Proceedings of the Workshop | |
dc.subject | Engineering | |
dc.title | Unsupervised learning of Turkish morphology with multiple codebook VQ-VAE | |
dc.type | Conference Proceeding | |
dspace.entity.type | Publication | |
local.publication.orgunit1 | Research Center | |
local.publication.orgunit2 | KUIS AI (Koç University & İş Bank Artificial Intelligence Center) | |
relation.isOrgUnitOfPublication | 77d67233-829b-4c3a-a28f-bd97ab5c12c7 | |
relation.isOrgUnitOfPublication.latestForDiscovery | 77d67233-829b-4c3a-a28f-bd97ab5c12c7 | |
relation.isParentOrgUnitOfPublication | d437580f-9309-4ecb-864a-4af58309d287 | |
relation.isParentOrgUnitOfPublication.latestForDiscovery | d437580f-9309-4ecb-864a-4af58309d287 |