Publication:
Modeling morphologically rich languages using splitwords and unstructured dependencies

dc.conference.dateAUG 2–7, 2009
dc.conference.locationSuntec, Singapore
dc.conference.organizerAssociation for Computational Linguistic
dc.contributor.departmentDepartment of Computer Engineering
dc.contributor.facultymemberYes
dc.contributor.kuauthorBiçici, Ergün
dc.contributor.kuauthorYüret, Deniz
dc.contributor.schoolcollegeinstituteCollege of Engineering
dc.date.accessioned2024-11-09T23:50:51Z
dc.date.issued2009
dc.description.abstractWe experiment with splitting words into their stem and suffix components for modeling morphologically rich languages. We show that using a morphological analyzer and disambiguator results in a significant perplexity reduction in Turkish. We present flexible n-gram models, Flex-Grams, which assume that the n-1 tokens that determine the probability of a given token can be chosen anywhere in the sentence rather than the preceding n-1 positions. Our final model achieves 27% perplexity reduction compared to the standard n-gram model.
dc.description.fulltextYes
dc.description.harvestedfromManual
dc.description.indexedbyScopus
dc.description.openaccessGold OA
dc.description.peerreviewstatusPeer-Reviewed
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuN/A
dc.description.studentonlypublicationNo
dc.description.studentpublicationYes
dc.description.versionPublished Version
dc.identifier.WoSQuartileN/A
dc.identifier.doi10.3115/1667583.1667690
dc.identifier.embargoNo
dc.identifier.endpage348
dc.identifier.isbn9781617382581
dc.identifier.scopus2-s2.0-84859062288
dc.identifier.startpage345
dc.identifier.urihttps://doi.org/10.3115/1667583.1667690
dc.identifier.urihttps://hdl.handle.net/20.500.14288/14609
dc.keywordsComputational linguistics
dc.keywordsNatural language processing systems
dc.keywordsText processing
dc.keywordsMorphological analyzer
dc.keywordsN-gram modeling
dc.keywordsN-gram models
dc.keywordsTurkishs
dc.keywords% reductions
dc.keywordsSplittings
dc.keywordsModeling languages
dc.language.isoeng
dc.publisherAssociation for Computational Linguistics
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofACL-IJCNLP 2009 - Joint Conf. of the 47th Annual Meeting of the Association for Computational Linguistics and 4th Int. Joint Conf. on Natural Language Processing of the AFNLP, Proceedings of the Conf.
dc.relation.openaccessYes
dc.rightsOther
dc.subjectComputer engineering
dc.subjectComputing
dc.subjectScience
dc.titleModeling morphologically rich languages using splitwords and unstructured dependencies
dc.typeConference Proceeding
dspace.entity.typePublication
local.contributor.kuauthorYüret, Deniz
local.contributor.kuauthorBiçici, Ergün
relation.isOrgUnitOfPublication89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isOrgUnitOfPublication.latestForDiscovery89352e43-bf09-4ef4-82f6-6f9d0174ebae
relation.isParentOrgUnitOfPublication8e756b23-2d4a-4ce8-b1b3-62c794a8c164
relation.isParentOrgUnitOfPublication.latestForDiscovery8e756b23-2d4a-4ce8-b1b3-62c794a8c164

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
IR08053.pdf
Size:
145.01 KB
Format:
Adobe Portable Document Format