Publication:
A deep learning approach for data driven vocal tract area function estimation

dc.conference.dateDEC 18-21, 2018
dc.conference.locationAthens, GREECE
dc.conference.organizer2018 IEEE Workshop on Spoken Language Technology (SLT 2018)
dc.contributor.departmentMVGL (Multimedia, Vision and Graphics Laboratory)
dc.contributor.facultymemberYes
dc.contributor.kuauthorAsadiabadi, Sasan
dc.contributor.kuauthorErzin, Engin
dc.contributor.schoolcollegeinstituteLaboratory
dc.date.accessioned2024-11-09T23:43:45Z
dc.date.issued2018
dc.description.abstractIn this paper we present a data driven vocal tract area function (VTAF) estimation using Deep Neural Networks (DNN). We approach the VTAF estimation problem based on sequence to sequence learning neural networks, where regression over a sliding window is used to learn arbitrary non-linear one-to-many mapping from the input feature sequence to the target articulatory sequence. We propose two schemes for efficient estimation of the VTAF; (1) a direct estimation of the area function values and (2) an indirect estimation via predicting the vocal tract boundaries. We consider acoustic speech and phone sequence as two possible input modalities for the DNN estimators. Experimental evaluations are performed over a large data comprising acoustic and phonetic features with parallel articulatory information from the USC-TIMIT database. Our results show that the proposed direct and indirect schemes perform the VTAF estimation with mean absolute error (MAE) rates lower than 1.65 mm, where the direct estimation scheme is observed to perform better than the indirect scheme.
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.openaccessYES
dc.description.publisherscopeInternational
dc.description.sponsoredbyTubitakEuN/A
dc.description.sponsorshipWe thank to NVIDIA for donating Titan XP within the GPU Grant Program.
dc.description.studentonlypublicationNo
dc.description.studentpublicationYes
dc.identifier.doi10.1109/SLT.2018.8639582
dc.identifier.endpage173
dc.identifier.isbn9781538643341
dc.identifier.issn2639-5479
dc.identifier.scopus2-s2.0-85063083027
dc.identifier.startpage167
dc.identifier.urihttps://hdl.handle.net/20.500.14288/13549
dc.identifier.urihttps://doi.org/10.1109/SLT.2018.8639582
dc.identifier.wos000463141800025
dc.keywordsSpeech articulation
dc.keywordsVocal tract area function
dc.keywordsDeep neural network
dc.keywordsConvolutional neural network articulatory movements
dc.keywordsNeural-networks
dc.keywordsSpeech
dc.keywordsShape
dc.language.isoeng
dc.publisherInstitute of Electrical and Electronics Engineers
dc.relation.ispartofIEEE Workshop on Spoken Language Technology
dc.subjectComputer science
dc.subjectArtificial intelligence
dc.subjectEngineering
dc.subjectElectrical electronic engineering
dc.titleA deep learning approach for data driven vocal tract area function estimation
dc.typeConference Proceeding
dspace.entity.typePublication
local.contributor.kuauthorAsadiabadi, Sasan
local.contributor.kuauthorErzin, Engin
relation.isOrgUnitOfPublicationcb6bbbf6-fd19-4052-b581-f591a9748d21
relation.isOrgUnitOfPublication.latestForDiscoverycb6bbbf6-fd19-4052-b581-f591a9748d21
relation.isParentOrgUnitOfPublication20385dee-35e7-484b-8da6-ddcc08271d96
relation.isParentOrgUnitOfPublication.latestForDiscovery20385dee-35e7-484b-8da6-ddcc08271d96

Files