<link rel="stylesheet" href="styles.f3b1fba60ec7970c.css">

Publication:
Joint audio-video processing for biometric speaker identification

Loading...
Thumbnail Image

Departments

School / College / Institute

Item type:Organizational Unit,

Program

Organization Authors

Co-Authors

Date

Language

Embargo Status

N/A

Journal Title

Journal ISSN

Volume Title

Alternative Title

Abstract

We present a bimodal audio-visual speaker identification system. The objective is to improve the recognition performance over conventional unimodal schemes. The proposed system exploits not only the temporal and spatial correlations existing in the speech and video signals of a speaker, but also the cross-correlation between these two modalities. Lip images extracted from each video frame are transformed onto an eigenspace. The obtained eigenlip coefficients are interpolated to match the rate of the speech signal and fused with Mel frequency cepstral coefficients (MFCC) of the corresponding speech signal. The resulting joint feature vectors are used to train and test a hidden Markov model (HMM) based identification system. Experimental results are included to demonstrate the system performance.

Source

Publisher

Institute of Electrical and Electronics Engineers

Citation

item.page.haspartof

Source

2003 International Conference on Multimedia and Expo, Vol III, Proceedings

item.page.ispartofseries

item.page.edition

DOI

item.page.datauri

item.page.link

Rights

N/A

Copyrights Note

Rights and licensing

N/A

Endorsement

Review

Supplemented By

Referenced By

Related Patent

Related Goal

Google Scholar
Scholar'da Ara ↗
5
Görüntülenme
0
İndirme
Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.