Non-Parallel Articulatory-to-Acoustic Conversion Using Multiview-based Time Warping

González López, José Andrés; Gómez Alanís, Alejandro; Pérez Córdoba, José Luis; Green, Phil D.

doi:10.3390/app12031167

dc.contributor.author	González López, José Andrés
dc.contributor.author	Gómez Alanís, Alejandro
dc.contributor.author	Pérez Córdoba, José Luis
dc.contributor.author	Green, Phil D.
dc.date.accessioned	2022-01-25T07:39:39Z
dc.date.available	2022-01-25T07:39:39Z
dc.date.issued	2022-01-23
dc.identifier.citation	Gonzalez-Lopez, J.A.; Gomez-Alanis, A.; Pérez-Córdoba, J.L.; Green, P.D. Non-Parallel Articulatory-to-Acoustic Conversion Using Multiview-Based Time Warping. Appl. Sci. 2022, 12, 1167. [https://doi.org/10.3390/app12031167]	es_ES
dc.identifier.uri	http://hdl.handle.net/10481/72464
dc.description	This work was supported in part by the Spanish State Research Agency (SRA) grant number PID2019-108040RB-C22/SRA/10.13039/501100011033, and the FEDER/Junta de AndalucíaConsejería de Transformación Económica, Industria, Conocimiento y Universidades project no. B-SEJ-570-UGR20.	es_ES
dc.description.abstract	In this paper, we propose a novel algorithm called multiview temporal alignment by dependence maximisation in the latent space (TRANSIENCE) for the alignment of time series consisting of sequences of feature vectors with different length and dimensionality of the feature vectors. The proposed algorithm, which is based on the theory of multiview learning, can be seen as an extension of the well-known dynamic time warping (DTW) algorithm but, as mentioned, it allows the sequences to have different dimensionalities. Our algorithm attempts to find an optimal temporal alignment between pairs of nonaligned sequences by first projecting their feature vectors into a common latent space where both views are maximally similar. To do this, powerful, nonlinear deep neural network (DNN) models are employed. Then, the resulting sequences of embedding vectors are aligned using DTW. Finally, the alignment paths obtained in the previous step are applied to the original sequences to align them. In the paper, we explore several variants of the algorithm that mainly differ in the way the DNNs are trained. We evaluated the proposed algorithm on a articulatory-to-acoustic (A2A) synthesis task involving the generation of audible speech from motion data captured from the lips and tongue of healthy speakers using a technique known as permanent magnet articulography (PMA). In this task, our algorithm is applied during the training stage to align pairs of nonaligned speech and PMA recordings that are later used to train DNNs able to synthesis speech from PMA data. Our results show the quality of speech generated in the nonaligned scenario is comparable to that obtained in the parallel scenario.	es_ES
dc.description.sponsorship	Spanish State Research Agency (SRA) PID2019-108040RB-C22/SRA/10.13039/501100011033	es_ES
dc.description.sponsorship	FEDER/Junta de AndalucíaConsejería de Transformación Económica, Industria, Conocimiento y Universidades project no. B-SEJ-570-UGR20.	es_ES
dc.language.iso	eng	es_ES
dc.publisher	MDPI	es_ES
dc.rights	Atribución 3.0 España
dc.rights.uri	http://creativecommons.org/licenses/by/3.0/es/
dc.subject	Deep learning	es_ES
dc.subject	Multiview learning	es_ES
dc.subject	Dynamic time warping	es_ES
dc.subject	Canonical correlation analysis	es_ES
dc.subject	Silent speech interface	es_ES
dc.subject	Latent embedding	es_ES
dc.title	Non-Parallel Articulatory-to-Acoustic Conversion Using Multiview-based Time Warping	es_ES
dc.type	journal article	es_ES
dc.rights.accessRights	open access	es_ES
dc.identifier.doi	10.3390/app12031167
dc.type.hasVersion	AM	es_ES

Ficheros en el ítem

Nombre:: applsci-1558375-proofreading.pdf
Tamaño:: 563.7Kb
Formato:: PDF

Este ítem aparece en la(s) siguiente(s) colección(ones)

DTSTC - Artículos

Mostrar el registro sencillo del ítem

Excepto si se señala otra cosa, la licencia del ítem se describe como Atribución 3.0 España