TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus

This paper describes a new Slovak speech recognition dedicated corpus built from TEDx talks and Jump Slovakia lectures. The proposed speech database consists of 220 talks and lectures in total duration of about 58 hours. Annotated speech database was generated automatically in an unsupervised manner by using acoustic speech segmentation based on principal component analysis and automatic speech transcription using two complementary speech recognition systems. The evaluation data consisting of 50 manually annotated talks and lectures in total duration of about 12 hours, has been created for evaluation of the quality of Slovak speech recognition. By unsupervised automatic annotation of TEDx talks and Jump Slovakia lectures we have obtained 21.26% of new speech segments with approximately 9.44% word error rate, suitable for retraining or adaptation of acoustic models trained beforehand.

eISSN:: 1338-4287
ISSN:: 0021-5597
Language:: English

Publication timeframe:: 2 times per year
Journal Subjects:: Linguistics and Semiotics, Theoretical Frameworks and Disciplines, Linguistics, other

Journal RSS Feed

TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus

Published Online: Jan 24, 2018

Page range: 346 - 354

DOI: https://doi.org/10.1515/jazcas-2017-0044

Keywords
automatic annotation, speech recognition, speech corpus

© 2017 Ján Staš et al., published by De Gruyter Open

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.

TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus

Published Online: Jan 24, 2018

Page range: 346 - 354

DOI: https://doi.org/10.1515/jazcas-2017-0044

Keywordsautomatic annotation, speech recognition, speech corpus

© 2017 Ján Staš et al., published by De Gruyter Open

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.

Keywords
automatic annotation, speech recognition, speech corpus