Primary tabs
2016
Van Heerden, C. J., Kleynhans, N. ., & Davel, M. H. (2016). Improving the Lwazi ASR baseline. In Interspeech. San Francisco, USA. http://doi.org/http://dx.doi.org/10.21437/Interspeech.2016-1412
We investigate the impact of recent advances in speech recognition techniques for under-resourced languages. Specifically, we review earlier results published on the Lwazi ASR corpus of South African languages, and experiment with additional acoustic modeling approaches. We demonstrate large gains by applying current state-of-the-art techniques, even if the data itself is neither extended nor improved. We analyze the various performance improvements observed, report on comparative performance per technique – across all eleven languages in the corpus – and discuss the implications of our findings for under-resourced languages in general.
@{153,
author = {Charl Van Heerden and Neil Kleynhans and Marelie Davel},
title = {Improving the Lwazi ASR baseline},
abstract = {We investigate the impact of recent advances in speech recognition techniques for under-resourced languages. Specifically, we review earlier results published on the Lwazi ASR corpus of South African languages, and experiment with additional acoustic modeling approaches. We demonstrate large gains by applying current state-of-the-art techniques, even if the data itself is neither extended nor improved. We analyze the various performance improvements observed, report on comparative performance per technique – across all eleven languages in the corpus – and discuss the implications of our findings for under-resourced languages in general.},
year = {2016},
journal = {Interspeech},
pages = {3529-3538},
month = {08/09-12/09},
address = {San Francisco, USA},
doi = {http://dx.doi.org/10.21437/Interspeech.2016-1412},
}
2015
Wissing, D. ., Pienaar, W. ., & Van Niekerk, D. R. (2015). Palatalisation of /s/ in Afrikaans. Stellenbosch Papers in Linguistics Plus, 48. http://doi.org/10.5842/48-0-688
This article reports on the investigation of the acoustic characteristics of the Afrikaans voiceless alveolar fricative /s/[1]. As yet, a palatal [ʃ] for /s/ has been reported only in a limited case, namely where /s/ is followed by palatal /j/, for example in the phrase is jy (‘are you’), pronounced as [ə-ʃəi]. This seems to be an instance of regressive coarticulation, resulting in coalescence of basic /s/ and /j/. The present study revealed that, especially in the pronunciation of young, white Afrikaans-speakers, /s/ is also palatalised progressively when preceded by /r/ in the coda cluster /rs/, and, to a lesser extent, also in other contexts where /r/ is involved, for example across syllable and word boundaries. Only a slight presence of palatalisation was detected in the production of /s/ in the speech of the white, older speakers of the present study. This finding might be indicative of a definite change in the Afrikaans consonant system. A post hoc reflection is offered here on the possible presence of /s/-fronting, especially in the speech of the younger females. Such pronunciation could very well be a prestige marker for affluent speakers of Afrikaans.
@article{293,
author = {Daan Wissing and Wikus Pienaar and Daniel Van Niekerk},
title = {Palatalisation of /s/ in Afrikaans},
abstract = {This article reports on the investigation of the acoustic characteristics of the Afrikaans voiceless alveolar fricative /s/[1]. As yet, a palatal [ʃ] for /s/ has been reported only in a limited case, namely where /s/ is followed by palatal /j/, for example in the phrase is jy (‘are you’), pronounced as [ə-ʃəi]. This seems to be an instance of regressive coarticulation, resulting in coalescence of basic /s/ and /j/. The present study revealed that, especially in the pronunciation of young, white Afrikaans-speakers, /s/ is also palatalised progressively when preceded by /r/ in the coda cluster /rs/, and, to a lesser extent, also in other contexts where /r/ is involved, for example across syllable and word boundaries. Only a slight presence of palatalisation was detected in the production of /s/ in the speech of the white, older speakers of the present study. This finding might be indicative of a definite change in the Afrikaans consonant system. A post hoc reflection is offered here on the possible presence of /s/-fronting, especially in the speech of the younger females. Such pronunciation could very well be a prestige marker for affluent speakers of Afrikaans.},
year = {2015},
journal = {Stellenbosch Papers in Linguistics Plus},
volume = {48},
pages = {137-158},
publisher = {Stellenbosch University},
doi = {10.5842/48-0-688},
}
Modipa, T. ., & Davel, M. H. (2015). Predicting vowel substitution in code-switched speech. In Pattern Recognition Association of South Africa (PRASA). Port Elizabeth, South Africa. http://doi.org/10.1109/RoboMech.2015.7359515
The accuracy of automatic speech recognition (ASR) systems typically degrades when encountering code-switched speech. Some of this degradation is due to the unexpected pronunciation effects introduced when languages are mixed. Embedded (foreign) phonemes typically show more variation than phonemes from the matrix language: either approximating the embedded language pronunciation fairly closely, or realised as any of a set of phonemic counterparts from the matrix language. In this paper we describe a technique for predicting the phoneme substitutions that are expected to occur during code-switching, using non-acoustic features only. As case study we consider Sepedi/English code switching and analyse the different realisations of the English schwa. A code-switched speech corpus is used as input and vowel substitutions identified by auto-tagging this corpus based on acoustic characteristics. We first evaluate the accuracy of our auto-tagging process, before determining the predictability of our auto-tagged corpus, using non-acoustic features.
@{292,
author = {Thipe Modipa and Marelie Davel},
title = {Predicting vowel substitution in code-switched speech},
abstract = {The accuracy of automatic speech recognition (ASR) systems typically degrades when encountering code-switched speech. Some of this degradation is due to the unexpected pronunciation effects introduced when languages are mixed. Embedded (foreign) phonemes typically show more variation than phonemes from the matrix language: either approximating the embedded language pronunciation fairly closely, or realised as any of a set of phonemic counterparts from the matrix language. In this paper we describe a technique for predicting the phoneme substitutions that are expected to occur during code-switching, using non-acoustic features only. As case study we consider Sepedi/English code switching and analyse the different realisations of the English schwa. A code-switched speech corpus is used as input and vowel substitutions identified by auto-tagging this corpus based on acoustic characteristics. We first evaluate the accuracy of our auto-tagging process, before determining the predictability of our auto-tagged corpus, using non-acoustic features.},
year = {2015},
journal = {Pattern Recognition Association of South Africa (PRASA)},
pages = {154-159},
month = {26/11-27/11},
address = {Port Elizabeth, South Africa},
isbn = {978-1-4673-7450-7, 978-1-4673-7449-1},
doi = {10.1109/RoboMech.2015.7359515},
}
Kleynhans, N. ., & Barnard, E. . (2015). Efficient data selection for ASR. Language Resources and Evaluation, 49(2). http://doi.org/10.1007/s10579-014-9285-0
Automatic speech recognition (ASR) technology has matured over the past few decades and has made significant impacts in a variety of fields, from assistive technologies to commercial products. However, ASR system development is a resource intensive activity and requires language resources in the form of text annotated audio recordings and pronunciation dictionaries. Unfortunately, many languages found in the developing world fall into the resource-scarce category and due to this resource scarcity the deployment of ASR systems in the developing world is severely inhibited. One approach to assist with resource-scarce ASR system development, is to select 'useful' training samples which could reduce the resources needed to collect new corpora. In this work, we propose a new data selection framework which can be used to design a speech recognition corpus. We show for limited data sets, independent of language and bandwidth, the most effective strategy for data selection is frequency-matched selection and that the widely-used maximum entropy methods generally produced the least promising results. In our model, the frequency-matched selection method corresponds to a logarithmic relationship between accuracy and corpus size; we also investigated other model relationships, and found that a hyperbolic relationship (as suggested from simple asymptotic arguments in learning theory) may lead to somewhat better performance under certain conditions.
@article{291,
author = {Neil Kleynhans and Etienne Barnard},
title = {Efficient data selection for ASR},
abstract = {Automatic speech recognition (ASR) technology has matured over the past few decades and has made significant impacts in a variety of fields, from assistive technologies to commercial products. However, ASR system development is a resource intensive activity and requires language resources in the form of text annotated audio recordings and pronunciation dictionaries. Unfortunately, many languages found in the developing world fall into the resource-scarce category and due to this resource scarcity the deployment of ASR systems in the developing world is severely inhibited. One approach to assist with resource-scarce ASR system development, is to select 'useful' training samples which could reduce the resources needed to collect new corpora. In this work, we propose a new data selection framework which can be used to design a speech recognition corpus. We show for limited data sets, independent of language and bandwidth, the most effective strategy for data selection is frequency-matched selection and that the widely-used maximum entropy methods generally produced the least promising results. In our model, the frequency-matched selection method corresponds to a logarithmic relationship between accuracy and corpus size; we also investigated other model relationships, and found that a hyperbolic relationship (as suggested from simple asymptotic arguments in learning theory) may lead to somewhat better performance under certain conditions.},
year = {2015},
journal = {Language Resources and Evaluation},
volume = {49},
pages = {327-353},
issue = {2},
publisher = {Springer Science+Business Media},
address = {Dordrecht},
doi = {10.1007/s10579-014-9285-0},
}
Kleynhans, N. ., De Wet, F. ., & Barnard, E. . (2015). Unsupervised acoustic model training: comparing South African English and isiZulu. In Pattern Recognition Association of South Africa (PRASA). Port Elizabeth, South Africa. http://doi.org/ 10.1109/RoboMech.2015.7359512
Large amounts of untranscribed audio data are generated every day. These audio resources can be used to develop robust acoustic models that can be used in a variety of speech-based systems. Manually transcribing this data is resource intensive and requires funding, time and expertise. Lightly-supervised training techniques, however, provide a means to rapidly transcribe audio, thus reducing the initial resource investment to begin the modelling process. Our findings suggest that the lightly-supervised training technique works well for English but when moving to an agglutinative language, such as isiZulu, the process fails to achieve the performance seen for English. Additionally, phone-based performances are significantly worse when compared to an approach using word-based language models. These results indicate a strong dependence on large or well-matched text resources for lightly-supervised training techniques.
@{290,
author = {Neil Kleynhans and Febe De Wet and Etienne Barnard},
title = {Unsupervised acoustic model training: comparing South African English and isiZulu},
abstract = {Large amounts of untranscribed audio data are generated every day. These audio resources can be used to develop robust acoustic models that can be used in a variety of speech-based systems. Manually transcribing this data is resource intensive and requires funding, time and expertise. Lightly-supervised training techniques, however, provide a means to rapidly transcribe audio, thus reducing the initial resource investment to begin the modelling process. Our findings suggest that the lightly-supervised training technique works well for English but when moving to an agglutinative language, such as isiZulu, the process fails to achieve the performance seen for English. Additionally, phone-based performances are significantly worse when compared to an approach using word-based language models. These results indicate a strong dependence on large or well-matched text resources for lightly-supervised training techniques.},
year = {2015},
journal = {Pattern Recognition Association of South Africa (PRASA)},
pages = {136-141},
address = {Port Elizabeth, South Africa},
isbn = {978-1-4673-7450-7, 978-1-4673-7449-1},
doi = {10.1109/RoboMech.2015.7359512},
}
Giwa, O. ., & Davel, M. H. (2015). Text-based Language Identification of Multilingual Names. In Pattern Recognition Association of South Africa (PRASA). Port Elizabeth, South Africa. http://doi.org/ 10.1109/RoboMech.2015.7359517
Text-based language identification (T-LID) of isolated words has been shown to be useful for various speech processing tasks, including pronunciation modelling and data categorisation. When the words to be categorised are proper names, the task becomes more difficult: not only do proper names often have idiosyncratic spellings, they are also often considered to be multilingual. We, therefore, investigate how an existing T-LID technique can be adapted to perform multilingual word classification. That is, given a proper name, which may be either mono- or multilingual, we aim to determine how accurately we can predict how many possible source languages the word has, and what they are. Using a Joint Sequence Model-based approach to T-LID and the SADE corpus - a newly developed proper names corpus of South African names - we experiment with different approaches to multilingual T-LID. We compare posterior-based and likelihood-based methods and obtain promising results on a challenging task.
@{289,
author = {Oluwapelumi Giwa and Marelie Davel},
title = {Text-based Language Identification of Multilingual Names},
abstract = {Text-based language identification (T-LID) of isolated words has been shown to be useful for various speech processing tasks, including pronunciation modelling and data categorisation. When the words to be categorised are proper names, the task becomes more difficult: not only do proper names often have idiosyncratic spellings, they are also often considered to be multilingual. We, therefore, investigate how an existing T-LID technique can be adapted to perform multilingual word classification. That is, given a proper name, which may be either mono- or multilingual, we aim to determine how accurately we can predict how many possible source languages the word has, and what they are. Using a Joint Sequence Model-based approach to T-LID and the SADE corpus - a newly developed proper names corpus of South African names - we experiment with different approaches to multilingual T-LID. We compare posterior-based and likelihood-based methods and obtain promising results on a challenging task.},
year = {2015},
journal = {Pattern Recognition Association of South Africa (PRASA)},
pages = {166-171},
address = {Port Elizabeth, South Africa},
isbn = {978-1-4673-7450-7, 978-1-4673-7449-1},
doi = {10.1109/RoboMech.2015.7359517},
}
Davel, M. H., Barnard, E. ., Van Heerden, C. J., Hartman, W. ., Karakos, D. ., Schwartz, R. ., & Tsakalidis, S. . (2015). Exploring minimal pronunciation modeling for low resource languages. In Interspeech. Dresden, Germany.
Pronunciation lexicons can range from fully graphemic (modeling each word using the orthography directly) to fully phonemic (first mapping each word to a phoneme string). Between these two options lies a continuum of modeling options. We analyze techniques that can improve the accuracy of a graphemic system without requiring significant effort to design or implement. The analysis is performed in the context of the IARPA Babel project, which aims to develop spoken term detection systems for previously unseen languages rapidly, and with minimal human effort. We consider techniques related to letter-to-sound mapping and language-independent syllabification of primarily graphemic systems, and discuss results obtained for six languages: Cebuano, Kazakh, Kurmanji Kurdish, Lithuanian, Telugu and Tok Pisin.
@{288,
author = {Marelie Davel and Etienne Barnard and Charl Van Heerden and William Hartman and Damianos Karakos and Richard Schwartz and Stavros Tsakalidis},
title = {Exploring minimal pronunciation modeling for low resource languages},
abstract = {Pronunciation lexicons can range from fully graphemic (modeling each word using the orthography directly) to fully phonemic (first mapping each word to a phoneme string). Between these two options lies a continuum of modeling options. We analyze techniques that can improve the accuracy of a graphemic system without requiring significant effort to design or implement. The analysis is performed in the context of the IARPA Babel project, which aims to develop spoken term detection systems for previously unseen languages rapidly, and with minimal human effort. We consider techniques related to letter-to-sound mapping and language-independent syllabification of primarily graphemic systems, and discuss results obtained for six languages: Cebuano, Kazakh, Kurmanji Kurdish, Lithuanian, Telugu and Tok Pisin.},
year = {2015},
journal = {Interspeech},
pages = {538-542},
address = {Dresden, Germany},
}
Badenhorst, J. ., & Davel, M. H. (2015). Synthetic triphones from trajectory-based feature distributions. In Pattern Recognition Association of South Africa (PRASA). Port Elizabeth, South Africa. http://doi.org/10.1109/RoboMech.2015.7359509
We experiment with a new method to create synthetic models of rare and unseen triphones in order to supplement limited automatic speech recognition (ASR) training data. A trajectory model is used to characterise seen transitions at the spectral level, and these models are then used to create features for unseen or rare triphones. We find that a fairly restricted model (piece-wise linear with three line segments per channel of a diphone transition) is able to represent training data quite accurately. We report on initial results when creating additional triphones for a single-speaker data set, finding small but significant gains, especially when adding additional samples of rare (rather than unseen) triphones.
@{287,
author = {Jaco Badenhorst and Marelie Davel},
title = {Synthetic triphones from trajectory-based feature distributions},
abstract = {We experiment with a new method to create synthetic models of rare and unseen triphones in order to supplement limited automatic speech recognition (ASR) training data. A trajectory model is used to characterise seen transitions at the spectral level, and these models are then used to create features for unseen or rare triphones. We find that a fairly restricted model (piece-wise linear with three line segments per channel of a diphone transition) is able to represent training data quite accurately. We report on initial results when creating additional triphones for a single-speaker data set, finding small but significant gains, especially when adding additional samples of rare (rather than unseen) triphones.},
year = {2015},
journal = {Pattern Recognition Association of South Africa (PRASA)},
pages = {118-122},
address = {Port Elizabeth, South Africa},
isbn = {978-1-4673-7450-7, 978-1-4673-7449-1},
doi = {10.1109/RoboMech.2015.7359509},
}
Pagination
- First page
- Previous page
- 1
- 2
- 3
- 4
- 5


