Academic literature on the topic 'Syllable-based ASR'

Create a spot-on reference in APA, MLA, Chicago, Harvard, and other styles

Select a source type:

Consult the lists of relevant articles, books, theses, conference reports, and other scholarly sources on the topic 'Syllable-based ASR.'

Next to every source in the list of references, there is an 'Add to bibliography' button. Press on it, and we will generate automatically the bibliographic reference to the chosen work in the citation style you need: APA, MLA, Harvard, Chicago, Vancouver, etc.

You can also download the full text of the academic publication as pdf and read online its abstract whenever available in the metadata.

Journal articles on the topic "Syllable-based ASR"

1

Galatang, Danny Henry, and Suyanto Suyanto. "Syllable-Based Indonesian Automatic Speech Recognition." International Journal on Electrical Engineering and Informatics 12, no. 4 (2020): 720–28. http://dx.doi.org/10.15676/ijeei.2020.12.4.2.

Full text
Abstract:
The syllable-based automatic speech recognition (ASR) systems commonly perform better than the phoneme-based ones. This paper focuses on developing an Indonesian monosyllable-based ASR (MSASR) system using an ASR engine called SPRAAK and comparing it to a phoneme-based one. The Mozilla DeepSpeech-based end-to-end ASR (MDSE2EASR), one of the state-of-the-art models based on character (similar to the phoneme-based model), is also investigated to confirm the result. Besides, a novel Kaituoxu SpeechTransformer (KST) E2EASR is also examined. Testing on the Indonesian speech corpus of 5,439 words sh
APA, Harvard, Vancouver, ISO, and other styles
2

Setiyaningsih, Timor, Mohd Sanusi Azmi, and Azah Kamilah Draman. "Syllable Segmentation with Vowel Detection on Verse Quranic Recitation." JOIV : International Journal on Informatics Visualization 8, no. 4 (2024): 2428. https://doi.org/10.62527/joiv.8.4.2663.

Full text
Abstract:
In speech recognition, segmentation involves partitioning a continuous audio signal containing speech into smaller units or segments, such as words, phonemes, or syllables. This process is paramount in speech recognition systems, as it delineates the boundaries between distinct speech elements, facilitating subsequent analysis and processing. Segmentation accuracy significantly impacts speech recognition systems' overall precision and performance, enabling more precise identification and processing of individual speech units. Moreover, proper segmentation empowers the automatic speech recognit
APA, Harvard, Vancouver, ISO, and other styles
3

Valizada, Alakbar. "DEVELOPMENT OF A REAL-TIME SPEECH RECOGNITION SYSTEM FOR THE AZERBAIJANI LANGUAGE." Problems of Information Society 14, no. 2 (2023): 55–60. http://dx.doi.org/10.25045/jpis.v14.i2.07.

Full text
Abstract:
This paper investigates the development of a real-time automatic speech recognition system dedicated to the Azerbaijani language, focusing on addressing the prevalent gap in speech recognition system for underrepresented languages. Our research integrates a hybrid acoustic modeling approach that combines Hidden Markov Model and Deep Neural Network to interpret the complexities of Azerbaijani acoustic patterns effectively. Recognizing the agglutinative nature of Azerbaijani, the ASR system employs a syllable-based n-gram model for language modeling, ensuring the system accurately captures the s
APA, Harvard, Vancouver, ISO, and other styles
4

Mahesha, P., and D. S. Vinod. "Gaussian Mixture Model Based Classification of Stuttering Dysfluencies." Journal of Intelligent Systems 25, no. 3 (2016): 387–99. http://dx.doi.org/10.1515/jisys-2014-0140.

Full text
Abstract:
AbstractThe classification of dysfluencies is one of the important steps in objective measurement of stuttering disorder. In this work, the focus is on investigating the applicability of automatic speaker recognition (ASR) method for stuttering dysfluency recognition. The system designed for this particular task relies on the Gaussian mixture model (GMM), which is the most widely used probabilistic modeling technique in ASR. The GMM parameters are estimated from Mel frequency cepstral coefficients (MFCCs). This statistical speaker-modeling technique represents the fundamental characteristic so
APA, Harvard, Vancouver, ISO, and other styles
5

Perrine, Brittany L., Ronald C. Scherer, and Jason A. Whitfield. "Signal Interpretation Considerations When Estimating Subglottal Pressure From Oral Air Pressure." Journal of Speech, Language, and Hearing Research 62, no. 5 (2019): 1326–37. http://dx.doi.org/10.1044/2018_jslhr-s-17-0432.

Full text
Abstract:
Purpose Oral air pressure measurements during lip occlusion for /pVpV/ syllable strings are used to estimate subglottal pressure during the vowel. Accuracy of this method relies on smoothly produced syllable repetitions. The purpose of this study was to investigate the oral air pressure waveform during the /p/ lip occlusions and propose physiological explanations for nonflat shapes. Method Ten adult participants were trained to produce the “standard condition” and were instructed to produce nonstandard tasks. Results from 8 participants are included. The standard condition required participant
APA, Harvard, Vancouver, ISO, and other styles
6

Rui, Xian Yi, Yi Biao Yu, and Ying Jiang. "Connected Mandarin Digit Speech Recognition Using Two-Layer Acoustic Universal Structure." Advanced Materials Research 846-847 (November 2013): 1380–83. http://dx.doi.org/10.4028/www.scientific.net/amr.846-847.1380.

Full text
Abstract:
Because of the single-syllable of Chinese words and the confusing nature of Chinese pronunciation, connected mandarin digit speech recognition (CMDSR) is a challenging task in the field of speech recognition. This paper applied a novel acoustic representation of speech, called the acoustic universal structure (AUS) where the non-linguistic variations such as vocal tract length, lines and noises are well removed. A two-layer matching strategy based on the AUS models of speech, including the digit and string AUS models, is proposed for connected mandarin digit speech recognition. The speech reco
APA, Harvard, Vancouver, ISO, and other styles
7

Rong, Panying. "Neuromotor Control of Speech and Speechlike Tasks: Implications From Articulatory Gestures." Perspectives of the ASHA Special Interest Groups 5, no. 5 (2020): 1324–38. http://dx.doi.org/10.1044/2020_persp-20-00070.

Full text
Abstract:
Purpose This study aimed to provide a preliminary examination of the articulatory control of speech and speechlike tasks based on a gestural framework and identify shared and task-specific articulatory factors in speech and speechlike tasks. Method Ten healthy participants performed two speechlike tasks (i.e., alternating motion rate [AMR] and sequential motion rate [SMR]) and three speech tasks (i.e., reading of “clever Kim called the cat clinic” at the regular, fast, and slow rates) that varied in phonological complexity and rate. Articulatory kinematics were recorded using an electromagneti
APA, Harvard, Vancouver, ISO, and other styles
8

Van der Burg, Erik, and Patrick T. Goodbourn. "Rapid, generalized adaptation to asynchronous audiovisual speech." Proceedings of the Royal Society B: Biological Sciences 282, no. 1804 (2015): 20143083. http://dx.doi.org/10.1098/rspb.2014.3083.

Full text
Abstract:
The brain is adaptive. The speed of propagation through air, and of low-level sensory processing, differs markedly between auditory and visual stimuli; yet the brain can adapt to compensate for the resulting cross-modal delays. Studies investigating temporal recalibration to audiovisual speech have used prolonged adaptation procedures, suggesting that adaptation is sluggish. Here, we show that adaptation to asynchronous audiovisual speech occurs rapidly. Participants viewed a brief clip of an actor pronouncing a single syllable. The voice was either advanced or delayed relative to the correspo
APA, Harvard, Vancouver, ISO, and other styles
9

Abimbola, Damilola W., and Akin Ademuyiwa. "Differences in Russian and English Pronunciation." Artha Journal of Social Sciences 22, no. 4 (2024): 23–45. https://doi.org/10.12724/ajss.67.2.

Full text
Abstract:
It might be difficult for English speakers to learn Russian since there are significant phonetic discrepancies between the two languages. The vowel system is one of the biggest distinctions between the two languages; Russian has 10 vowels while English only has five. Compared to English vowels, Russian vowels are uttered more clearly and for longer periods of time, which frequently produces a more nasalized sound. In contrast, diphthongs are found in English but not in Russian. Several Russian consonant sounds, including "Ж" and "Ш" are absent from the English language when it comes to consona
APA, Harvard, Vancouver, ISO, and other styles
10

Aye, Nyein Mon. "IMPROVING MYANMAR AUTOMATIC SPEECH RECOGNITION WITH OPTIMIZATION OF CONVOLUTIONAL NEURAL NETWORK PARAMETERS." January 30, 2019. https://doi.org/10.5281/zenodo.2553028.

Full text
Abstract:
Researchers of many nations have developed automatic speech recognition (ASR) to show their national improvement in information and communication technology for their languages. This work intends to improve the ASR performance for Myanmar language by changing different Convolutional Neural Network (CNN) hyperparameters such as number of feature maps and pooling size. CNN has the abilities of reducing in spectral variations and modeling spectral correlations that exist in the signal due to the locality and pooling operation. Therefore, the impact of the hyperparameters on CNN accuracy in ASR ta
APA, Harvard, Vancouver, ISO, and other styles
More sources

Dissertations / Theses on the topic "Syllable-based ASR"

1

DavidSarwono and 林立成. "A Syllable Cluster Based Weighted Kernel Feature Matrix for ASR Substitution Error Correction." Thesis, 2012. http://ndltd.ncl.edu.tw/handle/28550786350424171660.

Full text
Abstract:
碩士<br>國立成功大學<br>資訊工程學系碩博士班<br>100<br>Abstract A Syllable Cluster Based Weighted Kernel Feature Matrix for ASR Substitution Error Correction David Sarwono * Chung-Hsien Wu** Institute of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan, R.O.C. In recent years Automatic Speech Recognition (ASR) technology has become one of the most growing technologies in engineering science and research. However, the performance of ASR technology is still restricted in adverse environments. Errors in Automatic Speech Recognition outputs lead to low performance for
APA, Harvard, Vancouver, ISO, and other styles
2

Anoop, C. S. "Automatic speech recognition for low-resource Indian languages." Thesis, 2023. https://etd.iisc.ac.in/handle/2005/6195.

Full text
Abstract:
Building good models for automatic speech recognition (ASR) requires large amounts of annotated speech data. Recent advancements in end-to-end speech recognition have aggravated the need for data. However, most Indian languages are low-resourced and lack enough training data to build robust and efficient ASR systems. Despite the challenges associated with the scarcity of data, Indian languages offer some unique characteristics that can be utilized to improve speech recognition in low-resource settings. Most languages have an overlapping phoneme set and a strong correspondence between their cha
APA, Harvard, Vancouver, ISO, and other styles

Books on the topic "Syllable-based ASR"

1

Stein, Gabriele. Peter Levins’ description of word-formation (1570). Oxford University Press, 2017. http://dx.doi.org/10.1093/oso/9780198807377.003.0008.

Full text
Abstract:
One of the most original English lexicographical ventures in the sixteenth century was Peter Levins’ Manipulus vocabulorum (1570). This is the first English rhyming dictionary. Some nine thousand English words were arranged in the alphabetical order of their last syllable and then translated into Latin. Levins’ word selection will thus have been largely based on the sound structure of the lexical items. The long years spent by Levins on assembling and arranging the dictionary material inevitably drew his attention to English suffixes like -able, -er, -ish, -less, and -ness and such second elem
APA, Harvard, Vancouver, ISO, and other styles
2

Uffmann, Christian. World Englishes and Phonological Theory. Edited by Markku Filppula, Juhani Klemola, and Devyani Sharma. Oxford University Press, 2015. http://dx.doi.org/10.1093/oxfordhb/9780199777716.013.32.

Full text
Abstract:
The relationship between phonological theory and World Englishes is generally characterized by a mutual lack of interest. This chapter argues for a greater engagement of both fields with each other, looking at constraint-based theories of phonology, especially Optimality Theory (OT), as a case in point. Contact varieties of English provide strong evidence for synchronically active constraints, as it is substrate or L1 constraints that are regularly transferred to the contact variety, not rules. Additionally, contact varieties that have properties that are in some way ‘in between’ the substrate
APA, Harvard, Vancouver, ISO, and other styles
3

Özçelik, Öner. The Phonology of Turkish. Oxford University PressOxford, 2024. http://dx.doi.org/10.1093/oso/9780192869722.001.0001.

Full text
Abstract:
Abstract The Phonology of Turkish offers a comprehensive overview and analysis of the phonological structure of modern Turkish. While phenomena at both segmental and suprasegmental levels are discussed, the emphasis is on the latter, analyzing phonological processes extending over a number of different domains. Couched within a primarily constraint-based framework, lower-level prosodic constituents, including syllables, feet, and prosodic words, are incorporated into a general theory with higher-level constituents, the Phonological Phrase and the Intonational Phrase, assuming that phonological
APA, Harvard, Vancouver, ISO, and other styles

Book chapters on the topic "Syllable-based ASR"

1

Kim, Byeongchang, Junhwi Choi, and Gary Geunbae Lee. "ASR Error Management Using RNN Based Syllable Prediction for Spoken Dialog Applications." In Advances in Parallel and Distributed Computing and Ubiquitous Services. Springer Singapore, 2016. http://dx.doi.org/10.1007/978-981-10-0068-3_12.

Full text
APA, Harvard, Vancouver, ISO, and other styles
2

Schneider, Anne H., Johannes Hellrich, and Saturnino Luz. "Word, Syllable and Phoneme Based Metrics Do Not Correlate with Human Performance in ASR-Mediated Tasks." In Advances in Natural Language Processing. Springer International Publishing, 2014. http://dx.doi.org/10.1007/978-3-319-10888-9_39.

Full text
APA, Harvard, Vancouver, ISO, and other styles
3

Heslop, Kate. "A Poetry Machine." In Viking Mediologies. Fordham University Press, 2022. http://dx.doi.org/10.5422/fordham/9780823298242.003.0007.

Full text
Abstract:
Analysis of the two manuscript versions of the Second Grammatical Treatise reveals a common interest in musical performance, which is also reflected in the treatise’s tripartite division of sound, indebted to medieval music theory. Music and grammar meet in the ars rithmica, an analytical tradition devoted to syllable-counting, often rhyming kinds of poetry usually performed to musical accompaniment. The “new poetics” of the late twelfth and thirteenth centuries, influenced by ars rithmica, posits meter as a tool for the renewal of poetry based on the best of old traditions. The influence of a
APA, Harvard, Vancouver, ISO, and other styles

Conference papers on the topic "Syllable-based ASR"

1

Myo, Saw Kyaw Htin, Moe Thet, Hein Htun, Arkar Myint, La Min Htut, and Htut Shie. "Syllable-Based Myanmar Speech-to-Text ASR using Deep Learning and CTC." In 2024 Conference of Young Researchers in Electrical and Electronic Engineering (ElCon). IEEE, 2024. http://dx.doi.org/10.1109/elcon61730.2024.10468151.

Full text
APA, Harvard, Vancouver, ISO, and other styles
2

Ryu, Hyuksu, Minsu Na, and Minhwa Chung. "Pronunciation modeling of loanwords for Korean ASR using phonological knowledge and syllable-based segmentation." In 2015 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA). IEEE, 2015. http://dx.doi.org/10.1109/apsipa.2015.7415308.

Full text
APA, Harvard, Vancouver, ISO, and other styles
3

Liu, Chao-Hong, Chung-Hsien Wu, David Sarwono, and Jhing-Fa Wang. "Candidate generation for ASR output error correction using a context-dependent syllable cluster-based confusion matrix." In Interspeech 2011. ISCA, 2011. http://dx.doi.org/10.21437/interspeech.2011-488.

Full text
APA, Harvard, Vancouver, ISO, and other styles
4

Hromada, Daniel, and Hyungjoong Kim. "Digital Primer Implementation of Human-Machine Peer Learning for Reading Acquisition: Introducing Curriculum 2." In 10th International Conference on Human Interaction and Emerging Technologies (IHIET 2023). AHFE International, 2023. http://dx.doi.org/10.54941/ahfe1004027.

Full text
Abstract:
The aim of the digital primer project is cognitive enrichment and fostering of acquisition of basic literacy and numeracy of 5 – 10 year old children. Here, we focus on Primer's ability to accurately process child speech which is fundamental to the acquisition of reading component of the Primer. We first note that automatic speech recognition (ASR) and speech-to-text of child speech is a challenging task even for large-scale, cloud-based ASR systems. Given that the Primer is an embedded AI artefact which aims to perform all computations on edge devices like RaspberryPi or Nvidia Jetson, the ta
APA, Harvard, Vancouver, ISO, and other styles
5

Qu, Zhongdi, Parisa Haghani, Eugene Weinstein, and Pedro Moreno. "Syllable-based acoustic modeling with CTC-SMBR-LSTM." In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2017. http://dx.doi.org/10.1109/asru.2017.8268932.

Full text
APA, Harvard, Vancouver, ISO, and other styles
6

Moungsri, Decha, Tomoki Koriyama, and Takao Kobayashi. "Enhanced F0 generation for GPR-based speech synthesis considering syllable-based prosodic features." In 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 2017. http://dx.doi.org/10.1109/apsipa.2017.8282285.

Full text
APA, Harvard, Vancouver, ISO, and other styles
7

Sakamoto, Nagisa, Kazumasa Yamamoto, and Seiichi Nakagawa. "Combination of syllable based N-gram search and word search for spoken term detection through spoken queries and IV/OOV classification." In 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU). IEEE, 2015. http://dx.doi.org/10.1109/asru.2015.7404795.

Full text
APA, Harvard, Vancouver, ISO, and other styles
We offer discounts on all premium plans for authors whose works are included in thematic literature selections. Contact us to get a unique promo code!