To see the other types of publications on this topic, follow the link: Syllable-based ASR.

Journal articles on the topic 'Syllable-based ASR'

Create a spot-on reference in APA, MLA, Chicago, Harvard, and other styles

Select a source type:

Consult the top 20 journal articles for your research on the topic 'Syllable-based ASR.'

Next to every source in the list of references, there is an 'Add to bibliography' button. Press on it, and we will generate automatically the bibliographic reference to the chosen work in the citation style you need: APA, MLA, Harvard, Chicago, Vancouver, etc.

You can also download the full text of the academic publication as pdf and read online its abstract whenever available in the metadata.

Browse journal articles on a wide variety of disciplines and organise your bibliography correctly.

1

Galatang, Danny Henry, and Suyanto Suyanto. "Syllable-Based Indonesian Automatic Speech Recognition." International Journal on Electrical Engineering and Informatics 12, no. 4 (2020): 720–28. http://dx.doi.org/10.15676/ijeei.2020.12.4.2.

Full text
Abstract:
The syllable-based automatic speech recognition (ASR) systems commonly perform better than the phoneme-based ones. This paper focuses on developing an Indonesian monosyllable-based ASR (MSASR) system using an ASR engine called SPRAAK and comparing it to a phoneme-based one. The Mozilla DeepSpeech-based end-to-end ASR (MDSE2EASR), one of the state-of-the-art models based on character (similar to the phoneme-based model), is also investigated to confirm the result. Besides, a novel Kaituoxu SpeechTransformer (KST) E2EASR is also examined. Testing on the Indonesian speech corpus of 5,439 words sh
APA, Harvard, Vancouver, ISO, and other styles
2

Setiyaningsih, Timor, Mohd Sanusi Azmi, and Azah Kamilah Draman. "Syllable Segmentation with Vowel Detection on Verse Quranic Recitation." JOIV : International Journal on Informatics Visualization 8, no. 4 (2024): 2428. https://doi.org/10.62527/joiv.8.4.2663.

Full text
Abstract:
In speech recognition, segmentation involves partitioning a continuous audio signal containing speech into smaller units or segments, such as words, phonemes, or syllables. This process is paramount in speech recognition systems, as it delineates the boundaries between distinct speech elements, facilitating subsequent analysis and processing. Segmentation accuracy significantly impacts speech recognition systems' overall precision and performance, enabling more precise identification and processing of individual speech units. Moreover, proper segmentation empowers the automatic speech recognit
APA, Harvard, Vancouver, ISO, and other styles
3

Valizada, Alakbar. "DEVELOPMENT OF A REAL-TIME SPEECH RECOGNITION SYSTEM FOR THE AZERBAIJANI LANGUAGE." Problems of Information Society 14, no. 2 (2023): 55–60. http://dx.doi.org/10.25045/jpis.v14.i2.07.

Full text
Abstract:
This paper investigates the development of a real-time automatic speech recognition system dedicated to the Azerbaijani language, focusing on addressing the prevalent gap in speech recognition system for underrepresented languages. Our research integrates a hybrid acoustic modeling approach that combines Hidden Markov Model and Deep Neural Network to interpret the complexities of Azerbaijani acoustic patterns effectively. Recognizing the agglutinative nature of Azerbaijani, the ASR system employs a syllable-based n-gram model for language modeling, ensuring the system accurately captures the s
APA, Harvard, Vancouver, ISO, and other styles
4

Mahesha, P., and D. S. Vinod. "Gaussian Mixture Model Based Classification of Stuttering Dysfluencies." Journal of Intelligent Systems 25, no. 3 (2016): 387–99. http://dx.doi.org/10.1515/jisys-2014-0140.

Full text
Abstract:
AbstractThe classification of dysfluencies is one of the important steps in objective measurement of stuttering disorder. In this work, the focus is on investigating the applicability of automatic speaker recognition (ASR) method for stuttering dysfluency recognition. The system designed for this particular task relies on the Gaussian mixture model (GMM), which is the most widely used probabilistic modeling technique in ASR. The GMM parameters are estimated from Mel frequency cepstral coefficients (MFCCs). This statistical speaker-modeling technique represents the fundamental characteristic so
APA, Harvard, Vancouver, ISO, and other styles
5

Perrine, Brittany L., Ronald C. Scherer, and Jason A. Whitfield. "Signal Interpretation Considerations When Estimating Subglottal Pressure From Oral Air Pressure." Journal of Speech, Language, and Hearing Research 62, no. 5 (2019): 1326–37. http://dx.doi.org/10.1044/2018_jslhr-s-17-0432.

Full text
Abstract:
Purpose Oral air pressure measurements during lip occlusion for /pVpV/ syllable strings are used to estimate subglottal pressure during the vowel. Accuracy of this method relies on smoothly produced syllable repetitions. The purpose of this study was to investigate the oral air pressure waveform during the /p/ lip occlusions and propose physiological explanations for nonflat shapes. Method Ten adult participants were trained to produce the “standard condition” and were instructed to produce nonstandard tasks. Results from 8 participants are included. The standard condition required participant
APA, Harvard, Vancouver, ISO, and other styles
6

Rui, Xian Yi, Yi Biao Yu, and Ying Jiang. "Connected Mandarin Digit Speech Recognition Using Two-Layer Acoustic Universal Structure." Advanced Materials Research 846-847 (November 2013): 1380–83. http://dx.doi.org/10.4028/www.scientific.net/amr.846-847.1380.

Full text
Abstract:
Because of the single-syllable of Chinese words and the confusing nature of Chinese pronunciation, connected mandarin digit speech recognition (CMDSR) is a challenging task in the field of speech recognition. This paper applied a novel acoustic representation of speech, called the acoustic universal structure (AUS) where the non-linguistic variations such as vocal tract length, lines and noises are well removed. A two-layer matching strategy based on the AUS models of speech, including the digit and string AUS models, is proposed for connected mandarin digit speech recognition. The speech reco
APA, Harvard, Vancouver, ISO, and other styles
7

Rong, Panying. "Neuromotor Control of Speech and Speechlike Tasks: Implications From Articulatory Gestures." Perspectives of the ASHA Special Interest Groups 5, no. 5 (2020): 1324–38. http://dx.doi.org/10.1044/2020_persp-20-00070.

Full text
Abstract:
Purpose This study aimed to provide a preliminary examination of the articulatory control of speech and speechlike tasks based on a gestural framework and identify shared and task-specific articulatory factors in speech and speechlike tasks. Method Ten healthy participants performed two speechlike tasks (i.e., alternating motion rate [AMR] and sequential motion rate [SMR]) and three speech tasks (i.e., reading of “clever Kim called the cat clinic” at the regular, fast, and slow rates) that varied in phonological complexity and rate. Articulatory kinematics were recorded using an electromagneti
APA, Harvard, Vancouver, ISO, and other styles
8

Van der Burg, Erik, and Patrick T. Goodbourn. "Rapid, generalized adaptation to asynchronous audiovisual speech." Proceedings of the Royal Society B: Biological Sciences 282, no. 1804 (2015): 20143083. http://dx.doi.org/10.1098/rspb.2014.3083.

Full text
Abstract:
The brain is adaptive. The speed of propagation through air, and of low-level sensory processing, differs markedly between auditory and visual stimuli; yet the brain can adapt to compensate for the resulting cross-modal delays. Studies investigating temporal recalibration to audiovisual speech have used prolonged adaptation procedures, suggesting that adaptation is sluggish. Here, we show that adaptation to asynchronous audiovisual speech occurs rapidly. Participants viewed a brief clip of an actor pronouncing a single syllable. The voice was either advanced or delayed relative to the correspo
APA, Harvard, Vancouver, ISO, and other styles
9

Abimbola, Damilola W., and Akin Ademuyiwa. "Differences in Russian and English Pronunciation." Artha Journal of Social Sciences 22, no. 4 (2024): 23–45. https://doi.org/10.12724/ajss.67.2.

Full text
Abstract:
It might be difficult for English speakers to learn Russian since there are significant phonetic discrepancies between the two languages. The vowel system is one of the biggest distinctions between the two languages; Russian has 10 vowels while English only has five. Compared to English vowels, Russian vowels are uttered more clearly and for longer periods of time, which frequently produces a more nasalized sound. In contrast, diphthongs are found in English but not in Russian. Several Russian consonant sounds, including "Ж" and "Ш" are absent from the English language when it comes to consona
APA, Harvard, Vancouver, ISO, and other styles
10

Aye, Nyein Mon. "IMPROVING MYANMAR AUTOMATIC SPEECH RECOGNITION WITH OPTIMIZATION OF CONVOLUTIONAL NEURAL NETWORK PARAMETERS." January 30, 2019. https://doi.org/10.5281/zenodo.2553028.

Full text
Abstract:
Researchers of many nations have developed automatic speech recognition (ASR) to show their national improvement in information and communication technology for their languages. This work intends to improve the ASR performance for Myanmar language by changing different Convolutional Neural Network (CNN) hyperparameters such as number of feature maps and pooling size. CNN has the abilities of reducing in spectral variations and modeling spectral correlations that exist in the signal due to the locality and pooling operation. Therefore, the impact of the hyperparameters on CNN accuracy in ASR ta
APA, Harvard, Vancouver, ISO, and other styles
11

Truong Tien, Toan. "ASR - VLSP 2021: An Efficient Transformer-based Approach for Vietnamese ASR Task." VNU Journal of Science: Computer Science and Communication Engineering 38, no. 1 (2022). http://dx.doi.org/10.25073/2588-1086/vnucsce.325.

Full text
Abstract:
Various techniques have been applied to enhance automatic speech recognition during the last few years. Reaching auspicious performance in natural language processing makes Transformer architecture becoming the de facto standard in numerous domains. This paper first presents our effort to collect a 3000-hour Vietnamese speech corpus. After that, we introduce the system used for VLSP 2021 ASR task 2, which is based on the Transformer. Our simple method achieves a favorable syllable error rate of 6.72% and gets second place on the private test. Experimental results indicate that the proposed app
APA, Harvard, Vancouver, ISO, and other styles
12

Manohar, Kavya, Jayan A R, and Rajeev Rajan. "Improving speech recognition systems for the morphologically complex Malayalam language using subword tokens for language modeling." EURASIP Journal on Audio, Speech, and Music Processing 2023, no. 1 (2023). http://dx.doi.org/10.1186/s13636-023-00313-7.

Full text
Abstract:
AbstractThis article presents the research work on improving speech recognition systems for the morphologically complex Malayalam language using subword tokens for language modeling. The speech recognition system is built using a deep neural network–hidden Markov model (DNN-HMM)-based automatic speech recognition (ASR). We propose a novel method, syllable-byte pair encoding (S-BPE), that combines linguistically informed syllable tokenization with the data-driven tokenization method of byte pair encoding (BPE). The proposed method ensures words are always segmented at valid pronunciation bounda
APA, Harvard, Vancouver, ISO, and other styles
13

Thanh, Pham Viet, Le Duc Cuong, Dao Dang Huy, et al. "ASR - VLSP 2021: Semi-supervised Ensemble Model for Vietnamese Automatic Speech Recognition." VNU Journal of Science: Computer Science and Communication Engineering 38, no. 1 (2022). http://dx.doi.org/10.25073/2588-1086/vnucsce.332.

Full text
Abstract:
Automatic speech recognition (ASR) is gaining huge advances with the arrival of End-to-End architectures. Semi-supervised learning methods, which can utilize unlabeled data, have largely contributed to the success of ASR systems, giving them the ability to surpass human performance. However, most of the researches focus on developing these techniques for English speech recognition, which raises concern about their performance in other languages, especially in low-resource scenarios. In this paper, we aim at proposing a Vietnamese ASR system for participating in the VLSP 2021 Automatic Speech R
APA, Harvard, Vancouver, ISO, and other styles
14

Vitale, Vincenzo Norman, Francesco Cutugno, Antonio Origlia, and Gianpaolo Coro. "Exploring emergent syllables in end-to-end automatic speech recognizers through model explainability technique." Neural Computing and Applications, February 13, 2024. http://dx.doi.org/10.1007/s00521-024-09435-1.

Full text
Abstract:
AbstractAutomatic speech recognition systems based on end-to-end models (E2E-ASRs) can achieve comparable performance to conventional ASR systems while reproducing all their essential parts automatically, from speech units to the language model. However, they hide the underlying perceptual processes modelled, if any, and they have lower adaptability to multiple application contexts, and, furthermore, they require powerful hardware and an extensive amount of training data. Model-explainability techniques can explore the internal dynamics of these ASR systems and possibly understand and explain
APA, Harvard, Vancouver, ISO, and other styles
15

Haley, Katarina L., Adam Jacks, Soomin Kim, Marcia Rodriguez, and Lorelei P. Johnson. "Normative Values for Word Syllable Duration With Interpretation in a Large Sample of Stroke Survivors With Aphasia." American Journal of Speech-Language Pathology, August 18, 2023, 1–13. http://dx.doi.org/10.1044/2023_ajslp-22-00300.

Full text
Abstract:
Purpose: Slow speech rate and abnormal temporal prosody are primary diagnostic criteria for differentiating between people with aphasia who do and do not have apraxia of speech. We sought to identify appropriate cutoff values for abnormal word syllable duration (WSD) in a word repetition task, interpret them relative to a data set of people with chronic aphasia, and evaluate the extent to which manually derived measures could be approximated through an automated process that relied on commercial speech recognition technology. Method: Fifty neurotypical participants produced 49 multisyllabic wo
APA, Harvard, Vancouver, ISO, and other styles
16

"Speech Recognition System for Isolated Tamil Words using Random Forest Algorithm." International Journal of Recent Technology and Engineering 9, no. 1 (2020): 2431–35. http://dx.doi.org/10.35940/ijrte.a1467.059120.

Full text
Abstract:
ASR is the use of system software and hardware based techniques to identify and process human voice. In this research, Tamil words are analyzed, segmented as syllables, followed by feature extraction and recognition. Syllables are segmented using short term energy and segmentation is done in order to minimize the corpus size. The algorithm for syllable segmentation works by performing the STE function of the continuous speech signal. The proposed approach for speech recognition uses the combination of Mel-Frequency Cepstral Coefficients (MFCC) and Linear Predictive Coding (LPC). MFCC features
APA, Harvard, Vancouver, ISO, and other styles
17

Hirsch, Aron. "What is the domain for weight computation: the syllable or the interval?" Proceedings of the Annual Meetings on Phonology 1, no. 1 (2014). http://dx.doi.org/10.3765/amp.v1i1.21.

Full text
Abstract:
<p>The distribution of lexical stress is sensitive to the weight of rhythmic units such that heavier units more strongly attract stress. This paper addresses the question: what is the rhythmic unit relevant for weight computation? The traditional approach links weight to the <em>syllable</em>: weight is computed over the syllable rime (review in Blevins 1995), possibly with limited onset-sensitivity (Kelly 2004, Gordon 2005, Ryan 2013). I present experimental data which challenge this view, and support a recently proposed non-syllable-based alternative according to which weig
APA, Harvard, Vancouver, ISO, and other styles
18

Atiyah, Enas Hameed Nidah. "Stress in the Words of Home Tools in The Quranic Text: Doors And Ceiling As -A Sample-." Researcher Journal For Islamic Sciences 1, no. 2 (2022). http://dx.doi.org/10.37940/rjis.2022.1.2.8.

Full text
Abstract:
When the right of our early scholars confiscated and accusations attributed to them for not knowing some of the nature of the words, hence as researchers we must clarify and give evidences to decline these claims, so these facts become obvious through careful investigation of various resources. Therefore, the present applied study is conducted to give raise to a theory that some modernizers tried to dismantle from the ancient feild of study, which is the theory of (An-Nabr “stress”). This theory based on the principle of increasing air-pressure and focusing on the syllable, which is one of the
APA, Harvard, Vancouver, ISO, and other styles
19

Atiyah, Enas Hameed Nidah. "Stress in the Words of Home Tools in The Quranic Text: Doors And Ceiling As -A Sample-." Researcher Journal For Islamic Sciences 1, no. 2 (2022). http://dx.doi.org/10.37940/rjis.2022.1.2.8.

Full text
Abstract:
When the right of our early scholars confiscated and accusations attributed to them for not knowing some of the nature of the words, hence as researchers we must clarify and give evidences to decline these claims, so these facts become obvious through careful investigation of various resources. Therefore, the present applied study is conducted to give raise to a theory that some modernizers tried to dismantle from the ancient feild of study, which is the theory of (An-Nabr “stress”). This theory based on the principle of increasing air-pressure and focusing on the syllable, which is one of the
APA, Harvard, Vancouver, ISO, and other styles
20

Nawas, K. Khadar, A. Shahina, Keshav Balachandar, et al. "Recurrence plot embeddings as short segment nonlinear features for multimodal speaker identification using air, bone and throat microphones." Scientific Reports 14, no. 1 (2024). http://dx.doi.org/10.1038/s41598-024-62406-3.

Full text
Abstract:
AbstractSpeech is produced by a nonlinear, dynamical Vocal Tract (VT) system, and is transmitted through multiple (air, bone and skin conduction) modes, as captured by the air, bone and throat microphones respectively. Speaker specific characteristics that capture this nonlinearity are rarely used as stand-alone features for speaker modeling, and at best have been used in tandem with well known linear spectral features to produce tangible results. This paper proposes Recurrent Plot (RP) embeddings as stand-alone, non-linear speaker-discriminating features. Two datasets, the continuous multimod
APA, Harvard, Vancouver, ISO, and other styles
We offer discounts on all premium plans for authors whose works are included in thematic literature selections. Contact us to get a unique promo code!