Kliknij ten link, aby zobaczyć inne rodzaje publikacji na ten temat: Automatic speech recognition – Statistical methods.

Rozprawy doktorskie na temat „Automatic speech recognition – Statistical methods”

Utwórz poprawne odniesienie w stylach APA, MLA, Chicago, Harvard i wielu innych

Wybierz rodzaj źródła:

Sprawdź 36 najlepszych rozpraw doktorskich naukowych na temat „Automatic speech recognition – Statistical methods”.

Przycisk „Dodaj do bibliografii” jest dostępny obok każdej pracy w bibliografii. Użyj go – a my automatycznie utworzymy odniesienie bibliograficzne do wybranej pracy w stylu cytowania, którego potrzebujesz: APA, MLA, Harvard, Chicago, Vancouver itp.

Możesz również pobrać pełny tekst publikacji naukowej w formacie „.pdf” i przeczytać adnotację do pracy online, jeśli odpowiednie parametry są dostępne w metadanych.

Przeglądaj rozprawy doktorskie z różnych dziedzin i twórz odpowiednie bibliografie.

1

Wu, Jian, and 武健. "Discriminative speaker adaptation and environmental robustness in automatic speech recognition." Thesis, The University of Hong Kong (Pokfulam, Hong Kong), 2004. http://hub.hku.hk/bib/B31246138.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
2

黃伯光 and Pak-kwong Wong. "Statistical language models for Chinese recognition: speech and character." Thesis, The University of Hong Kong (Pokfulam, Hong Kong), 1998. http://hub.hku.hk/bib/B31239456.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
3

Chan, Oscar. "Prosodic features for a maximum entropy language model." University of Western Australia. School of Electrical, Electronic and Computer Engineering, 2008. http://theses.library.uwa.edu.au/adt-WU2008.0244.

Pełny tekst źródła
Streszczenie:
A statistical language model attempts to characterise the patterns present in a natural language as a probability distribution defined over word sequences. Typically, they are trained using word co-occurrence statistics from a large sample of text. In some language modelling applications, such as automatic speech recognition (ASR), the availability of acoustic data provides an additional source of knowledge. This contains, amongst other things, the melodic and rhythmic aspects of speech referred to as prosody. Although prosody has been found to be an important factor in human speech recognitio
Style APA, Harvard, Vancouver, ISO itp.
4

Fu, Qiang. "A generalization of the minimum classification error (MCE) training method for speech recognition and detection." Diss., Georgia Institute of Technology, 2008. http://hdl.handle.net/1853/22705.

Pełny tekst źródła
Streszczenie:
The model training algorithm is a critical component in the statistical pattern recognition approaches which are based on the Bayes decision theory. Conventional applications of the Bayes decision theory usually assume uniform error cost and result in a ubiquitous use of the maximum a posteriori (MAP) decision policy and the paradigm of distribution estimation as practice in the design of a statistical pattern recognition system. The minimum classification error (MCE) training method is proposed to overcome some substantial limitations for the conventional distribution estimation methods. In t
Style APA, Harvard, Vancouver, ISO itp.
5

Seward, Alexander. "Efficient Methods for Automatic Speech Recognition." Doctoral thesis, KTH, Tal, musik och hörsel, 2003. http://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-3675.

Pełny tekst źródła
Streszczenie:
This thesis presents work in the area of automatic speech recognition (ASR). The thesis focuses on methods for increasing the efficiency of speech recognition systems and on techniques for efficient representation of different types of knowledge in the decoding process. In this work, several decoding algorithms and recognition systems have been developed, aimed at various recognition tasks. The thesis presents the KTH large vocabulary speech recognition system. The system was developed for online (live) recognition with large vocabularies and complex language models. The system utilizes weight
Style APA, Harvard, Vancouver, ISO itp.
6

Clarkson, P. R. "Adaptation of statistical language models for automatic speech recognition." Thesis, University of Cambridge, 1999. http://ethos.bl.uk/OrderDetails.do?uin=uk.bl.ethos.597745.

Pełny tekst źródła
Streszczenie:
Statistical language models encode linguistic information in such a way as to be useful to systems which process human language. Such systems include those for optical character recognition and machine translation. Currently, however, the most common application of language modelling is in automatic speech recognition, and it is this that forms the focus of this thesis. Most current speech recognition systems are dedicated to one specific task (for example, the recognition of broadcast news), and thus use a language model which has been trained on text which is appropriate to that task. If, ho
Style APA, Harvard, Vancouver, ISO itp.
7

Wei, Yi. "Statistical methods on automatic aircraft recognition in aerial images." Thesis, University of Strathclyde, 2002. http://ethos.bl.uk/OrderDetails.do?uin=uk.bl.ethos.248947.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
8

Wong, Pak-kwong. "Statistical language models for Chinese recognition : speech and character /." Hong Kong : University of Hong Kong, 1998. http://sunzi.lib.hku.hk/hkuto/record.jsp?B20158725.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
9

McGreevy, Michael. "Statistical language modelling for large vocabulary speech recognition." Thesis, Queensland University of Technology, 2006. https://eprints.qut.edu.au/16444/1/Michael_McGreevy_Thesis.pdf.

Pełny tekst źródła
Streszczenie:
The move towards larger vocabulary Automatic Speech Recognition (ASR) systems places greater demands on language models. In a large vocabulary system, acoustic confusion is greater, thus there is more reliance placed on the language model for disambiguation. In addition to this, ASR systems are increasingly being deployed in situations where the speaker is not conscious of their interaction with the system, such as in recorded meetings and surveillance scenarios. This results in more natural speech, which contains many false starts and disfluencies. In this thesis we investigate a novel ap
Style APA, Harvard, Vancouver, ISO itp.
10

McGreevy, Michael. "Statistical language modelling for large vocabulary speech recognition." Queensland University of Technology, 2006. http://eprints.qut.edu.au/16444/.

Pełny tekst źródła
Streszczenie:
The move towards larger vocabulary Automatic Speech Recognition (ASR) systems places greater demands on language models. In a large vocabulary system, acoustic confusion is greater, thus there is more reliance placed on the language model for disambiguation. In addition to this, ASR systems are increasingly being deployed in situations where the speaker is not conscious of their interaction with the system, such as in recorded meetings and surveillance scenarios. This results in more natural speech, which contains many false starts and disfluencies. In this thesis we investigate a novel approa
Style APA, Harvard, Vancouver, ISO itp.
11

Doulaty, Bashkand Mortaza. "Methods for addressing data diversity in automatic speech recognition." Thesis, University of Sheffield, 2017. http://etheses.whiterose.ac.uk/17096/.

Pełny tekst źródła
Streszczenie:
The performance of speech recognition systems is known to degrade in mismatched conditions, where the acoustic environment and the speaker population significantly differ between the training and target test data. Performance degradation due to the mismatch is widely reported in the literature, particularly for diverse datasets. This thesis approaches the mismatch problem in diverse datasets with various strategies including data refinement, variability modelling and speech recognition model adaptation. These strategies are realised in six novel contributions. The first contribution is a data
Style APA, Harvard, Vancouver, ISO itp.
12

Gayvert, Robert T. "A statistical approach to formant tracking /." Online version of thesis, 1988. http://hdl.handle.net/1850/10499.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
13

Salvi, Giampiero. "Mining Speech Sounds : Machine Learning Methods for Automatic Speech Recognition and Analysis." Doctoral thesis, Stockholm : KTH School of Computer Science and Comunication, 2006. http://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-4111.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
14

Whittaker, Edward William Daniel. "Statistical language modelling for automatic speech recognition of Russian and English." Thesis, University of Cambridge, 2000. http://ethos.bl.uk/OrderDetails.do?uin=uk.bl.ethos.621936.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
15

Singh-Miller, Natasha 1981. "Neighborhood analysis methods in acoustic modeling for automatic speech recognition." Thesis, Massachusetts Institute of Technology, 2010. http://hdl.handle.net/1721.1/62450.

Pełny tekst źródła
Streszczenie:
Thesis (Ph. D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2010.<br>Cataloged from PDF version of thesis.<br>Includes bibliographical references (p. 121-134).<br>This thesis investigates the problem of using nearest-neighbor based non-parametric methods for performing multi-class class-conditional probability estimation. The methods developed are applied to the problem of acoustic modeling for speech recognition. Neighborhood components analysis (NCA) (Goldberger et al. [2005]) serves as the departure point for this study. NCA is a non-paramet
Style APA, Harvard, Vancouver, ISO itp.
16

Delmege, James W. "CLASS : a study of methods for coarse phonetic classification /." Online version of thesis, 1988. http://hdl.handle.net/1850/10449.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
17

Dambreville, Samuel. "Statistical and geometric methods for shape-driven segmentation and tracking." Diss., Atlanta, Ga. : Georgia Institute of Technology, 2008. http://hdl.handle.net/1853/22707.

Pełny tekst źródła
Streszczenie:
Thesis (Ph. D.)--Electrical and Computer Engineering, Georgia Institute of Technology, 2008.<br>Committee Chair: Allen Tannenbaum; Committee Member: Anthony Yezzi; Committee Member: Marc Niethammer; Committee Member: Patricio Vela; Committee Member: Yucel Altunbasak.
Style APA, Harvard, Vancouver, ISO itp.
18

Ravindran, Sourabh. "Physiologically Motivated Methods For Audio Pattern Classification." Diss., Georgia Institute of Technology, 2006. http://hdl.handle.net/1853/14066.

Pełny tekst źródła
Streszczenie:
Human-like performance by machines in tasks of speech and audio processing has remained an elusive goal. In an attempt to bridge the gap in performance between humans and machines there has been an increased effort to study and model physiological processes. However, the widespread use of biologically inspired features proposed in the past has been hampered mainly by either the lack of robustness across a range of signal-to-noise ratios or the formidable computational costs. In physiological systems, sensor processing occurs in several stages. It is likely the case that signal features and bio
Style APA, Harvard, Vancouver, ISO itp.
19

Melin, Håkan. "Automatic speaker verification on site and by telephone: methods, applications and assessment." Doctoral thesis, KTH, Tal, musik och hörsel, TMH, 2006. http://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-4242.

Pełny tekst źródła
Streszczenie:
Speaker verification is the biometric task of authenticating a claimed identity by means of analyzing a spoken sample of the claimant's voice. The present thesis deals with various topics related to automatic speaker verification (ASV) in the context of its commercial applications, characterized by co-operative users, user-friendly interfaces, and requirements for small amounts of enrollment and test data. A text-dependent system based on hidden Markov models (HMM) was developed and used to conduct experiments, including a comparison between visual and aural strategies for prompting claimants
Style APA, Harvard, Vancouver, ISO itp.
20

Yaman, Sibel. "A multi-objective programming perspective to statistical learning problems." Diss., Atlanta, Ga. : Georgia Institute of Technology, 2008. http://hdl.handle.net/1853/26470.

Pełny tekst źródła
Streszczenie:
Thesis (Ph.D)--Electrical and Computer Engineering, Georgia Institute of Technology, 2009.<br>Committee Chair: Chin-Hui Lee; Committee Member: Anthony Yezzi; Committee Member: Evans Harrell; Committee Member: Fred Juang; Committee Member: James H. McClellan. Part of the SMARTech Electronic Thesis and Dissertation Collection.
Style APA, Harvard, Vancouver, ISO itp.
21

Berry, Jeffrey James. "Machine Learning Methods for Articulatory Data." Diss., The University of Arizona, 2012. http://hdl.handle.net/10150/223348.

Pełny tekst źródła
Streszczenie:
Humans make use of more than just the audio signal to perceive speech. Behavioral and neurological research has shown that a person's knowledge of how speech is produced influences what is perceived. With methods for collecting articulatory data becoming more ubiquitous, methods for extracting useful information are needed to make this data useful to speech scientists, and for speech technology applications. This dissertation presents feature extraction methods for ultrasound images of the tongue and for data collected with an Electro-Magnetic Articulograph (EMA). The usefulness of these featu
Style APA, Harvard, Vancouver, ISO itp.
22

Khodai-Joopari, Mehrdad Information Technology &amp Electrical Engineering Australian Defence Force Academy UNSW. "Forensic speaker analysis and identification by computer : a Bayesian approach anchored in the cepstral domain." Awarded by:University of New South Wales - Australian Defence Force Academy. School of Information Technology and Electrical Engineering, 2007. http://handle.unsw.edu.au/1959.4/38715.

Pełny tekst źródła
Streszczenie:
This thesis advances understanding of the forensic value of the automatic speech parameters by addressing the following question: what is the potentiality of the speech cepstrum as a forensic-acoustic parameter? Despite many advances in automatic speech and speaker recognition, robust and unconstrained progress in technical forensic speaker identification has been partly impeded by our incomplete understanding of the interaction and relation between forensic phonetics and the techniques employed in state-of-the-art automatic speech and speaker recognition. The posed question underlies the rec
Style APA, Harvard, Vancouver, ISO itp.
23

Yoshino, Koichiro. "Spoken Dialogue System for Information Navigation based on Statistical Learning of Semantic and Dialogue Structure." 京都大学 (Kyoto University), 2014. http://hdl.handle.net/2433/192214.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
24

Zamora, Martínez Francisco Julián. "Aportaciones al modelado conexionista de lenguaje y su aplicación al reconocimiento de secuencias y traducción automática." Doctoral thesis, Universitat Politècnica de València, 2012. http://hdl.handle.net/10251/18066.

Pełny tekst źródła
Streszczenie:
El procesamiento del lenguaje natural es un área de aplicación de la inteligencia artificial, en particular, del reconocimiento de formas que estudia, entre otras cosas, incorporar información sintáctica (modelo de lenguaje) sobre cómo deben juntarse las palabras de una determinada lengua, para así permitir a los sistemas de reconocimiento/traducción decidir cual es la mejor hipótesis �con sentido común�. Es un área muy amplia, y este trabajo se centra únicamente en la parte relacionada con el modelado de lenguaje y su aplicación a diversas tareas: reconocimiento de secuencias mediante modelos
Style APA, Harvard, Vancouver, ISO itp.
25

Le, Hai Son. "Continuous space models with neural networks in natural language processing." Phd thesis, Université Paris Sud - Paris XI, 2012. http://tel.archives-ouvertes.fr/tel-00776704.

Pełny tekst źródła
Streszczenie:
The purpose of language models is in general to capture and to model regularities of language, thereby capturing morphological, syntactical and distributional properties of word sequences in a given language. They play an important role in many successful applications of Natural Language Processing, such as Automatic Speech Recognition, Machine Translation and Information Extraction. The most successful approaches to date are based on n-gram assumption and the adjustment of statistics from the training data by applying smoothing and back-off techniques, notably Kneser-Ney technique, introduced
Style APA, Harvard, Vancouver, ISO itp.
26

Bezůšek, Marek. "Objektivizace Testu 3F - dysartrický profil pomocí akustické analýzy." Master's thesis, Vysoké učení technické v Brně. Fakulta elektrotechniky a komunikačních technologií, 2021. http://www.nusl.cz/ntk/nusl-442568.

Pełny tekst źródła
Streszczenie:
Test 3F is used to diagnose the extent of motor speech disorder – dysarthria for czech speakers. The evaluation of dysarthric speech is distorted by subjective assessment. The motivation behind this thesis is that there are not many automatic and objective analysis tools that can be used to evaluate phonation, articulation, prosody and respiration of speech disorder. The aim of this diploma thesis is to identify, implement and test acoustic features of speech that could be used to objectify and automate the evaluation. These features should be easily interpretable by the clinician. It is assum
Style APA, Harvard, Vancouver, ISO itp.
27

Sodré, Bruno Ribeiro. "Reconhecimento de padrões aplicados à identificação de patologias de laringe." Universidade Tecnológica Federal do Paraná, 2016. http://repositorio.utfpr.edu.br/jspui/handle/1/2013.

Pełny tekst źródła
Streszczenie:
As patologias que afetam a laringe estão aumentando consideravelmente nos últimos anos devido à condição da sociedade atual onde há hábitos não saudáveis como fumo, álcool e tabaco e um abuso vocal cada vez maior, talvez por conta do aumento da poluição sonora, principalmente nos grandes centros urbanos. Atualmente o exame utilizado pela endoscopia per-oral, direcionado a identificar patologias de laringe, são a videolaringoscopia e videoestroboscopia, ambos invasivos e por muitas vezes desconfortável ao paciente. Buscando melhorar o bem estar e minimizar o desconforto dos pacientes que necessi
Style APA, Harvard, Vancouver, ISO itp.
28

Goecke, Roland. "A stereo vision lip tracking algorithm and subsequent statistical analyses of the audio-video correlation in Australian English." Phd thesis, 2004. http://hdl.handle.net/1885/149999.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
29

Ghane, Parisa. "Silent speech recognition in EEG-based brain computer interface." Thesis, 2015. http://hdl.handle.net/1805/9886.

Pełny tekst źródła
Streszczenie:
Indiana University-Purdue University Indianapolis (IUPUI)<br>A Brain Computer Interface (BCI) is a hardware and software system that establishes direct communication between human brain and the environment. In a BCI system, brain messages pass through wires and external computers instead of the normal pathway of nerves and muscles. General work ow in all BCIs is to measure brain activities, process and then convert them into an output readable for a computer. The measurement of electrical activities in different parts of the brain is called electroencephalography (EEG). There are lots of sens
Style APA, Harvard, Vancouver, ISO itp.
30

May, Avner. "Kernel Approximation Methods for Speech Recognition." Thesis, 2018. https://doi.org/10.7916/D8D80P9T.

Pełny tekst źródła
Streszczenie:
Over the past five years or so, deep learning methods have dramatically improved the state of the art performance in a variety of domains, including speech recognition, computer vision, and natural language processing. Importantly, however, they suffer from a number of drawbacks: 1. Training these models is a non-convex optimization problem, and thus it is difficult to guarantee that a trained model minimizes the desired loss function. 2. These models are difficult to interpret. In particular, it is difficult to explain, for a given model, why the computations it performs make accu
Style APA, Harvard, Vancouver, ISO itp.
31

Chandrasekaran, Aravind. "Efficient methods for rapid UBM training (RUT) for robust speaker verification /." 2008. http://proquest.umi.com/pqdweb?did=1650508671&sid=2&Fmt=2&clientId=10361&RQT=309&VName=PQD.

Pełny tekst źródła
Style APA, Harvard, Vancouver, ISO itp.
32

Wang, Qi. "Nonlinear noise compensation in feature domain for speech recognition with numerical methods /." 2004. http://wwwlib.umi.com/cr/yorku/fullcit?pMQ99403.

Pełny tekst źródła
Streszczenie:
Thesis (M.Sc.)--York University, 2004. Graduate Programme in Computer Science.<br>Typescript. Includes bibliographical references (leaves 60-65). Also available on the Internet. MODE OF ACCESS via web browser by entering the following URL: http://wwwlib.umi.com/cr/yorku/fullcit?pMQ99403
Style APA, Harvard, Vancouver, ISO itp.
33

"The statistical evaluation of minutiae-based automatic fingerprint verification systems." Thesis, 2006. http://library.cuhk.edu.hk/record=b6074180.

Pełny tekst źródła
Streszczenie:
Basic technologies for fingerprint feature extraction and matching have been improved to such a stage that they can be embedded into commercial Automatic Fingerprint Verification Systems (AFVSs). However, the reliability of AFVSs has kept attracting concerns from the society since AFVSs do fail occasionally due to difficulties like problematic fingers, changing environments, and malicious attacks. Furthermore, the absence of a solid theoretical foundation for evaluating AFVSs prevents these failures from been predicted and evaluated. Under the traditional empirical AFVS evaluation framework, r
Style APA, Harvard, Vancouver, ISO itp.
34

"Robust methods for Chinese spoken document retrieval." 2003. http://library.cuhk.edu.hk/record=b5896122.

Pełny tekst źródła
Streszczenie:
Hui Pui Yu.<br>Thesis (M.Phil.)--Chinese University of Hong Kong, 2003.<br>Includes bibliographical references (leaves 158-169).<br>Abstracts in English and Chinese.<br>Abstract --- p.2<br>Acknowledgements --- p.6<br>Chapter 1 --- Introduction --- p.23<br>Chapter 1.1 --- Spoken Document Retrieval --- p.24<br>Chapter 1.2 --- The Chinese Language and Chinese Spoken Documents --- p.28<br>Chapter 1.3 --- Motivation --- p.33<br>Chapter 1.3.1 --- Assisting the User in Query Formation --- p.34<br>Chapter 1.4 --- Goals --- p.34<br>Chapter 1.5 --- Thesis Organization --- p.35<br>Chapter 2 ---
Style APA, Harvard, Vancouver, ISO itp.
35

Τσιλφίδης, Αλέξανδρος. "Signal processing methods for enhancing speech and music signals in reverberant environments." Thesis, 2011. http://nemertes.lis.upatras.gr/jspui/handle/10889/4710.

Pełny tekst źródła
Streszczenie:
This thesis presents novel signal processing algorithms for speech and music dereverberation. The proposed algorithms focus on blind single-channel suppression of late reverberation; however binaural and semi-blind methods have also been introduced. Late reverberation is a particularly harmful distortion, since it significantly decreases the perceived quality of the reverberant signals but also degrades the performance of Automatic Speech Recognition (ASR) systems and other speech and music processing algorithms. Hence, the proposed deverberation methods can be either used as standalone enhanc
Style APA, Harvard, Vancouver, ISO itp.
36

Medeiros, Henrique Rodrigues Barbosa de. "Automatic detection of disfluencies in a corpus of university lectures." Master's thesis, 2014. http://hdl.handle.net/10071/8683.

Pełny tekst źródła
Streszczenie:
This dissertation focuses on the identification of disfluent sequences and their distinct structural regions. Reported experiments are based on audio segmentation and prosodic features, calculated from a corpus of university lectures in European Portuguese, containing about 32 hours of speech and about 7.7% of disfluencies. The set of features automatically extracted from the forced alignment corpus proved to be discriminant of the regions contained in the production of a disfluency. The best results concern the detection of the interregnum, followed by the detection of the interruption
Style APA, Harvard, Vancouver, ISO itp.
Oferujemy zniżki na wszystkie plany premium dla autorów, których prace zostały uwzględnione w tematycznych zestawieniach literatury. Skontaktuj się z nami, aby uzyskać unikalny kod promocyjny!