Conceptio
›
speech-recognition
Topic
speech-recognition
Knowledge-graph topic
· documents ABOUT speech-recognition across the archive
22
Documents about speech-recognition
Documents about speech-recognition
Automatic Speech Recognition and Large Language Models for Multilingual Pathology Report Generation: Proof-of-Concept Study.
#182361
NCBI PubMed Central
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
#216862
arXiv CS
Dictionary-Augmented Large Language Model Postprocessing for Bilingual Code-Switched Medical Speech Recognition: Development and Evaluation Study.
#354669
NCBI PubMed Central
Visual Speech Recognition for Languages with Limited Labeled Data using Automatic Labels from Whisper
#430499
KOASAS: KAIST Open Access Self-Archiving System
TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models
#821230
arXiv (All)
Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
#927434
arXiv (All)
Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models
#1013590
arXiv (All)
Realtime-Venus: A full-duplex interaction system with asynchronous delegation
#1036766
arXiv (All)
Cochlear Implant Recipients: Comprehensive Longitudinal Evaluation
#888815
ClinicalTrials.gov
WhisperPipe: Source Code and Implementation for Real-Time ASR
#30190
DataCite
WhisperPipe: Source Code and Implementation for Real-Time ASR
#30191
DataCite
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
#131330
arXiv (OAI)
Voice-Based Structured Nursing Documentation Using Automatic Speech Recognition and Large Language Models: Development and Evaluation Study.
#265310
NCBI PubMed Central
Automatic Speech Recognition and Acoustic Analysis for Dysarthria Assessment in Telerehabilitation: User-Centered Design and Usability Study.
#340478
NCBI PubMed Central
Automatic speech recognition for Telugu: a comparative analysis of Wav2Vec 2.0 model variants and hyperparameter tuning.
#354589
NCBI PubMed Central
Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis
#789246
arXiv (OAI Expanded)
Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis
#793250
arXiv (All)
DuoGesture: Motion-Grounded Semantic Conditioning and Biomechanical Beat Priors for Co-Speech Gesture Generation
#920346
arXiv (OAI Expanded)
DuoGesture: Motion-Grounded Semantic Conditioning and Biomechanical Beat Priors for Co-Speech Gesture Generation
#924578
arXiv (All)
Positive Sample Propagation along the Audio-Visual Event Line
#972833
arXiv (All)
THE CONCEPT OF ARTIFICIAL INTELLIGENCE AND PHILOSOPHICAL ANALYSIS OF ITS ROLE IN SOCIETY
#26013
Zenodo (CERN)
GCDance: Genre-Controlled Music-Driven 3D Full Body Dance Generation
#153031
arXiv (OAI)
A Machine Learning Approach to Voice-Based Parkinson Disease Screening Using Multiview Spectrogram and Speech Recognition Features: Diagnostic Study.
#275595
NCBI PubMed Central
Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
#611395
arXiv (OAI Expanded)
Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
#612872
arXiv (All)
MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening
#1012138
arXiv (All)
The effects of ASR-based practice on learners’ perception of L2 sounds
#274586
Iowa State University Digital Repository
Assessing the Effectiveness of Automatic Speech Recognition Technology in Emergency Medicine Settings: a Comparative Study of Four AI-Powered Engines.
#630933
NCBI PubMed Central
POLARIS: Training-Free Audio Fingerprinting with Saliency-Based Landmarks and Delaunay Grouping
#1011235
arXiv (All)
T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition
#1011888
arXiv (All)
Two’s company, three’s a crowd: speech recognition with competing talkers: normally-hearing, cochlear implant and CI simulation subjects
#182796
Southampton
AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
#618877
arXiv (OAI Expanded)
AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
#620556
arXiv (All)
Effect of Donepezil on Speech Recognition in Cochlear Implant Users
#837645
ClinicalTrials.gov
Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
#921087
arXiv (OAI Expanded)
Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
#925323
arXiv (All)
Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
#1033859
arXiv (All)
Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
#1036074
arXiv (All)
A Multi-task Learning Balanced Attention Convolutional Neural Network Model for Few-shot Underwater Acoustic Target Recognition
#128371
arXiv (OAI)
Resolving Phonological Uncertainty in Reading: Orthography-to-Phonology Computations in Hebrew
#434997
OSF
Efficacy of Digital Noise Reduction Strategies: A Hearing Aid Trial
#807466
ClinicalTrials.gov
Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
#1012340
arXiv (All)
A comparison of the effects of soundfield amplification on acoustical characteristics and word recognition performance in relocatable and permanent classrooms
#957779
Louisiana Tech Digital Commons
RTP Payload Formats for European Telecommunications Standards Institute (ETSI) European Standard ES 202 050, ES 202 211, and ES 202 212 Distributed Speech Recognition Encoding
#500468
IETF RFCs
Multifaceted engagement in social interaction with a machine: the JOKER project
#955807
Koc University Digital Collections
ARTPARK-IISc/Vaani
#786975
Hugging Face Datasets
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
#6500
OpenAlex
The Effect of Deep Neural Network Implementation on Speech Recognition, Listening Effort, and Sound Quality in Older Adults With Mild to Moderately Severe Hearing Loss.
#195995
NCBI PubMed Central
Results of a survey of Special Educational Needs Coordinators (SENCOs) into their pupils' use of speech recognition software in mainstream secondary schools, to support significant writing difficulties.
#207820
University of Reading Research Data Archive
Intelligent operating architecture for audio-visual breast self-examination multimedia training system
#210409
Animo Repository
← Previous
Page 4 of 5
Next →
Topic record
· derived from the Conceptio knowledge graph (shared subject terms across the corpus)
Conceptio Open Knowledge Archive — topic hubs link to canonical document pages with full provenance.