Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Cogn Psychol . Author manuscript; available in PMC: 2026 Apr 11. Published in final edited form as: Cogn Psychol. 2025 Oct 13;161:101766. doi: 10.1016/j.cogpsych.2025.101766 Search in PMC Search in PubMed View in NLM Catalog Add to search The intelligibility of consonants in American English infant-directed speech Daniel Swingley Daniel Swingley 1 University of Pennsylvania, United States of America Find articles by Daniel Swingley 1, * Author information Article notes Copyright and License information 1 University of Pennsylvania, United States of America * Correspondence to: Department of Psychology, University of Pennsylvania, 425 S. University Ave., Philadelphia PA 19104, United States of America. [email protected] . Issue date 2025 Nov. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/by-nc-nd/4.0/ ). PMC Copyright notice PMCID: PMC13066829 NIHMSID: NIHMS2163362 PMID: 41086777 The publisher's version of this article is available at Cogn Psychol Abstract To begin learning their language, infants must locate words in the speech signal. Some models of word discovery presuppose that the discovery process depends on identifying phonetic segments (phones) in speech. To test the plausibility of models arguing that infants can reliably categorize consonants in speech, adult native speakers were asked to identify the consonant in vowel-consonant-vowel sequences extracted from spontaneous English infant-directed speech. Listeners could consistently identify some instances of consonants (for example, correctly indicating that an /s/ was an /s/). But many tokens (about half) were not consistently identifiable. Performance was significantly worse for codas than onsets. Providing the full utterance context in low-pass-filtered form did not aid recognition, nor did familiarization with the talker. In a second task, listeners were barely above chance in guessing whether a consonant was a word onset or a word-final coda. Performance on infant-directed speech was not markedly better than performance on a comparison set of adult-directed speech consonants. Erroneous responses frequently had little systematic resemblance to the correct answer. The results suggest that it is not plausible that infants can parse most utterances exhaustively into strings of uttered speech sounds and feed those strings into a statistical clustering mechanism. Keywords: Language acquisition, Categorization, Speech, Phonetics, Unsupervised learning, Infant development When infants come into the world and start taking stock of what is going on around them, they will find people talking: talking to one another, and sometimes to the infant. For the baby, some aspects of spoken language are familiar, because of auditory exposure in the womb; others are new, like the visible connection between speech and the movement of the mouth, or the contingency between the talker’s speech and the infant’s behavior. It seems reasonable to suppose that infants find speech interesting and intuit that it is important. But how do they start making sense of what is happening when people talk? How do infants use their experience with parental speech to grasp how language uses sound to convey meaning? To do this, infants will need to discover words. The pathway to discovering words that has been the focus of the most empirical attention starts with perceptual categorization of the consonants and vowels of the language. Infants have perceptual biases that are partially aligned with the demands of phonetic categorization (e.g., Eimas & Miller, 1980 ; Eimas, Siqueland, Jusczyk, & Vigorito, 1971 ; Kuhl, 1985 ), and over the course of the first year, infants improve in the categorization of clear instances of the sounds of their native language, and decline in categorization of sounds of other languages (e.g., Kuhl, Stevens, Hayashi, Deguchi, Kiritani, & Iverson, 2006 ; Tsuji & Cristia, 2014 ). Early demonstrations of such developmental changes were described using the metaphor of “magnets” (i.e., attractors), the notion being that language-specific category prototypes pull the perceptual experience of speech sounds toward their category centers, making atypical instances of sounds difficult to distinguish from typical ones. These sorts of effects are viewed as automatic; for example, one does not try to hear vowels in terms of category prototypes, it just happens. The mandatory, automatic nature of this process explains the difficulty that adults and infants have in distinguishing certain foreign sound contrasts. When the sounds fall into the same native category, they are hard to tell apart ( Best, McRoberts, & Goodell, 2001 ). From these experiments, in which atypical instances of phones are perceived as though they were prototypical instances, it seems a small step to extend such results to the interpretation of speech in daily life. Adaptation toward the native language is unquestionably a product of real-life language experience. If daily language experience in all its variety can teach infants about their native language enough to cause language-specific speech-sound categorization in the lab, it seems likely that at least some of the time, infants categorize speech sounds outside the lab too, undergoing the same pull toward prototypes that is observed in experiments. So, perhaps we should think of infants hearing speech as converting the continuous phonetic signal into a sequence of consonants and vowels, where these sound categories are aligned, in large part, with the phonological categories they are learning. If this is true, then the problem of discovering words in speech becomes the problem of assembling phonological category labels into words, or of drawing word boundaries between these labels. Thus, when researchers discuss the problem of speech segmentation, they often start from the point where this analog-to-digital (or phonetic-to-phonological) conversion has already been done. The problem is frequently characterized informally using text of the form “therearenospacesbetweenwords.” with the letters being given, and the boundaries between them left to discover. Statistical models of the segmentation problem in the psycholinguistics literature almost always take this perspective (e.g., Batchelder, 2002 ; Brent & Cartwright, 1996 ; Cabiddu, Bott, & Jones, 2023 ; Cairns, Shillcock, Chater, & Levy, 1997 ; Daland & Pierrehumbert, 2011 ; Perruchet & Vinter, 1998 ; Swingley, 1999 ). A corpus of infant-directed speech is converted into a set of estimated phonological strings, and these strings are submitted to an algorithm or neural network that implements statistical heuristics sensitive to the cohesiveness or repetition structure of sub-sequences within the utterances (e.g., Bernard et al., 2020 ). For some models, the inputs are sequences of phones, and for other models, the inputs are sequences of syllables; in both cases, the starting point is a phonological transcription. The outcome of the model is an estimate of the child’s “protolexicon” ( Swingley, 2005 ), a set of phonological strings comprising words, parts of words, and portmanteaux in some proportion. The proposal is that over time, the correct segmentations will become meaningful words in the lexicon, and the missegmentations will be recognized as errors and discarded. Experimental studies of infants reveal that by 8 months of age, children can detect the repetition of frequently occurring syllables or syllable sequences, even with just a minute or two of exposure (e.g., Jusczyk & Aslin, 1995 ). This discovery process relies on the elevated frequency of the repeating speech chunks, but also relies on statistical properties of recent speech experience. Frequent chunks that are cohesive, because their components tend to occur together, are more likely to be remembered by infants than equally frequent chunks that are not so cohesive ( Aslin, Saffran, & Newport, 1998 ; Pelucchi, Hay, & Saffran, 2009a , 2009b ). This phenomenon is consistent with, and indeed motivated, most of the computational models of word segmentation cited above; if infants are counting relative frequencies of something, should it not be phones or syllables that they count, given what appears to be evidence that they automatically convert speech (analog) into categories (digital, countable)? Although narratives along these lines have been dominant in the developmental psycholinguistics literature for some time, their presuppositions about the learner and about the speech environment are debatable. As noted, they assume that learners break the flow of speech into discrete consonant and vowel categories, label the categories, and compute statistics (like transitional probabilities) over this sequence. But in fact, infants might not start out creating categories at this temporal scale. They might retain larger, looser units (e.g., Beckman & Edwards, 2000 ; Nittrouer, 2021 ) or smaller ones ( Schatz, Feldman, Goldwater, Cao, & Dupoux, 2021 ). Or they might start interpreting speech at multiple scales and along multiple dimensions all at once (e.g., Vihman & Majorano, 2017 ). Recent computational work has shown that the results of some of the infant experiments supporting the supposition that infants compute transitional probabilities over phones or syllables are also consistent with models in which discrete phone or syllable categories are not explicitly represented at all ( Swingley & Algayres, 2024 ). Apart from whatever infants are disposed to compute, the traditional account also assumes that the speech to infants is, at least to a first approximation, articulated clearly enough for the consonants and vowels to be correctly identified. Computing transitional probabilities over speech segments, or over syllables defined by those segments, is only possible if speech segments exist in the signal and can be categorized. This latter assumption is tested in the present research. There are prior indications in the literature that in fact infant-directed speech is not especially clear, phonetically. For example, Bard and Anderson (1983) recorded conversations between parents and an experimenter, and between parents and their one- and two-year-olds. Ten words from the conversation of each of the 24 dyads were selected at random, with no type repetitions, and excised acoustically from their constituent sentences. These clips were played to adult listeners for identification. Child-directed words were found to be significantly less identifiable than adult-directed words, and were correctly identified only about 25% of the time (compared to about 40% for adult-directed speech). A similar effect was found in a follow-up experiment that selected matched word types across the child-directed and adult directed stimulus sets so that the addressee differences could be attributable only to the way words were said, and not which words were said. In this study, identification performance varied substantially over items, but correct identifications of child-directed words averaged about 38%, and adult-directed words about 45%. This study suggests that the well-known findings of Pollack and Pickett (e.g., Pollack & Pickett, 1963 ), in which identification of words excised from conversational content was revealed to be very error-prone, cannot be dismissed for the infant case on the grounds that parents must hyperarticulate to their children. A handful of other studies have found infant-directed and adult-directed speech to be similar in their intelligibility, with some analyses finding minor advantages to one register or the other, but the effects often being dominated by individual differences or aspects of context (e.g., Buckler & Goy, 2018 ; Dilley, Millett, McAuley, & Bergeson, 2014 ; Lahey & Ernestus, 2014 ). Phonetic (and phonological) reduction phenomena found in adult conversation are also found, with similar frequencies of occurrence, in speech to children. A related series of investigations has used phonetic measurements rather than word identification to assess whether infant-directed speech is clarified relative to adult-directed speech. These studies do not assess whether speech sounds are identifiable, but adopt a comparative approach, reasoning that if infant-directed speech is acoustically clarified relative to adult-directed speech, infants’ phonetic category learning can be viewed as a consequence of parental adaptation to the infant’s needs ( Eaves, Jr., Feldman, Griffiths, & Shafto, 2016 ) or the infant’s potential (e.g., Kalashnikova, Goswami, & Burnham, 2020 ; Lovcevic, Kalashnikova, & Burnham, 2020 ). There is some debate about whether speech sounds are usually hyperarticulated in infant-directed speech relative to adult-directed speech and what the implications would be for learning (e.g., Cristia, Seidl, Junge, Soderstrom, & Hagoort, 2014 ; Kuhl et al., 1997 ; Liu, Kuhl, & Tsao, 2003 ; Martin et al., 2015 ). Part of the debate stems from the fact that hyperarticulation effects are not large, and are therefore susceptible to many sources of variability: individual variation among parents, situational context, the child’s age or linguistic competence, the language being spoken, and the kind of speech sound being measured. The present study will speak to this question, but does not attempt to resolve it. Rather, our goal was to help evaluate how often consonants of English infant-directed speech are realized clearly enough to be identifiable to adults, when the adults do not reliably know what word they are hearing. Our procedure was simple: adult native speakers of English were presented with consonants drawn from corpora of English infant-directed speech and adult-directed speech, and were asked to identify them. The consonants were taken from vowel-consonant-vowel (VCV) environments, and presented within the context of these surrounding vowels. The VCVs were selected with the constraint that the C was always either the final sound of a word, or the initial sound of a word. So, on each trial, judges heard a VCV sound clip, and chose the consonant they had heard, selecting from a list. Judges were also asked to indicate whether they thought the consonant was word-initial or word-final. This query was included to allow us to estimate how often local phonetic cues to syllable boundaries are present. Our primary questions were as follows. How often are word onset and offset consonants identifiable by adult listeners? When consonants are mis-identified, are the errors predictable? How often are syllable (or word) boundaries identifiable from local phonetic cues? If consonants are only infrequently identifiable by adult native speakers, it is likely that they are not usually identifiable by infants either, casting doubt on the applicability of quantitative word-segmentation models that presuppose the identifiability of speech sounds. The further implications of this depend partly on whether sounds are often misidentified as phonologically similar consonants. We envisioned two possibilities: first, that identification errors would manifest most often as sounds differing from the canonical target by one phonological feature, less often by two features, and so on. If so, it would suggest that indistinct parental pronunciations might lead to lexical representations in the right phonetic “ballpark” but with only probabilistic or vague feature specifications. Another possibility is that some tokens would be identifiable (many correct judgments) whereas other tokens would be completely unidentifiable (many errors, with little commonality among different judges’ responses). This outcome might suggest that some tokens of words in infant-directed speech are simply not well suited to the task of learning words. Considering the perceptual availability of syllable boundaries, if intervocalic consonants cannot be identified as onsets or offsets, it would indicate that models starting from the syllable as the elemental unit of statistical computation should not presuppose that infants work with the right units, at least in the case of VCV word boundaries. Note that this task, explicit identification or labeling of consonant sounds and syllable positions, is different from the task of recognizing words. In speech comprehension, one does not try to label consonants. But given that many models of word discovery presuppose the categorization of phonological segments, expert explicit categorization can help us to evaluate the plausibility or potential generality of naive implicit categorization. Any model that holds that phones are identified in word recognition or word learning requires that the listener have stable categories in memory, to which new instances can be assigned. In the case of consonants, and among skilled readers, these mental categories should be approximated well by the labels we have used in our response set, e.g., “b” being the consonant in the word “bib.” In the first study, speech samples were taken from the Brent Corpus of infant-directed speech ( Brent & Siskind, 2001 ). The second study tested consonant identification when the sounds were presented in filtered sentential contexts, to test whether performance in the first study would be improved by presenting judges with additional phonetic information. The third study examined whether prior exposure to the talker would enhance performance. The fourth study explored the degree to which successful identifications might have been based on identifying the word containing the consonant. The fifth study directly compared infant-directed speech and adult- directed speech. We will conclude with some observations about the implications of the results for our understanding of how infants begin to understand their language. 1. Experiment 1: Identifying infant-directed speech consonants 1.1. Methods 1.1.1. Speech materials Extracting VCV sequences at word boundaries requires annotated speech databases that have been aligned to the word and phone level. Our infant-directed speech samples came from the Brent corpus ( Brent & Siskind, 2001 ). The Brent corpus is a publicly available corpus downloadable from the CHILDES project database ( MacWhinney, 2000 ). Recordings from three mother–child dyads were used: mothers F1 (sessions f10jan97 [10 months, 3 days], m20jan97 [10;13], and r13feb97 [11;06]), W1 (f21jun96 [10;25]), and D1 (f17jan97 [11;06]). These dyads and sessions were selected for annotation because samples of each session were found to have relatively good recording quality, freedom from noise, and a small number of talkers in the session. Annotation began by segmenting the session-length audio files into utterances according to the utterance-boundary specifications given in the Brent and Siskind transcriptions, and then estimating word and phone boundaries using the Penn Phonetics Forced Aligner ( Yuan, Liberman, et al., 2008 ). These boundaries were then corrected by phonetically trained research assistants following a protocol modeled on the conventions of the Buckeye Corpus annotation manual ( Pitt et al., 2007 ). Because the aligner worked with a dictionary of canonical pronunciations, which frequently do not correspond to actual phonetic realizations, annotators corrected the phonological sequence when words were not pronounced canonically. The resulting phonological transcription was closer to the phonetic truth than the original canonical sequence was, but was still biased by annotators’ knowledge of the language and the sentence context (as we will see, annotation without lexical knowledge would be impossible). Annotators did not mark stress, but did distinguish /ә/ and /ɨ/ from other vowels. Flaps (/ɾ/) and glottal stops (/Ɂ/) were marked, as part of the protocol, though it is probably the case that annotators under-marked them to some degree, retaining the canonical sounds. These annotations were done between 2007 and 2016. From these annotations, vowel-consonant-vowel (VCV) sequences were selected, and filtered to include only those for which the consonant was at a word boundary. The consonant could therefore either be the coda consonant of the first word, or the onset consonant of the second word. For dyad D1, all such instances in the annotated session recording were retained for evaluation in the present work. In the case of dyads F1 and W1, we took advantage of the fact that each of the vowels in the annotated sessions had been screened as part of another study about vowel identity. This screening consisted of the present author listening to each vowel with 60 ms of phonetic context on either side. The goal of the screening was to exclude vowels that were too short, were too noisy, or (in the limit) did not even sound like human speech, because it would be unreasonable to ask listeners to judge those vowels’ identity. Note that vowels were not marked as such simply for being hard to identify; they had to be difficult to contemplate even as being vowels. Here, any VCV containing a vowel so marked was excluded. This excluded set amounted to 183 of 1100 potential VCVs from dyad F1 (16.6%), and 156 of 654 from dyad W1 (23.9%). The rationale for applying this exclusion was to ensure that if consonants turned out to be difficult to identify, the primary reason for this could not be that the context was especially uninterpretable or the recording quality especially poor. Table 1 indicates the number of VCV tokens that were evaluated from each dyad and for each syllable position. The greater representation of onset consonants is a reflection of the fact that onsets are more common than codas in English, including at word boundaries. Note that word position refers to the canonical position, as defined by the word boundary; phonetically, some codas may have been pronounced as onsets. Table 1. Study 1 (Brent corpus) tokens by dyad and word position. Counts Proportion dyad d1 f1 w1 d1 f1 w1 onset 201 657 351 .73 .72 .70 coda 73 260 147 .27 .28 .30 Open in a new tab To produce the auditory materials, the VCV sequences were excised from their context using a script calling the unix utility sox . Each raw soundfile’s amplitude was normalized using sox’s norm function, but not otherwise processed. 1689 consonant tokens were evaluated. The distribution of sounds according to their transcribed consonant is given in the Supporting Online Materials . Their median duration was 290 ms (range, 59–1766; first quartile 219, third quartile 400). 1.1.2. Procedures Consonant identifications and word-position judgments were made by paid judges recruited through the online service Prolific. To keep the task within reasonable duration limits, the tokens were divided into sets of 235 or 236 (F1, W1 sets) or 137 or 138 (D1 sets). Sets were created in such a way that the number of tokens of each transcribed consonant in each word position was balanced across sets, where each set reflected the consonants’ distribution in the full sample. Participants judged only one set. The first cohort of participants was tested on speech samples from dyads F1 and W1, and a second cohort was tested on speech samples from dyad D1. Data collection took place between June, 2020 and April, 2021. To qualify for the study, participants had to assert that they were native English speakers. To help evaluate whether they had been accurate about this, they were given a brief multiple choice test that included 8 items about English vocabulary and grammar. Only participants scoring at 4 items or better were included in the study. This was a low bar, but raising the cutoff did not improve average judgment performance, and lowering it reduced judgment performance, with performance operationalized as the mean proportion of judgments matching the canonical consonant. As a further precaution against including responses from inattentive judges, participants were dropped from the sample if their level of matching the canonical consonant fell below 20%. Of the 89 participants tested, 13 failed this criterion, and 11 failed the English test; most of these were the same people, leaving a final sample of 72 judges. At the start of the study, participants were given a brief description and training session concerning the American English flap consonant (the alveolar tap). The training was based on a set of clear examples, and a follow-up test that required them to differentiate flaps from medial /t/ and /d/, using recorded words like tradition and bitter . To carry on with the study, participants were required to answer 6 of 8 judgments of flaps correctly; if they answered 5 or fewer correctly, they were presented with the 8 judgments again, until they succeeded at 6 or more. Details of these procedures are given in the Supporting Online Materials . Judges were also given five practice trials using VCV clips taken from the Buckeye corpus ( Pitt et al. 2007 ). These were easy items, for which many participants identified all five items correctly. Judges repeated any trial they got wrong until they landed on the correct answer. These items were used with each participant group as a way of ensuring that there were no gross differences in the abilities of the judges from one experiment to another. The main portion of the task consisted of a series of trials on which judges saw a warning message, then heard the VCV recording, and were shown the text, “What sound did you hear? (Click on the letter)” with an array of letter-word pairs, like “b (bib)” and “d (did),” exemplifying all of the English consonants. For most consonants, the written examples on the screen showed the sound both at the start and end of the word. Unvoiced and voiced “th” were indicated as “th1, (sympathy)” and “th2 (weather);” flap was indicated as “flap (butter).” Once a choice of consonant was made, the VCV was played again, this time with instructions for choosing the onset or coda (“Was the sound you heard at the beginning or end of the word?” with the options being “beginning (toe deer)” or “end (toad ear).” Judges were not put under time pressure. 1.2. Results: Brent corpus 1.2.1. Consonant identification Our first analyses concern the identifiability of the consonants. Fig. 1 shows the mean proportion of each possible response for each of the target consonants. For this plot, the target is defined according to the canonical pronunciation within its host word; thus, for the column showing onset /t/ (right panel), the plot shows how often a sound that “should” be a /t/ was heard that way, and how often participants thought they heard another consonant. Tiles on the diagonal are the cases where judges accurately identified the sound. Among the onset consonants, in every case the modal response was the correct one. Among the coda consonants, this was often the case, with a few exceptions. However, many tokens were not identified reliably. Considering the correct judgments, and weighting each consonant equally, the mean accuracy over coda consonants was 28.8% (sd 15.2, median 30.8), and over onset consonants, 55.4% (sd 13.7, median 59.2). Similar accuracy averages emerged when comparing listener judgments to the transcribed consonants rather than the canonical consonants: codas, mean 30.3% (sd 17.6, median 29.8); onsets, mean 52.6% (sd 16.2, median 56.7). Fig. 1. Open in a new tab Confusion matrix for consonants drawn from the Brent corpus. The left panel shows coda consonants, the right panel onsets. Color indicates the proportion of times that a given canonical consonant (the columns) received a given interpretation (the rows). The number in each cell also reveals this proportion, with columns summing to one. Figures in blue are estimates computed over fewer than 10 items; in magenta, 10 or more. Rows and columns are ordered so that similar phones tend to be near one another. Fig. 2 shows the mean proportion of correct responses for each transcribed consonant for each of the speakers. For codas, difficulty ranged continuously from just over zero up to almost 75%; for onsets, a few sounds were substantially worse than the others (ɾ, ð, θ, ʤ), and the remainder all fell between 50% and 75%. These differences being noted, the mean accuracy proportions for each sound in coda and onset position were significantly correlated (excluding items occurring only in one position), at r = .605, t (18) = 3.23, p = .0047; the same was true considering ranks rather than means (spearman rho = .573, p = .009). Fig. 2. Open in a new tab Proportion of responses matching the gold-standard transcription, by sound, for each speaker (green) and the average (brown solid line). The left panel shows coda consonants, the right panel onsets. Considering individual items, some consonant exemplars were easy to identify, even among many of the consonants for which overall performance was poor. Weighting each of the speakers equally, 25.6% of onset consonants were identified as their transcribed “gold standard” phone by at least 80% of listeners; 5.9% of coda consonants reached this standard. On the other hand, some consonant exemplars were extremely difficult to identify: 15.0% of onsets and 42.6% of codas were guessed correctly by fewer than 20% of judges. Fig. 3 shows the proportion of items for which each level of correct responding was achieved over a full range of thresholds, from at least 10% of listeners up to at least 90% of listeners. Fig. 3. Open in a new tab Proportions of items for which Brent-corpus judges achieved a range of levels of agreement on the transcribed consonant. The left panel shows coda consonants, the right panel onsets. Color indicates dyad (mother-child pair). The line indicates the proportion of items for which judges chose the correct transcribed consonant at least as often as the percentage given on the x axis. The wrong answers for each consonant were diverse. One way to measure this is to consider the most frequently chosen wrong answer for each canonical consonant. Among onset consonants, averaging over items to arrive at means for each phone, the most common wrong choice was picked, on average, just 21.2% of the time (median, 19.6%); among codas, 24.9% (median, 22.3%). This means that the advantage of the correct answer over the next most popular answer was very substantial for the onsets, the average advantage being 31.2 percentage points (median, 38.1). Among codas, for which correct responses were much less frequent, their advantage over the most popular wrong answer was smaller (average, 6.2 percentage points; median, 6.3). As we have seen, there were plenty of wrong answers, but on a consonant by consonant basis, they were not consistently the same wrong answers. Diversity in the most common wrong answers for a given consonant could emerge if, in general, judges who failed to identify a sound were entirely at a loss, and had to guess. But it could also emerge if the individual items were different from one another in their phonetics: perhaps some /d/s sound a little like /b/, and others sound like /t/. To address this question, we can evaluate the amount of agreement in the wrong answers to each item. For example, for an item with two judges making errors, and both guessing at random from among the 23 incorrect consonant options, we could expect them to agree only 1/23 of the time; with three making errors by guessing randomly, we would expect all three to guess the same 1/23 3 of the time and exactly two to guess the same 3 2 ∗ 1 / 23 ∗ 1 / 23 ∗ 22 / 23 . Using this logic, we can estimate for each item, given its number of incorrect responses, how often judges would be expected to agree on the most popular wrong answer if they were simply guessing. 1 Fig. 4 shows every item, displaying the number of wrong responses to it on the x -axis, and the number of judges who picked the most frequently-chosen wrong answer on the y -axis. Note that the number of wrong answers is partly a function of the number of correct answers, and partly of how many judges evaluated each item, which ranged from 5 to 12 (mean 7.7). Naturally, as more judges responded, there were more opportunities to match by chance; but for every count of wrong answers above one, the mean number of agreeing judges (black circles) exceeded the mean number expected by chance (white circles). The red lines at the foot of each panel are density plots indicating the distribution of the items according to the number of wrong answers (tokens) each stimulus received. Fig. 4. Open in a new tab For each item in the Brent corpus, the counts of the modal incorrect responses. Coda consonants are shown in the left panel, onsets in the right panel. Color indicates the proportion of correct answers for each item, with darker, more violet colors showing poorer performance. Means over items at each level of x are plotted as black circles. The expected means if wrong answers were all random guesses are plotted as white circles, and the theoretical maximum (all judges in agreement about the top wrong answer) is illustrated with white triangles. The red lines below each scatterplot show the shape of the distribution of total error counts (with arbitrary y -axis scale), independently of how diverse the errors were. Error counts were strongly skewed toward relatively fewer wrong answers for onsets, but for codas, peaking at about half wrong answers. Fig. 4 makes clear that judgments that did not match the gold-standard transcription were only rarely clear (high-agreement) instances of some other consonant. Although judges making wrong answers agreed with one another considerably more often than would be expected by guessing, fewer than half of judges agreed with the most frequent incorrect answer, even looking item by item. These results appear to refute two plausible ideas. We might imagine that when a consonant is fairly clear (few errors), any mistakes would be drawn from an especially narrow set of phonetic neighbors. But mistaken judges’ errors were diverse even when the other judges were correct. We might also imagine that when a consonant receives a majority of wrong answers, it signals a transcription error, which is then set right by our judges who agree with one another. If this were true, then in the plot we would see that as target performance gets worse (dots less yellow and more purple), agreement on the wrong answers would approach perfect agreement more (white triangles). But instead, agreement on wrong answers maintains a similar proportion throughout (black dots). What sorts of errors did judges make? If they were able to recover some of the phonological specifications of the consonants they were hearing, we would expect that many of the wrong answers would differ from the gold-standard transcription by one phonological feature in many cases, by two features in fewer cases, and so on. For example, a voiced stop might be heard as its unvoiced counterpart, or a bilabial nasal as an alveolar nasal. Figs. 5 and 6 break down, for each consonant, which errors were most frequent. The pie charts show proportions drawn over items. Each stimulus was categorized according to the phonological voice, place, and manner features that were implied by listeners’ incorrect responses more than 50% of the time. The categories shown on the plot are exhaustive, illustrating the eight (2 × 2 × 2) response patterns possible for the phonological features of an item, considering only the wrong answers. For example, if judges were presented with a given /t/ stimulus, and 40% of them responded /t/, 30% responded /p/, and 30% responded /b/, then among the errors (60% of all answers), 100% matched in manner, because all responses were stop consonants; only 30% matched the target in voicing, and only 30% matched in place. This item would be represented as falling into the “_ _ m” category, bright blue on the plot, because only the manner feature (“stop”) figured in the majority of erroneous responses. This discrete classification of each item is what lets us consider each of the items separately in the plot; for the n items testing a given consonant, the portion of the pie chart each stimulus occupies is 1/ n . 2 Fig. 5. Open in a new tab Dominant features present in the incorrect responses to onset consonants in the Brent corpus. Wrong answers to each sound were categorized according to whether they matched the correct response in voicing (v), place of articulation (p), or manner (m), for most of the judges. For each item, the features that usually matched are given in color. If only one feature matched, the pie wedge is light blue (a manner match), light green (place), or red (voicing). If two features matched, the wedge is indicated with a blend of the colors; for example, if listeners’ responses usually implied a match to voicing and place, the wedge for that item is olive green (vp_). If all three features matched but the response was still wrong, the wedge is gray; if no features matched, black. The number of items displayed for each consonant is shown at the lower left of each circle. Fig. 6. Open in a new tab Dominant features present in the incorrect responses to coda consonants in the Brent corpus. Conventions are the same as Fig. 5 . The first notable result is that sounds varied widely in the set of phonological features they conveyed when listeners failed to identify them correctly. For example, the plosives (/p,t,k,b,d,g/) were identified correctly about 61% of the time (onsets) or 19% of the time (codas). If we consider item by item how often judges’ wrong guesses were stops more than 50% of the time, this turns out to be just 36% of the items for onsets, and 41% of the items for codas. Thus, when judges were wrong in identifying a plosive, fewer than two fifths of the guesses were consistently any stop consonant. To take another example, the onsets that were successfully identified most often were the bilabial and coronal nasals, /m/ and /n/—about 78%. When judges guessed wrong, only 11% of the items consistently (more than 50%) led judges to choose any nasal when making an error. Considering the dataset generally, what this suggests is that quite often judges did not behave as though they had some independent probability of identifying each phonological feature, where errors come from one or more wrongly extracted features. Rather, they seemed to either lock in on the right answer, or guess according to criteria that are not easy to define. A second surprising result, related to the first, is that a nontrivial proportion of responses to each sound, marked in black in the plots, preserved none of the basic phonological features of the target. For example, among the onset approximants (/l,r,y,w/), about 57% of onsets and 20% of codas were correctly identified. Among the remaining tokens, 20% of the onsets and 28% of the codas did not usually elicit responses matching voicing, place, or manner of these approximants; they were often heard as unvoiced, non-approximants, with the wrong place of articulation. Broadly speaking, then, intervocalic consonants in infant-directed speech were very heterogeneous in their identifiability. Consonants at the ends of words were substantially harder to recognize. As expected based on prior results (e.g., Redford & Diehl, 1999 ), this was one of the largest effects in the study. On a phone by phone basis, wrong answers were quite variable, with the confusion matrix illustrating surprisingly little concentration in the errors produced to each sound (cf. Cutler, Weber, Smits, & Cooper, 2004 ). 1.2.2. Syllable position judgments After making consonant identifications, listeners were also asked to evaluate whether a given sound was word-final, or word-initial. Judges tended to call the consonants onsets, even when they were codas. This bias toward onsets was found for nearly all of the tested coda consonants, as shown in Fig. 7 (left panel, blue diamonds). The tendency to interpret the consonants as onsets may come about because onsets are more frequent in English, or because some intervocalic codas may have been phonetically resyllabified by the speaker as onsets. We do not have an independent assessment of phonetic resyllabification. Of course, in our dataset, when a coda is phonetically resyllabified as an onset, it creates a mismatch between the phonetic syllable boundary and the word boundary. Whether this would lead infants to make poor guesses about word boundaries is an open question (e.g., Mattys & Jusczyk, 2001 ). If we consider syllable onset detection as a signal detection problem, listener by listener the mean d-prime was 0.245 (median, 0.219; sd, 0.413), which is certainly above zero (t(66) = 4.86, p<.0001) but patently quite poor. Fig. 7. Open in a new tab Accuracy in identifying the syllable position in consonants from the Brent corpus. Each point shows results from a different VCV token, with the y axis indicating proportion correct. Points are arrayed along the x axis according to the transcribed, gold-standard consonant category, ordered according to manner of articulation. Blue diamonds indicate the mean proportion correct for each consonant, sized proportionally to the number of items contributing to that mean. Darker (reddish) points indicate items for which judges correctly identified the consonant less than half of the time; lighter (green) points, more than half of the time. The left panel shows data from (gold-standard) coda consonants; the right panel, onsets. The plot reveals that items varied greatly in how identifiable syllable position was (i.e., large spread on the y axis), and that listeners tended to report “onset” for both codas and onsets (i.e., points are mainly below 0.50 for codas, above 0.50 for onsets). Even within coda and onset items, syllable position identification was significantly more accurate when the consonant in question was easier to identify, reflected here as light green points being higher than brown ones, but this effect was not strong enough to be readily visible here. Overall, judges’ decisions were informed by the canonical syllabic position of the intervocalic consonant, in spite of this bias. A linear mixed-effects regression predicting the proportion of judges choosing “onset” for each item indicated a significant effect of the true consonantal position. To evaluate whether syllable position judgments were more accurate when the consonant was easier to identify, mean proportion of correct identifications of each item were included in the regression. Random effects were speaker (which mother) and phone (which consonant). The results are given below in Table 2 . Table 2. Regression analysis predicting selection of onset as the consonant location. Predictors included the true consonant position (coda as baseline), the proportion of judges who correctly identified the consonant for that item (meancentered), and their interaction. Random effects sd: phone, 0.017; mother, 0.026; residual, 0.173. Model R 2 0.090, marginal R 2 (fixed effects only) 0.061. β coef. Std. Err df t value p value (Intercept) 0.569 (0.019) 4.35 30.50 <0.0001 true C position is onset 0.099 (0.012) 516.7 8.22 <0.0001 % correct C i.d. (centered) −0.092 (0.030) 317.6 −3.03 =0.0026 correct % i.d. and is onset 0.168 (0.035) 582.5 4.85 <0.0001 Open in a new tab The analysis shows that judges were significantly more likely to choose “onset” when the consonant was, indeed, a word-initial consonant, though as we have observed, their accuracy in doing so was quite poor. In the baseline case where the consonant was a coda, the more identifiable sounds were significantly less likely to be judged as onsets (third row of table); in the interaction case where the consonant was an onset, more identifiable sounds were significantly more likely to be chosen as onsets. This indicates that the clearer instances of intervocalic consonants sometimes provide information about syllable or word position, and this is true for both onsets and codas. At the same time, the data do not support a presupposition that in general infants would have access to word or syllable boundaries as a consequence of (for example) allophonic variation in the characteristics of onset and coda consonants. 1.3. Discussion Adult native English speakers hearing vowel-consonant-vowel sequences extracted from English infant-directed speech were able to consistently identify around 3/5 of onset consonants and a quarter of coda consonants. Their erroneous guesses were scattered and often hardly resembled the gold-standard transcribed consonant, in terms of featural overlap. Listeners had a significant but again modest ability to differentiate word onsets from codas. If the three mothers evaluated here are representative speakers, the results suggest that over the VCV tokens in their language experience, infants cannot rely on consonants as categorically identifiable units, for example as inputs to word learning or to the calculation of conditional probabilities. However, it is also possible that our task underestimated the identifiability of consonants because broader segments of speech are required for sounds to be identifiable at normal performance levels. It is well known that the realization of speech sounds is partly a function of the surrounding speech sounds, although most demonstrations of this fact concern variability attributable to the immediately preceding or following vowel (e.g. Delattre, Liberman, & Cooper, 1955 ), and that information was provided here. Although the ability of infants to use relatively long-distance cues to phone identity is unknown, it is possible that access to local prosodic features, or some calibration to particulars of the speaker’s vocal tract, would improve consonant identification. This leads to a methodological quandary: if we were to present significantly more context in the clear, listeners would identify words and use this lexical knowledge to guide their phone judgments, just as our transcribers did in creating the gold-standard transcription. As a compromise, in Experiment 2 we presented VCV sequences embedded in low-pass-filtered versions of their original contexts. Low-pass filtering preserves many elements of the speech signal, but renders it much more difficult to understand, because higher-frequency information helps specify consonants in particular (e.g., Baer, Moore, & Kluk, 2002 ; Pollack, 1948 ; Pollack & Pickett, 1964 ). As a result, low-pass filtering has been employed in numerous psycholinguistic studies in which listeners’ sensitivity to speech features needed to be dissociated from lexical influences on interpretation (e.g., Morton & Trehub, 2001 ). Here, VCVs were presented in the clear, but with their sentential contexts low-pass filtered. 2. Experiment 2: Low-pass-filtered sentence contexts 2.1. Methods The goal of Experiment 2 was to attempt to replicate the findings of Experiment 1 while providing more extensive phonetic context. This was done by first playing to listeners the entire sentence containing the VCV clip, with all portions of the utterance except the VCV clip low-pass filtered at 400 Hz. Filtering was done using the sox utility’s sinc function. Low-pass filtering gives the speech a muffled aspect that causes speech segments (and therefore words) to be difficult to identify, but preserves the overall pitch contour and information about amplitude and speech rhythm. After listeners heard the sentence, the VCV clip was played again on its own, and listeners judged the consonant’s identity. Then, as in Experiment 1 , they heard the VCV again and were asked whether they thought the consonant was a word onset or a word offset. Listeners began the task with the five practice trials that had been used in Experiment 1 , hearing a VCV clip and responding (as before) to the consonant. After this practice period, judges were trained on the low-pass-filtered stimulus task using five utterances from the Buckeye Corpus (see Experiment 5 ). As in the subsequent trials, judges heard each clip in a low-pass-filtered context, and then the clip again, before making their judgments. Because the Buckeye corpus is not explicitly divided into utterances, the low-pass-filtered context segments for this training were the one-second portions that preceded and followed the VCV in the original recording. All other procedural features of the study were retained as in Experiment 1 , including the recruiting procedure, the training on flap consonants, and the tests of basic English knowledge, and the participant exclusion procedures. Testing was done in May–June of 2023. 2.1.1. Speech materials A subset of the same clips from the first experiment were tested, including 597 of the original 1689 VCVs. They were selected in a quasirandom fashion aiming to sample the larger stimulus set as uniformly as possible, maintaining ratios of speakers, onsets to codas, and approximate numbers of each consonant from each speaker and in each position. Items were selected from a range of empirical difficulty levels (based on Experiment 1 ) such that the 597 clips’ accuracy levels would be similar to the accuracy levels obtained from the full set; that is, the subset contained items that in aggregate had a difficulty distribution similar to that of the full Experiment 1 dataset. The clips were divided into 4 exclusive sets of 149 or 150, and each judge heard materials from one of these sets. 2.1.2. Participant exclusion As in Experiment 1 , judges were excluded from the sample if they scored below 4 of 8 on the test of English vocabulary and grammar, or if their mean percentage of canonical consonant selections fell below 20%. Of the 35 judges recruited, three failed these criteria; thus, the final sample included 32 listeners. 2.2. Results: Low-pass-filtering Judges in Experiments 1 and 2 all made judgments of a small (n=5) set of VCVs under the same conditions (i.e., no sentence context in either group) before participating in the main trials. To confirm that the two participant groups were not grossly different from one another in performing the task, these trials were compared across groups. The mean percentage of correct consonant identifications was similar, with each group’s median number of attempts required to enter the correct consonant on these trials equal to 1.2 (Expt 1) or 1.0 (Expt. 2); mean attempts were similar as well (Expt. 1, 2.52 (sd 2.66); Expt. 2, 2.12 (sd 1.98)). Experiment 2 had two goals: to determine whether the main results of the first study would replicate in a new sample of judges, and to evaluate whether performance would increase when judges had access to aspects of the phonetic context of the VCV clips. Fig. 8 shows the mean proportion of each possible response for each of the target consonants. As in Fig. 1 , the gold-standard target was defined as the canonical consonant for its host word. Weighing each consonant equally, the mean accuracy over coda consonants was 25.0% (sd 15.8, median 21.6) and over onset consonants, 49.5% (sd 14.6, median 50.2). Comparing listener judgments to transcribed consonants rather than the canonical consonants, mean identification performance was slightly higher: codas, mean 29.7% (sd 16.3, median 25.9); onsets, 48.4 (sd 16.7, median 49.4). These performance levels are similar to those observed in Experiment 1 . Fig. 8. Open in a new tab Confusion matrix for consonants drawn from the Brent corpus, with VCV clips played in low-pass-filtered phonetic context. The left panel shows coda consonants, the right panel onsets. Color indicates the proportion of times that a given canonical consonant (the columns) received a given interpretation (the rows), as does the number in each cell. The font color of the number signals how many items contributed to that proportion. Comparison of Figs. 1 and 8 shows that judges’ responses were similar across the two experiments. Once again, performance on the onset consonants far exceeded performance on the coda consonants, and it was rare for judges to converge on particular wrong answers for any consonants, at least averaging over the items testing a given consonant. Performance in the two experiments can be visualized more directly by comparing the number of items for which judges reached various levels of agreement, as in Fig. 3 . Here, Fig. 9 presents the same data for Experiment 2 (in color), along with the analogous results for the matching items of Experiment 1 . Fig. 9. Open in a new tab Proportions of items for which Brent-corpus judges achieved a range of levels of agreement on the transcribed consonant. The left panel shows coda consonants, the right panel onsets. Color indicates dyad (mother-child pair) for Experiment 2 , in which judges heard the full sentence context in low-pass filtered form. Gray shows the same results from Experiment 1 , for the subset of items also tested in Experiment 2 . Each line indicates the proportion of items for which judges chose the correct transcribed consonant at least as often as the percentage given on the x axis. Because the audio clips in Experiment 2 were also tested in Experiment 1 , it is possible to evaluate the effect of presenting low-pass-filtered contexts within items. A regression analysis comparing the proportion of correct answers on each item in each experiment revealed that, contrary to expectation, judges were somewhat worse at the task when provided with the surrounding context, as shown in Table 3 . 3 Table 3. Regression analysis predicting accuracy in selecting the transcribed consonant, with experiment, true syllable position, and their interaction as fixed effects, and phone as a random effect. The full model R 2 was 0.277; the R 2 for fixed effects only, 0.188. β coef. Std. Err z value p value (Intercept) −1.644 (0.250) −6.57 <0.0001 expt. is expt. 2 −0.632 (0.317) −1.99 <0.046 consonant is an onset 1.900 (0.242) 7.84 <0.0001 expt. 2 and is onset 0.093 (0.349) 0.27 <0.789 Open in a new tab Phone by phone, performance on Experiments 1 and 2 was correlated. Considering the 34 onsets or codas for which there were at least 3 tested instances, the median correlation between the two experiments was 0.56 (25th percentile, .37, 75th, .68). Turning to participants’ identification of syllable position, there was no indication that presentation of the sentence context helped listeners to tell whether a consonant was an onset or a coda. The pattern of results was similar to that found in Experiment 1 , as shown in Table 4 . Table 4. Regression analysis predicting selection of onset as the consonant location. Predictors included the true consonant position (coda as baseline), the proportion of judges who correctly identified the consonant for that item (mean-centered), and their interaction. Random effects sd: phone, 0.00; mother, 0.023; residual, 0.180. β coef. Std. Err df t value p value (Intercept) 0.556 (0.018) 593.0 30.69 <0.0001 true C position is onset 0.135 (0.020) 593.0 6.60 <0.0001 % correct C i.d. (centered) −0.163 (0.058) 593.0 −2.77 =0.0057 correct % i.d. and is onset 0.304 (0.066) 593.0 4.60 <0.0001 Open in a new tab Thus, presentation of full sentential contexts in low-pass-filtered form did not indicate that the surprisingly variable consonant identifiability demonstrated in Experiment 1 could be attributed to the fact that the VCV clips being judged provided insufficient phonetic context, except inasmuch as the filtered phonetic contexts provided in Experiment 2 did not permit identification of the words. 2.3. Discussion Low-pass filtered speech impairs identifiability of speech sounds while preserving access to pitch contour, syllable rhythm, and overall amplitude envelope. We found here that providing the phonetic context of our VCV target stimuli in low-pass-filtered form did not render the consonants more identifiable and did not contribute positively to judgment of syllable position. This result indicates that the generally poor performance of our participants in Experiment 1 cannot be ascribed to the presentation of VCV clips in the absence of pitch, rhythm, amplitude envelope, or even position within the utterance. Of course, it is still possible that the high-frequency phonetic elements of context we removed here could contribute to consonant identification across an intervening vowel. Such “long-distance coarticulation” effects are better known, and probably stronger, among vowels than among consonants (e.g., Recasens, 2018 ). Consonants do affect the realization of immediately adjacent consonants (e.g., Mann & Repp, 1980 ), and it seems reasonable to expect that in some cases a consonant might have phonetically detectable consequences across an intervening vowel. For such effects to be relevant here, though, they would have to have an influence beyond what listeners could already take into account when hearing the interconsonantal vowel, which was of course presented in the clear here. We cannot rule out the possibility that in ordinary speech perception, where consonants are given in full phonetic contexts, there is some aspect of this context that is useful for consonant identification, is removed by low-pass-filtering, and does not achieve its useful effect via the lexicon. However, to our knowledge there is no evidence that infants could also make use of this information to help them identify consonants. In sum, the results of Experiment 2 lend confidence to our approach, and reinforce the conclusion that intervocalic consonants in infant-directed speech are frequently very difficult to identify. 3. Experiment 3: Speaker familiarity Adult native speakers of American English have more than an order of magnitude more experience with English speech than do infants under 12 month of age. It is on this basis that we consider these adults’ ability to recognize English consonants to estimate an upper bound on infants’ ability to do the same implicitly in the learning and recognition of words. However, our adult judges were, potentially, hindered in one respect relative to the infants they are intended to model: infants have more experience with the specific talkers in their environment than our adult judges did with the speakers voicing our stimuli. It is therefore possible that our judges were handicapped in their judgments by this lack of familiarity, and thus cannot estimate the upper bound on identifiability as intended. Although in American English both dialectal differences and intrinsic (biologically-based) talker variation would be expected to have a much larger impact on vowels than consonants, we cannot rule out a priori the possibility that talker experience could play a role in our consonant identification task. In Experiment 3 , we estimated the effect of greater experience with specific talkers on consonant identification. Judges participated in a version of Experiment 1 , performing exactly the same task with a subset of the Experiment 1 stimuli. Each of the judges evaluated VCVs from two of the three talkers, blocked by talker. For one of the two talkers, judges first listened to about ten minutes’ worth of unfiltered speech from that talker, and then responded to a series of VCVs. For the other talker, judges had no listening experience except the series of VCV judgment trials. To ensure that participants were listening carefully during the talker familiarization phase, they performed a semantic judgment task as they heard the sentences. By comparing consonant identification performance for the familiarized talker and the unfamiliarized talker, we can estimate the impact of familiarity with the talker on consonant identification. Of course, ten minutes of exposure is not of the same order as the several months of exposure experienced by the infants we wish to model. But native speakers are adept at talker adaptation; indeed, studies of adaptation have revealed significant adaptation even within one sentence (e.g., Ladefoged & Broadbent, 1957 ). One study using a reaction time measure to compare multi-talker and single-talker conditions found that about half of the hindrance caused by multi-talker interpretation was removed by presenting target stimuli in a two-syllable carrier phrase rather than no carrier phrase, and even more by a four-syllable carrier phrase ( Choi & Perrachione, 2019 ). The fact that native speakers can rapidly overcome the challenge posed by variable talkers suggests that the influence of talker-specific phonetic attributes on consonant identification should be significantly attenuated with modest amounts of experience (e.g., Kleinschmidt & Jaeger, 2015 ). Experiment 3 therefore directly compares consonant identification without prior exposure to the talker, and with several minutes’ speech experience with the talker. 3.1. Methods 3.1.1. Speech materials Familiarization sentences were extracted from the Brent Corpus. To produce a familiarization experience of 10–12 min long, between 240 and 250 utterances were selected from each mother among the three who had contributed VCVs for the test phase. The total duration of all speech material (not counting brief pauses inserted between each utterance) ranged from 9 min 7 s (mother d1) to 10 min 0 s (mother w1). The number of transcribed words ranged from 1554 (mother w1) to 1736 (mother f1). Audio amplitude was normalized sentence by sentence using sox . For each corpus, 72 sentences contained the name of an animal (for example, bear, chicken, bunny, horsie ). VCV sequences were a subset of the test materials of Experiment 1 . Items for which the consonant was a flap or a glottal stop were excluded, because participants in the prior experiments rarely chose these categories; here, we were more interested in optimizing the granularity of the comparison between familiarization conditions than in maximizing the coverage of the dataset. Test items were randomly sampled, but stratified such that Experiment-1 identification performance in the subset quantitatively matched Experiment-1 identification performance in the full set. Considering judges’ average match to the canonical consonant for each dyad and syllable position, the subset and the full set differed by less than 0.2%. Thus, in terms of overall performance, the materials of Experiment 3 were representative of the materials of Experiment 1 . 3.1.2. Procedure Each of 24 participants in the final sample responded to test materials from two of the three mothers’ speech. Presentation was counterbalanced across twelve stimulus orders. Four orders presented dyads f1 and w1; four presented f1 and d1; and four presented w1 and d1. Test trials were blocked by speaker. In each order, one of the two speaker blocks was preceded by the 10 min of familiarization speech; the other was not. For half of the trial orders, participants did the familiarization–test block first and then the unfamiliarized test block; for the other half, participants did the unfamiliarized test block first and then the familiarization–test block. This counterbalancing arrangement meant that for each test item, an equal number of participants heard it with prior talker familiarization and without; and an equal number heard it in their first block of trials and in their second. These test blocks were preceded by the English vocabulary and grammar screening also used in Experiment 1 . No participants were excluded for falling below threshold on this screening. Before the familiarization phase, participants were told that they would hear a series of sentences, and that they should listen carefully. If they heard a word for an animal, they were to click on a button labeled “animal” on the screen. This go/no-go task was used to help ensure attention to the sentences. Participants were excluded if their d-prime in distinguishing animal-containing utterances from the others fell below 0.75. This removed two participants (0.48, 0.58; the next lowest scored 0.93). One participant appeared to have followed the reverse pattern and achieved a d-prime of −.96, and was retained in the sample. For the consonant identification trials, participants were instructed just as they had been in Experiment 1 . Participants were recruited on the Prolific platform using the same procedures as in the prior experiments. As in Experiment 1 , participants were excluded if their overall accuracy in matching the transcribed consonant across all test trials fell below 20%. This resulted in the exclusion and replacement of one participant. Testing was done in March–June of 2024. 3.2. Results: Speaker familiarity Participants did not identify consonants significantly more accurately after actively listening to the talker for ten minutes. The overall results are shown in Fig. 10 , which follows the conventions of Fig. 9 . Fig. 10. Open in a new tab Proportions of items for which Brent-corpus judges achieved a range of levels of agreement on the transcribed consonant. The left panel shows coda consonants, the right panel onsets. Color indicates whether the speaker’s voice had been familiarized to the judge or not. Each line indicates the proportion of items for which judges chose the correct transcribed consonant at least as often as the percentage given on the x axis. Performance was evaluated formally using logistic regression analysis. The initial model predicted match to the transcribed consonant (a binary variable) from the fixed-effects predictors condition (familiarized or unfamiliarized talker), syllable position (onset or coda), and block (first or second), with random effects phone, judge, and speaker, each with random slopes for condition. The random slopes led to singularity and accounted for minuscule portions of variance, and so were removed. Clearly nonsignificant (p > 0.10) interactions were removed, with the exception of the condition:syllable-position interaction which was retained for theoretical interest (see Table 5 ). Table 5. Regression analysis predicting accuracy in selecting the transcribed consonant, with condition, true syllable position, block, and the interaction of condition and position as fixed effects, and phone, judge, and speaker as random effects. The full model R 2 was 0.271; the R 2 for fixed effects only, 0.077. β coef. Std. Err. z_value p_value Intercept −1.203 (0.257) −4.676 <0.0000 cond. is unfamiliar −0.141 (0.137) −1.027 0.3046 position is onset 1.264 (0.124) 10.170 <0.0000 block is 2nd 0.331 (0.076) 4.360 <0.0000 cond. unfamil. and pos. onset −0.212 (0.164) −1.292 0.1965 Open in a new tab As the analysis table shows, familiarity with the speaker did not cause significant improvements in judges’ ability to identify intervocalic consonants. Judging an unfamiliar talker led to small and inconsistent decrements that were numerically greater in onset than coda position, but in no case were they significant. As expected, performance was considerably better for onsets than codas. Performance was also better in the second block, presumably a practice effect as this improvement did not interact with speaker familiarity. 3.3. Discussion The fact that about 10 minutes’ exposure had a negligible impact on participants’ identification of consonants in our test VCVs indicates that judges’ highly variable performance in Experiment 1 (and subsequent experiments) should not be attributed to a lack of familiarity with the voice of the talkers: providing extensive familiarization to a talker’s voice, far more than is typically available in speaker adaptation studies, did not improve performance. This is not to deny the existence of speaker adaptation and the associated cognitive costs; it only indicates that when our judges performed very well in identifying certain consonant tokens, and very poorly in identifying others, it is unlikely that the poor tokens were difficult because of idiosyncrasies of the talker’s overall way of speaking or anything peculiar in how they normally say specific consonants. It is more likely that tokens that are difficult to identify are simply reduced beyond recognition, rather than being products of an unfamiliar accent or hard-to-assimilate vocal tract structure. Of course, it is conceivable that even more familiarization practice would have led listeners to register extremely subtle featural characteristics of the talkers, and furthermore that such characteristics are routinely available to infants in the earliest stages of language learning, but this seems unlikely. The fact that the corpus is full of easily interpretable sentences seems to conflict with the dubious availability of the intervocalic consonants for categorization. How can this be? We will return to this question in the General Discussion, but the short answer is that speech understanding, like all of perception, depends on background knowledge. We can understand what people say because we already have expectations about what people are likely to say and how they will say it. This fact is part of what makes language acquisition so interesting as an engineering problem: how can infants acquire this system before they have the necessary “top down” knowledge? In the remaining two experiments, we consider two different ways to understand this problem. First, we consider the items that adults were relatively successful with, and estimate how often this may have been because they were using lexical knowledge even over our bare VCV stimuli. Second, we evaluate whether infant-directed speech differs from adult-directed speech in the clarity of its consonants. 4. Experiment 4: Lexical bias When listeners were successful in identifying consonants, was it because the consonants themselves were very clearly realized, or was it because the VCV provided enough information to make clear which words were involved, allowing listeners to use a lexical route to deduce what the consonant must have been? The premise of testing only VCVs rather than longer stretches of speech is that VCVs should usually be insufficient for identifying words. Yet this may not always have been the case. To help improve our intuitions about this, in Experiment 4 we asked listeners to guess which words they were hearing, rather than guessing which consonants they were hearing. Because we were interested in what leads judges to get consonants right, we selected the items for which at least 80% of guesses were correct. Very few of these were coda consonants, so we restricted the set to the onset consonants. We reasoned that if judges hearing a given VCV responded with a great diversity of words that all started with the transcribed consonant, it would indicate that consonant was identifiable independently of a particular word. On the other hand, if several judges responded with the same word (and that word started with the test consonant, of course) then listeners in prior experiments may have arrived at that consonant at least in part by recognizing the word and following the lexical route. Of course, this does not show that the consonant was hard to identify–perhaps it was indeed easy to identify and this contributed heavily to the identification of the most likely word. But this outcome would nevertheless imply that the stimulus as a whole might not have been short enough (for example) to force a bottom-up category judgment. 4.1. Methods 4.1.1. Speech materials Starting from the stimuli used in Experiment 1 , the items for which correct identification of the transcribed consonant met or exceeded 80% of judgments were selected. A few stimuli originally from words deemed unlikely to be guessable by participants were excluded (“d” as a letter-name, “dillon”, “huh”, “poo”, and “poopoos”). This resulted in 332 tokens, deriving from 162 unique words. One instance of each word was randomly chosen, producing 162 stimulus items. Due to an error, two of these were omitted from the experimental orders, leaving 160 items for evaluation by each participant. 4.1.2. Procedure Participants were recruited over Prolific as in the previous studies. They were told that they would hear short audio clips made up of a vowel, a consonant, and a vowel, all taken from recordings of mothers talking to their babies. The instructions indicated that the consonant was always the first sound in a word, and that they would usually not hear the whole word. We also explained that there might be more than one word that fits, and they should just pick one; they also had a “no word” option they could click on if they had no idea what the word might be. Participants were told “Don’t click ‘no word’ too often, because we might think you aren’t trying.” Then they heard two practice items that were taken from speech in the Buckeye Corpus. As in the previous experiments, judges were asked to identify the English word for a few pictured objects and were given a few multiple-choice questions involving English syntax. Then they were presented with the 160 trials in a random order. To achieve a desired sample size of 16, 19 participants were tested. Three were excluded for choosing the “no word” option on more than a quarter of the trials (far more than any other participants). Two of these three were also anomalous in entering a large number of nonwords or a large number of multiword sequences. Testing was done in December of 2023. 4.2. Results: Lexical bias The free responses were examined for anomalies. When judges entered two words, the one that started with the target consonant was retained as the response. For the purpose of analysis, words were converted to lemmas, so that (for example) “chewed” and “chewing” were converted to “chew,” or “dog” and “dogs” to “dog.” The rationale for this collapsing procedure was that listeners in the consonant choice experiments could not have had any way to know whether the word was singular or plural (and so forth) in most or all cases, so if their choice were driven by the lexicon, it would have been by the same lemma, whether Experiment 4 ’s judges wound up typing it as “dogs” or “dog.” For this reason, words whose morphological variants involved a different vowel, like “could” and “can,” were not converted to the same lemma; “could” and “can” were not considered the same judgment. Consider the case in which listeners in Experiment 1 heard the VCV [aIkæ], taken from “I can.” If a listener in the current experiment nominated “could” as the word that the CV initiates, this seems a quite different thing from a listener in the current experiment nominating “can,” for our purposes in understanding how lexical activation might feed back to phone identification. Responses therefore fell into four kinds: (a) matching the source lemma; (b) containing the target consonant but from a different word; (c) failing to contain the target consonant; and (d) the asterisk marking “no idea.” Fig. 11 shows for each item the breakdown of responses into these categories. Fig. 11. Open in a new tab Experiment 4 : Correspondence between word guesses and the actual source word of each VCV, where the VCVs were selected as the items with the easiest-to-identify consonants. Each column represents a stimulus VCV from Expt. 4. Each of these stimuli was rated by 16 listeners (the y axis). Colors show how many listeners guessed the same word that the CV had come from in the corpus (light green); how many listeners guessed a different word, but a different word starting with the same consonant (medium blue); how many guessed a word starting with a different consonant from the VCV (red); and how many failed to guess (dark brown). For some VCVs (to the right in the plot), listeners hearing the VCV could guess the true original word; but for most VCVs, listeners guessed a different word with the correct consonant (i.e., the plot is largely blue). Gray circles indicate for each item the number of distinct lemmas entered by the 16 participants. Items were sorted from left to right according to the proportions of hits (word responses starting with the target consonant) and misses (other word responses). Participants entered a word that started with the target phone 85.5% of the time. This is as expected, given that these VCVs were chosen for evaluation because prior participants had identified the consonant correctly at least 80% of the time. Entries’ lemmas matched the source word of the VCV 29.5% of the time, over all trials. Taking these statistics together, then, 34.5% of the word responses that contained the correct consonant were also the original word; conversely, almost 2/3 of the times participants apparently arrived at the right consonant judgment, they did not get there by identifying the original word. Considering only responses that contained the target consonant, over items the mean number of different words entered was 6.2 (25th percentile, 4.0; 75th percentile, 8.0). As Fig. 11 shows, these statistics varied considerably from item to item. There were a few cases (toward the right in the plot) where most listeners came up with the word that had originated the VCV. There were many more items for which most listeners produced a word with the target consonant but not the original VCV word (the middle of the plot). Also visible in the Figure is the substantial diversity in listeners’ responses to most items. Over all items, the mean number of different lemmas produced was 8.1 (25th percentile, 6.0; 75th percentile, 10.0). Perhaps the strongest evidence for the independent identifiability of the target consonant would come from cases in which participants entered a word that started with that consonant, but then followed with a different vowel. For example, if a VCV came from the sequence “a doll,” and a participant entered “daddy,” it is hard to escape the conclusion that the /d/ sound was identifiable independently of the following lexical context. We can evaluate this by comparing the vowel of the canonical pronunciation of the participants’ word responses, and the second vowel of the stimulus VCV. Among the cases where the response matched the target consonant, the vowel also matched 61.7% of the time. If we exclude the cases where participants identified the source word, the vowel of their lexical response matched the target VCV’s second vowel less than half of the time (44.6%). This indicates that among the “easy” items selected for testing, many consonants did not depend upon recovering lexical context for their identifiability. Although it is clear that there was substantial heterogeneity among our stimulus items, these results allow for some general conclusions. The most important is that when consonants were easy to identify within VCVs, it seems to have been because the consonant itself was clearly identifiable, and not because listeners first identified the word specified by the CV and worked their way around from the lexicon to the phonology. This might have happened sometimes (we cannot rule it out, particularly when listeners identified the original target word), but even in those cases the results are still compatible with a bottom-up account in which the consonant is first identified and then passed to the lexicon. This may be viewed as good news for the language-learning infant who is working to build a lexicon. Although the primary message of these experiments is that much of infant-directed speech is difficult to interpret bottom-up because of phonetic reduction and other sources of variability, there is another side to the story. Some tokens are, in fact, readily identifiable, at least to adult native speakers. The evidence suggests that listeners responding to these “easy” cases were not “cheating” by taking advantage of severe lexical constraints to identify otherwise inscrutable consonants in a top-down manner. Some consonant tokens are clear. They are just not the typical case. 5. Experiment 5: Adult-directed speech Most research evaluating the intelligibility of infant-directed speech is grounded in a comparison of infant-directed speech and adult conversation. Infant-directed speech sometimes bears more of the acoustic hallmarks of hyperarticulation, such as vowels’ greater dispersion from the center in formant space ( Kalashnikova & Carreiras, 2022 ; Kuhl et al., 1997 ) but this is not a settled conclusion, partly owing to variability in methods and in corpora ( Cox et al., 2023 ). The availability of the Buckeye Corpus of adult conversation ( Pitt et al. 2007 ) makes it feasible to apply the present methods to adult conversation, thereby introducing a new way to compare infant-directed speech and adult-directed speech. The Buckeye Corpus was annotated to the phone level, with trained transcribers taking care to note actual phonetic realizations rather than canonical, “dictionary” pronunciations of words. The Buckeye Corpus is also large enough that it was possible to test a more uniform sample of onset and coda consonants than were available from our annotations of the Brent Corpus. Apart from the change of source materials, our procedures matched those of Experiment 1 . 5.1. Methods 5.1.1. Speech materials Buckeye interview recordings from four speakers were selected: talkers S04, S08, S12, and S39. These were chosen because they were adult women under 30 years of age, and therefore the closest match to the three Brent Corpus talkers. The Buckeye Corpus is delivered with meta-data indicating the onset and offset time of each word and phone in each conversational turn of the participating speaker. From these data, the full set of word-onset and word-offset VCV sequences was selected from all of the recording sessions of these four speakers. Items were excluded if they were marked with tags indicating laughing, vocal noises, interviewer interruptions, and the like; and were excluded if the intervocalic consonant was the entire word. This yielded a pool of about 2900 tokens from each speaker. Tokens were then excluded if the dataset included fewer than five tokens of that consonant in a given syllable position for every talker. This excluded four onsets (ŋ, ɾ, Ʒ, z) and several codas (e.g., b, ð, g, h, w). From the tokens pool, up to 12 instances from each available set of talker X syllable-position X phone were randomly sampled, yielding a set of 1916 stimulus items. Table 6 indicates the number of VCV items that were evaluated from each speaker in each syllable position. The set included a greater proportion of coda consonants (43.5%) than were included in Experiment 1 (27.7%), because of limitations of the Brent pool of possible VCVs. The median duration of the VCVs was 238 ms (range, 74–930; first quartile 188, third quartile 308). Audio files were prepared as in Experiment 1 . The stimuli were divided into nine sets with approximately equivalent numbers of items corresponding to each phone, consonant position, and speaker. Table 6. Study 5 (Buckeye corpus) tokens by speaker and word position. Counts Proportion Speaker 04 08 12 39 04 08 12 39 onset 274 263 269 273 .58 .56 .56 .56 coda 199 208 215 215 .42 .44 .44 .44 Open in a new tab 5.1.2. Procedure As in prior Experiments, judgments were made by listeners recruited on Prolific. 71 judges participated. Nine were excluded for extremely poor performance in the task. These participants tended to do worse on the English test. One more participant was excluded only for very poor performance on the English test. The final sample thus included 61 judges, with an attrition rate (14%) similar to the rate in Experiment 1 (19%). Instructions were the same as those used in Experiment 1 . An average of 6.8 judges responded to each VCV soundfile. Testing was done in April of 2021. 5.2. Results: Adult-directed speech Because our primary concern is the comparison of responses to the adult-directed Buckeye and infant-directed Brent corpora, a complete description of the Buckeye judgments alone, following the analysis strategy of Experiment 1 , is given in supporting online materials, including the confusion matrix for canonical consonants. To determine whether accuracy in identifying consonants was greater for items from the Brent corpus than the Buckeye corpus, accuracy was compared in a binomial multilevel logistic regression analysis with the judge’s match to the transcribed consonant identity as the dependent variable, and corpus and syllable position as the treatment-coded fixed-effects predictors, with listener as one random effect, and item nested within phone as another random effect. This analysis revealed significantly lower accuracy for codas than for onsets in the (regression base level) Brent corpus, as expected; a nonsignificant decrement in performance for the Buckeye corpus when considering onset consonants (the base level of syllable position); and a significant interaction between corpus and syllable position, indicating that performance was significantly better than expected when consonants were codas from the Buckeye corpus. The coefficients and other fixed-effects analysis elements are shown in Table 7 . Table 7. Study 5: Analysis of Buckeye and Brent identification accuracy. Input data were at the trial level. Syllable position and corpus were treatment coded, with Onset and Brent as the base levels, and Coda and Buckeye as the treatment levels. See text. The full model r 2 was 0.640; the r 2 for fixed effects only, 0.047. β coef. Std. Err z value p value (Intercept) −0.372 0.379 −0.981 0.3268 syll. position [coda] −1.557 0.099 −15.790 0.0000 corpus [Buckeye] −0.159 0.135 −1.180 0.2380 position [coda]:corpus [Buckeye] 0.383 0.120 3.193 0.0014 Open in a new tab To help simplify interpretation, we computed multilevel logistic regression analyses separately for each syllable position, with Corpus as the fixed effects predictor and including the same random effects as in the previous analysis. In the analysis of only onsets, there was not a significant effect of corpus (coefficient, −0.167, z −1.193, p = 0.23, i.e. a nonsignificant trend toward worse performance on Buckeye items). In the analysis of only codas, again there was not a significant effect of corpus (coefficient, .206, z 1.364, p = 0.17, i.e. a nonsignificant trend toward better performance on Buckeye items). It is not clear why syllable position affected the relative identifiability of adult-directed and infant-directed consonants. In principle, it is possible that resyllabification is more common in adult-directed speech, so that a larger proportion of codas were produced as onsets and that this led to more readily identifiable consonants among the nominal codas. But there is no independent evidence that this took place, and as indicted below, when performing the syllable position judgment task, listeners were not more inclined to report onsets in the Buckeye dataset (in a signal detection analysis, mean criterion c = −0.300, favoring onsets) than in the Brent dataset (mean c = −0.282, also favoring onsets; t(125.8) = 0.245, p = 0.81 ns ). The distribution of items with various frequencies of correct identification is displayed in Fig. 12 . Similar information given as a hazard plot analogous to Fig. 3 is presented in the supporting online materials, showing that the identifiability curves among Brent and Buckeye speakers overlapped, with variation among speakers being similar to the variation among corpora. Fig. 12. Open in a new tab Experiment 5 . Comparison of item by item consonant identification probabilities for the Brent corpus materials (infant-directed speech; green) and the Buckeye materials (adult–adult conversation; brown). The area under each curve is one. As in Experiment 1 , listeners were asked to decide whether each consonant was the start or end of a word. Their percent correct was 59.0% (sd by subjects, 10.6) for the Brent stimuli, and 50.9% (sd, 10.9) for the Buckeye stimuli; however, given that the Brent stimulus set comprised 72.4% onsets and the Buckeye 58% onsets, the advantage among Brent participants can be attributed in part to the overall bias to respond “onset.” To address the onset bias head-on, we computed a d-prime score for each listener. The mean d’ for Brent items was 0.245, as noted above; the mean d’ for Buckeye items was 0.0100 (median, −.0439; sd, 0.579). The latter was not different from chance levels (t(60)=0.135, ns). Sensitivity to positional cues was significantly greater for Brent participants than for Buckeye participants, based on a t-test over d’ values (t(107.5) = 2.620, p = 0.0101). A multilevel logistic regression analysis predicting correct responses from syllable position, corpus, and their interaction, showed a significant effect of syllable position in the reference level of Corpus (Brent), no effect of Corpus at the reference level of Syllable position (coda), and a significant interaction, indicating that Buckeye listeners were worse at judging onsets than would be expected based on the other predictors ( Table 8 ). This would seem to offer some hope for the thesis that infant-directed speech offers phonetic segmentation cues that could be helpful to infants in locating words. However, this assistance is very limited; in fact, given that the ratio of intervocalic singleton consonants that are word onsets to those that are word codas is about two and a half to one, infants might be best off simply assuming they are all onsets, and perform much better than our judges here. Table 8. Syllable position judgments (onset versus coda) over stimuli from the Brent and Buckeye corpora. The table gives results from a logistic regression analysis predicting correct syllable position judgments based on true syllable position (defined lexically) and corpus (Brent vs. Buckeye) and their interaction as fixed effects, and listener and audio file (nested within phone) as random effects. β coef. Std. Err. z value p value intercept −0.332 0.076 −4.366 0.0000 syll. position [onset] 0.981 0.052 18.823 0.0000 corpus [Buckeye] −0.093 0.096 −0.968 0.3331 position [onset]:corpus [Buckeye] −0.154 0.065 −2.378 0.0174 Open in a new tab 5.3. Discussion A comparison of responses to consonants taken from the Brent infant-directed speech corpus, and consonants from the Buckeye adult–adult conversational corpus, revealed that among word onsets, infant-directed consonants were not significantly more identifiable, though there was a modest trend in that direction. Among word codas, infant-directed consonants were somewhat less identifiable than adult-directed consonants, though again that trend was not significant. These minor differences between the corpora enter into a wider debate about whether infant-directed speech is, generally, phonetically clearer than adult-directed speech. The majority of studies making this comparison have examined vowels, primarily by analyzing acoustic measurements: duration, pitch and pitch movement, and most importantly the dispersion of formant values (e.g., Kuhl et al., 1997 ). Some studies find that infant-directed vowels, in particular the point vowels [i, a, u], collectively occupy a greater expanse of phonetic space than the adultdirected vowels (e.g., Hartman, Ratner, & Newman, 2017 ; Kalashnikova & Burnham, 2018 ). Other studies find that the vowels are not more distinct from one another (e.g., Cox et al., 2023 ; Cristia & Seidl, 2014 ). Studies of consonants are fewer in number, and also present a rather mixed picture ( Bernstein-Ratner, 1984 ; Khlystova, Chong, & Sundara, 2023 ; Shockey & Bond, 1980 ; Sundberg & Lacerda, 1999 ). Part of the difficulty in coming to a firm general characterization of the difference in the two registers is that the infant-directed speech register is highly variable—some of its attributes may vary across child ages, among languages, and according to the goals of the parent in the conversation (e.g., Beech & Swingley, 2024 ; Hartman et al., 2017 ; Stern, Spieker, Barnett, & MacKain, 1983 ; for a meta-analysis focusing mainly on vowels, see Cox et al., 2023 ). What we found here, at least for our small sample of mothers and other adult women, is that the distribution of “easy” and “difficult” intervocalic consonants was quite similar over our infant- directed and adult-directed samples. 6. General discussion The Brent corpus is, on the whole, interpretable by ear. Its sentences are usually easy to understand. When adults nonetheless perform poorly in identifying speech sounds extracted from these sentences, the contrast reveals the importance of listeners’ knowledge of the language in supporting speech comprehension. This knowledge of the language is available when samples are long enough to include words, but highly restricted otherwise. We can understand speech despite pervasive hypoarticulation because we are good at guessing what the talker is likely to have said: we grasp the discourse context, we know the words of the language, we know which word combinations are frequent, and we can disfavor interpretations that violate the language’s grammar. Native speakers also know the language’s typical forms of phonological reduction and assimilation, to the extent that they are regular enough to be learned and reproduced in the dialect. This is why adults can make sense of speech in spite of rampant phonetic reduction: speech comprehension is not strictly “bottom up.” But infants come into the world without any of these language-specific sources of knowledge. How can they break into a system that is optimized to balance the needs of adult talkers and listeners, but not infant learners? The usual answer is that infant-directed speech is special ( Ferguson, 1964 ). In talking with their infants, parents (and other discourse partners) modify their pronunciation in ways that support learning. They adopt exaggerated pitch contours, lengthen some sounds, and emphasize consonants and vowels to make them easier to categorize. In support of this notion, individual variation in the tendency to produce exaggerated forms of words and sounds has been found to be correlated with measures of language development in infants ( Hartman et al., 2017 ; Kalashnikova & Burnham, 2018 ; Liu et al., 2003 ; Song, Demuth, & Morgan, 2010 ). As discussed above, though, even if infant-directed speech were found to be somewhat clearer than adult-directed speech much of the time, this would not itself mean that linguistic structure is, ipso facto , readily available in infant-directed speech. Demonstrating an infant-directed-speech clarification effect is not the same thing as explaining how infants achieve the learning they do. To explain this achievement, we need learning models that can inform us about how the data that is plausibly accessible to infants could support the acquisition of structure (e.g., Ludusan, Mazuka, & Dupoux, 2021 ). Until about ten years ago, infants’ ability to detect unfamiliar words in continuous speech was usually modeled as a problem of identifying statistically coherent chunks in strings of consonants and vowels. Often traced to Brent and Cartwright’s ( Brent & Cartwright, 1996 ) work, these models emerged in response to demonstrations of infant statistical learning in the lab ( Aslin et al., 1998 ). They included Daland and Pierrehumbert (2011) , Perruchet and Vinter (1998) and Swingley (2005) , or Monaghan and Christiansen (2010) , among others. They presupposed (somewhat apologetically) that infants have access to a sequence of consonant and vowel categories derived from the canonical pronunciations of the words in the corpus. The current results help to identify just how Panglossian this assumption was. Although it is very difficult to make the relevant measurements in infants, it is unlikely that infants exceed adults’ performance here during their everyday apprehension of the speech signal. Our participants were mature native speakers of the language, deciding on only one consonant at a time, and given full VCV environments without any time pressure. When infants are confronted with sentences, their performance in categorizing speech sounds is almost certainly worse than adults.’ Although we do not present a reanalysis of the word segmentation models here, it is likely that altering the input so that only about half of the intervocalic consonants are consistently identifiable would have a substantial impact on model performance, even if all the other sounds were left intact (see, relatedly, Beech & Swingley, 2023 ). In the first experiment presented here, we found that about 60% of word-onset consonants were identifiable by most listeners, and about 25% of word-final (coda) consonants were. In Experiment 2 , we found that performance did not improve when the VCVs were placed in their sentential context, with this context low-pass-filtered to hinder word identification. Poor performance was still found even in the presence of at least some cues to utterance position, sentence rhythm, sentential pitch patterns, and global features of the speaker’s vocal tract. In Experiment 3 , we compared performance when listeners had no prior experience with the talker’s voice, and when they had about 10 min of continuous exposure to the talker’s voice. Familiarization with the talker’s voice did not improve performance in recognizing consonants. In Experiment 4 , we asked whether good performance on some items might be due to a lexical route in which the VCV led to word recognition, which then strongly constrained the choice of consonant. Listeners were presented with “easy” VCVs in which the consonant had been identified with high accuracy in Experiment 1 , and were asked to guess which word the CV started. Among trials in which they got the consonant right, they also guessed the source word about a third of the time. Frequently they got the consonant right but their lexical guess was not only wrong, but had the wrong vowel. Because they got the right consonant, but the wrong word, we know that some of the easier-to-identify consonants are easy to identify in and of themselves, and not because they strongly suggest their source words and therefore permit a “top down” route to identifying the consonant. Finally, Experiment 5 compared consonant identification performance on adult-directed speech VCVs and infant-directed speech VCVs. Performance on adult-directed and infant-directed speech was similar. In sum, we have found that the consonants of infant-directed speech are extremely heterogeneous in their identifiability. Some can be readily identified. Many cannot. And when consonants are misidentified, very often the guesses are not even close ( Figs. 4 , 5 , and 6 ). Throughout, we also tested adults’ ability to decide whether a consonant was the first in a word, or the last. The identifiability of consonants’ syllable affiliation is useful to estimate, because statistically speaking, onset consonants are good guesses for word onsets, and codas are not. If infants had access to syllable boundary information, it could offer them a head start in identifying word boundaries. Performance in this task was above chance, but quite poor by any standard (e.g., d-prime about 0.25). It should not be assumed that infants have access to syllable boundaries. The results suggest that it is simply not tenable to presuppose that infants convert the entire speech stream into a sequence of consonants and vowels, group these sounds into syllables, and segment the speech into words by counting syllable types and computing conditional probabilities over these types. Nor is it tenable to make similar assumptions but skip the syllable identification step and proceed directly to phones. In recent years, several research groups have begun to abandon the presupposition that infants encounter the continuous speech stream as a sequence of consonants and vowels at all, let alone the consonants and vowels that would be present if talkers were to speak with dictionary pronunciations. Models that take raw speech samples as inputs and learn structure by seeking regularities in the auditory signal have displayed behaviors analogous to those demonstrated by infants in speech-sound categorization experiments and word segmentation experiments ( Khorrami, K., & Räsänen, 2025 ; Lavechin et al., 2024 ; Schatz et al., 2021 ; Swingley & Algayres, 2024 ). One theoretical response to these modern models is to deny that infants’ phonetic adaptation to native-language phonology ( Kuhl, Williams, Lacerda, Stevens, & Lindblom, 1992 ; Werker & Tees, 1984 ) depends on learning phone-sized phonetic categories, given that analogous learning effects can apparently arise without the construction of a phone level of interpretation (e.g., Matusevych, Schatz, Kamper, Feldman, & Goldwater, 2023 ). The phone-free models do seek out statistically coherent portions of the signal, but this coherence is not defined over discrete phone categories. Although our results suggest that the phone-based models made overly optimistic assumptions about phone identifiability, they do not speak directly to the debate about whether infants interpret speech as a sequence of consonant- or vowel-sized segments—they still might, computational (non)existence proofs notwithstanding. Infants might have an innate or early-developing tendency to consider consonants and vowels as separable kinds of sound, and therefore be predisposed to break up the speech signal into roughly phone-sized units. If so, our results indicate that infants get much better data from some tokens than from others. Could infants learn primarily from the better data? Perhaps, if they can detect it. The portions of the signal that exemplify categories more distinctly, whether phonological or lexical, might be marked in ways that are accessible to infants: they may fall on pitch peaks, be longer in duration, coincide with parent touch or other physical movement, synchronize with distinctive facial signals, or take place during discourse moments of greater referential clarity ( Adriaans & Swingley, 2017 ; Beech & Swingley, 2024 ; Gogate & Bahrick, 2001 ; Tincoff, Seidl, Buckley, Wojcik, & Cristia, 2019 ; Vihman & Majorano, 2017 ). If we view the infant’s primary task as learning the language, their failure to grasp some instances of sounds or words may be less important than their ability to learn from infrequent, but highly informative, instances. In our task, judges made categorical decisions by choosing from a set of orthographic labels. Of course, infants do not explicitly connect speech sounds to written letters. Still, when infants are simulated in traditional word segmentation models, all of the computations take discrete categories as their operands, where each category label serves as a marker subsuming all tokens of a given type. Thus, both the standard models, and our task, require that speech categories be identified: for the models, the consonant must be identified with a symbolic code that thenceforth stands for that portion of the signal; for the listeners, the same is true, with the proviso that symbolic code is the English orthography. This is the code that is automatic and natural in literate English-speaking adults—so much so that adults have considerable difficulty thinking of speech sounds any other way (e.g., Treiman & Cassar, 1997 ). What we contend, then, is that adult responses in our task delineate an upper bound on what infants, or for that matter any listeners, might extract from the speech signal independently of knowledge of the lexicon. If anything, the picture painted here is too optimistic, given that our listeners probably did use lexical knowledge at least some of the time. A consequence of this conclusion is that text-based computational models of word segmentation probably overestimate their success rates substantially. There are surely differences between adults performing in the present tasks, and infants listening to their parents’ speech. Adults know some 40,000 lemmas ( Brysbaert, Stevens, Mandera, & Keuleers, 2016 ). Adults have some conscious grasp of categorical phonology earned through reading, learning morphology, and differentiating words. Could all this experience paradoxically make them worse at identifying consonants? The mechanism by which experience and knowledge might hinder performance would be intrusion from the lexicon, and perhaps intrusion from other phonotactic knowledge. Adults have a strong set of intuitions derived from implicit knowledge of linguistic probabilities at multiple levels. In principle they could hear a [b] in [aibo] and (implicitly) override their [b] judgment based on knowledge that in, say, an [ai o] context, [d] is more frequent because of phrases like “I don’t.” But for this knowledge to have set adults astray here, it would have to mismatch the probabilities present in the experimental materials. This seems unlikely, because the materials were themselves sampled from actual conversational corpora, whether parent-infant conversation (the Brent corpus) or adult–adult conversation (the Buckeye). To the extent that the test materials exemplified probabilities characteristic of English, knowledge of English should make adults better at the task, not worse. Ultimately, children need to learn the phonology of their language. This will be required for making good inferences about minimal pairs, for relating production and comprehension, and for isolating and learning morphological regularities that are defined using single sounds and their allophonic variants. Where does it come from? The developmental psycholinguistics literature suggests two main ways to think about the emergence of receptive phonology. One account holds that infants are born with an innate first pass at the problem. They have a phonetic similarity space evolved for language, and a set of provisional categories or perceptual boundaries within which experienced speech tokens fall, in a format conducive to supporting the perceptual learning that ultimately results in language-specific phonetic categories (e.g., Kuhl, Conboy, Coffey-Corina, Padden, Rivera-Gaxiola, & Nelson, 2008 ). This hypothesis has been supported by experiments about phonological categorization and adaptation in infants (the loci classici being Eimas et al., 1971 and Werker & Tees, 1984 ), and by crosslinguistic studies showing that these language-specific adaptations are relevant for lexical representation in toddlers (e.g., Dietrich, Swingley, & Werker, 2007 ; Ramon-Casas, Swingley, Bosch, & Sebastián-Gallés, 2009 ). The other account holds that speech segments emerge primarily from the lexicon: first, infants learn words as analog acoustic forms, and then, as children learn more words that sound similar to one another, they begin to learn phonological categories—to help them keep similar-sounding words apart ( Metsala & Walley, 1998 ) or because knowing more words means having a larger dataset over which to infer similarities (e.g., Lindblom, 1992 ; for a thorough review, see Nittrouer, 2021 ). Since the time when these accounts were first proposed, the former account has gradually approached the latter, maintaining the original commitment to important phonological developments in the first year of life, but also admitting various influences from the structure of the lexicon and acknowledging substantial additional development through and beyond the toddler period ( McMurray, Danelz, Rigler, & Seedorff, 2018 ). For example, computational modeling suggests that the early “protolexicon” and the distributional structure of phone tokens jointly constrain phonetic learning ( Feldman, Griffiths, Goldwater, & Morgan, 2013 ; Swingley, 2009 , 2019 ; see also Hitczenko & Feldman, 2022 ; Werker & Curtin, 2005 ). Some early word knowledge may be anchored around infants’ own first productions, as well; analysis of children’s speech patterns very often reveals consistencies of form that do not align well with segmental descriptions (e.g., Vihman & Majorano, 2017 ). On these modern accounts, the learning of words and segments is complementary, and quite possibly has to be, because of the extreme variability present in the speech data ( Bion, Miyazawa, Kikuchi, & Mazuka, 2013 ; de Seyssel, Lavechin, & Dupoux, 2023 ; Swingley & Alarcon, 2018 ). In addition, toddlers have turned out to be surprisingly reluctant to learn phonological minimal pairs under some conditions ( Kemp, Scott, Bernhardt, Johnson, Siegel, & Werker, 2017 ; Stager & Werker, 1997 ; Swingley, 2016 ; Swingley & Aslin, 2007 ), raising the possibility that even if words are represented in terms of phonological categories, children are wary of ascribing lexical distinctions to small phonological differences ( Beech & Swingley, 2023 ; White & Morgan, 2008 ). What the present work contributes to this debate is a way to characterize the dataset from which infants draw their inferences about phonetic and lexical structures. The proposal that infants can derive their language’s phonetic categories through distributional analysis of the speech sounds they experience has been difficult to support using computational or statistical models given realistic training data (e.g., Antetomaso, Miyazawa, Feldman, Elsner, Hitczenko, & Mazuka, 2017 ; Swingley & Alarcon, 2018 ). The fact that many consonants of infant-directed speech are not interpretable to adults suggests that the explanation for distributional clustering models’ failure to yield accurate phonetic categories may not only be a matter of deficient talker normalization (which our judges presumably can do well) or poor simulation of the basic human perceptual space (which is available to our human judges, naturally), but also a matter of needing to use information more efficiently, possibly by filtering out low-quality tokens that are inscrutable to native speakers, and employing other sources of constraint ( Adriaans & Swingley, 2017 ; Hitczenko & Feldman, 2022 ). Exactly what infants stand to gain from the most reduced tokens is not clear; they might weight the clearer instances more heavily without completely ignoring the very reduced ones. The present studies have important limitations. First, although many consonant tokens were evaluated, the primary dataset came from the speech of three American English-speaking mothers. Although the results were similar among these three talkers, it is possible that future work will reveal that they were not representative in some way. Although the risk of this seems small (for example, their children’s vocabulary sizes were not unusual among the children in the Brent corpus), our results would be more secure with a broader base of speakers, particularly in the case of the comparison between infant-directed and adult-directed speech. The dataset also only concerns consonants in VCVs at word boundaries. Consonants in other environments may be easier, or more likely harder, to identify. Second, it is possible that adults’ ability to identify speech sounds from a VCV context is not representative of adults’ ability to identify sounds given more extensive phonetic contexts. That is, it is possible that for some consonant tokens, recognition requires phonetic information that goes beyond the immediately adjacent vowels. In Experiment 2 , listeners heard the sounds within the context of the entire utterance, but with this context (outside the VCV) low-pass filtered to restrict access to lexical content. This did not improve performance. But it could be that the phonetic context required by adults to identify consonants both extends beyond the immediately adjacent vowels, and is not captured in low-pass-filtered speech. Although it is not clear how to characterize this phonetic information precisely, there is a substantial literature showing, for example, measurable influences of one consonant on another, over the span of a vowel (e.g., Sawusch & Newman, 2000 ; for a general review, Shattuck-Hufnagel, 2014 ). At present, very little is known about infants’ ability to use this kind of broadly distributed information to recover phonological structure. Our ability to explain how children learn language depends on quantitative characterization of child’s language environment. We cannot explain the problem children are solving if we do not specify the problem in terms of elements the child plausibly has access to ( Dupoux, 2018 ). At the same time, it is also useful to characterize the language environment in terms of the descriptive vocabulary of phonology—even if to reveal that the connection between the physical reality of the acoustic signal is frequently difficult to connect to the linguistic labels. Children do learn the canonical forms of spoken words. To achieve that goal, they will need to separate wheat from chaff. Both are present, but the chaff may be more plentiful than we thought. Supplementary Material 1 NIHMS2163362-supplement-1.pdf (485.9KB, pdf) Acknowledgments This work was supported by the National Science Foundation grants 1917608 and 2444175, and uses corpus resources that were created under National Institutes of Health grant R01-HD049681 to D. Swingley. Portions of the research reported here were presented at the Workshop on Infant Speech Perception (Donostia) in 2022 and at the Boston University Conference on Language Development in 2024. Appendix A. Supplementary data Supplementary material related to this article can be found online at https://doi.org/10.1016/j.cogpsych.2025.101766 . Footnotes Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. 1 We make the simplifying assumption that correct answers were not guesses. 2 Some responses matched in all three features but were still incorrect, because of the coarse granularity of the feature sets. For example, /s/ and /∫/ are unvoiced coronal fricatives (v,p,m) but are not the same sound. 3 The analysis was carried out using the glmer function in R, with the formula percent_correct ~ experiment * syllable_position + (1 |phone), family=“binomial”. Data availability Trial-level data and stimuli at author’s OSF site. References Adriaans F, & Swingley D (2017). Prosodic exaggeration within infant-directed speech: Consequences for vowel learnability. Journal of the Acoustical Society of America, 141, 3070–3078. 10.1121/1.4982246. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Antetomaso S, Miyazawa K, Feldman N, Elsner M, Hitczenko K, & Mazuka R (2017). Modeling phonetic category learning from natural acoustic data. In Proceedings of the annual boston university conference on language development. [ Google Scholar ] Aslin RN, Saffran JR, & Newport EL (1998). Computation of conditional probability statistics by 8-month-old infants. Psychological Science, 9, 321–324. 10.1111/1467-9280.00063. [ DOI ] [ Google Scholar ] Baer T, Moore BC, & Kluk K (2002). Effects of low pass filtering on the intelligibility of speech in noise for people with and without dead regions at high frequencies. The Journal of the Acoustical Society of America, 112(3), 1133–1144. 10.1121/1.1498853. [ DOI ] [ PubMed ] [ Google Scholar ] Bard EG, & Anderson AH (1983). The unintelligibility of speech to children. Journal of Child Language, 10, 265–292. [ DOI ] [ PubMed ] [ Google Scholar ] Batchelder EO (2002). Bootstrapping the lexicon: A computational model of infant speech segmentation. Cognition, 83(2), 167–206. 10.1016/S0010-0277(02)00002-1, https://www.sciencedirect.com/science/article/pii/S0010027702000021 . [ DOI ] [ PubMed ] [ Google Scholar ] Beckman ME, & Edwards J (2000). The ontogeny of phonological categories and the primacy of lexical learning in linguistic development. Child Development, 71, 240–249. 10.1111/1467-8624.00139. [ DOI ] [ PubMed ] [ Google Scholar ] Beech C, & Swingley D (2023). Consequences of phonological variation for algorithmic word segmentation. Cognition, 235, Article 105401. 10.1016/j.cognition.2023.105401. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Beech C, & Swingley D (2024). Relating referential clarity and phonetic clarity in infant-directed speech. Developmental Science, 27, 334–341. 10.1111/desc.13442. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bernard M, Thiolliere R, Saksida A, Loukatou GR, Larsen E, Johnson M, .… Cristia A (2020). WordSeg: Standardizing unsupervised word form segmentation from textardizing unsupervised word form segmentation from text. Behavior Research Methods, 52, 264–278. [ DOI ] [ PubMed ] [ Google Scholar ] Bernstein-Ratner N (1984). Patterns of vowel modification in mother-child speech. Journal of Child Language, 11, 557–578. [ PubMed ] [ Google Scholar ] Best CT, McRoberts GW, & Goodell E (2001). Discrimination of non-native consonant contrasts varying in perceptual assimilation to the listener’s native phonological system. The Journal of the Acoustical Society of America, 109(2), 775–794. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Bion RA, Miyazawa K, Kikuchi H, & Mazuka R (2013). Learning phonemic vowel length from naturalistic recordings of Japanese infant-directed speech. PLoS One, 8(2), Article e51594. 10.1371/journal.pone.0051594. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Brent MR, & Cartwright TA (1996). Distributional regularity and phonotactic constraints are useful for segmentation. Cognition, 61, 93–125. [ DOI ] [ PubMed ] [ Google Scholar ] Brent MR, & Siskind JM (2001). The role of exposure to isolated words in early vocabulary development. Cognition, 81, B33–B44. 10.1016/s0010-0277(01)00122-6. [ DOI ] [ PubMed ] [ Google Scholar ] Brysbaert M, Stevens M, Mandera P, & Keuleers E (2016). How many words do we know? Practical estimates of vocabulary size dependent on word definition, the degree of language input and the participant’s age. Frontiers in Psychology, 7, 1116. 10.3389/fpsyg.2016.01116. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Buckler H, & Goy EK (2018). What infant-directed speech tells us about the development of compensation for assimilation. Journal of Phonetics, 66, 45–62. 10.1016/j.wocn.2017.09.004. [ DOI ] [ Google Scholar ] Cabiddu F, Bott L, & Jones C (2023). CLASSIC utterance boundary: A chunking-based model of early naturalistic word segmentation. Language Learning, 73(3), 942–975. 10.1111/lang.12559. [ DOI ] [ Google Scholar ] Cairns P, Shillcock R, Chater N, & Levy J (1997). Bootstrapping word boundaries: a bottom-up corpus-based approach to segmentation. Cognitive Psychology, 33, 111–153. [ DOI ] [ PubMed ] [ Google Scholar ] Choi JY, & Perrachione TK (2019). Time and information in perceptual adaptation to speech. Cognition, 192, Article 103982. 10.1016/j.cognition.2019.05.019. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Cox C, Bergmann C, Fowler E, Keren-Portnoy T, Roepstorff A, Bryant G, et al. (2023). A systematic review and Bayesian meta-analysis of the acoustic features of infant-directed speech. Nature Human Behaviour, 7(1), 114–133. 10.1038/s41562-022-01452-1. [ DOI ] [ PubMed ] [ Google Scholar ] Cristia A, & Seidl A (2014). The hyperarticulation hypothesis of infant-directed speech. Journal of Child Language, 41(4), 913–934. [ DOI ] [ PubMed ] [ Google Scholar ] Cristia A, Seidl A, Junge C, Soderstrom M, & Hagoort P (2014). Predicting individual variation in language from infant speech perception measures. Child Development, 85(4), 1330–1345. 10.1111/cdev.12193. [ DOI ] [ PubMed ] [ Google Scholar ] Cutler A, Weber A, Smits R, & Cooper N (2004). Patterns of english phoneme confusions by native and non-native listeners. The Journal of the Acoustical Society of America, 116(6), 3668–3678. 10.1121/1.1810292. [ DOI ] [ PubMed ] [ Google Scholar ] Daland R, & Pierrehumbert JB (2011). Learning diphone-based segmentation. Cognitive Science, 35(1), 119–155. 10.1111/j.1551-6709.2010.01160.x. [ DOI ] [ PubMed ] [ Google Scholar ] de Seyssel M, Lavechin M, & Dupoux E (2023). Realistic and broad-scope learning simulations: first results and challenges. Journal of Child Language, 50, 1294–1317. 10.1017/S0305000923000272. [ DOI ] [ PubMed ] [ Google Scholar ] Delattre PC, Liberman AM, & Cooper FS (1955). Acoustic loci and transitional cues for consonants. The Journal of the Acoustical Society of America, 27(4), 769–773. [ Google Scholar ] Dietrich C, Swingley D, & Werker JF (2007). Native language governs interpretation of salient speech sound differences at 18 months. Proceedings of the National Academy of Sciences of the USA, 104, 16027–16031. 10.1073/pnas.0705270104. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Dilley LC, Millett AL, McAuley JD, & Bergeson TR (2014). Phonetic variation in consonants in infant-directed and adult-directed speech: The case of regressive place assimilation in word-final alveolar stops. Journal of Child Language, 41(1), 155–175. 10.1017/S0305000912000670. [ DOI ] [ PubMed ] [ Google Scholar ] Dupoux E (2018). Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner. Cognition, 173, 43–59. 10.1016/j.cognition.2017.11.008. [ DOI ] [ PubMed ] [ Google Scholar ] Eaves BS Jr., Feldman NH, Griffiths TL, & Shafto P (2016). Infant-directed speech is consistent with teaching. Psychological Review, 123(6), 758. 10.1037/rev0000031. [ DOI ] [ PubMed ] [ Google Scholar ] Eimas PD, & Miller JL (1980). Contextual effects in infant speech perception. Science, 209, 1140–1141. [ DOI ] [ PubMed ] [ Google Scholar ] Eimas PD, Siqueland ER, Jusczyk PW, & Vigorito J (1971). Speech perception in infants. Science, 171, 303–306. [ DOI ] [ PubMed ] [ Google Scholar ] Feldman NH, Griffiths TL, Goldwater S, & Morgan JL (2013). A role for the developing lexicon in phonetic category acquisition. Psychological Review, 120, 751–778. 10.1037/a0034245. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Ferguson CA (1964). Baby talk in six languages. American Anthropologist, 66, 103–114. [ Google Scholar ] Gogate LJ, & Bahrick LE (2001). Intersensory redundancy and 7-month-old infants’ memory for arbitrary syllable-object relations. Infancy, 2, 219–231. [ Google Scholar ] Hartman KM, Ratner NB, & Newman RS (2017). Infant-directed speech (IDS) vowel clarity and child language outcomes. Journal of Child Language, 44, 1140–1162. 10.1017/S0305000916000520. [ DOI ] [ PubMed ] [ Google Scholar ] Hitczenko K, & Feldman NH (2022). Naturalistic speech supports distributional learning across contexts. Proceedings of the National Academy of Sciences, 119, Article e2123230119. 10.1073/pnas.2123230119. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Jusczyk PW, & Aslin RN (1995). Infants’ detection of the sound patterns of words in fluent speech. Cognitive Psychology, 29, 1–23. [ DOI ] [ PubMed ] [ Google Scholar ] Kalashnikova M, & Burnham D (2018). Infant-directed speech from seven to nineteen months has similar acoustic properties but different functions. Journal of Child Language, 45(5), 1035–1053. 10.1017/50305000917000629. [ DOI ] [ PubMed ] [ Google Scholar ] Kalashnikova M, & Carreiras M (2022). Input quality and speech perception development in bilingual infants’ first year of life. Child Development, 93(1), e32–e46. 10.1111/cdev.13686. [ DOI ] [ PubMed ] [ Google Scholar ] Kalashnikova M, Goswami U, & Burnham D (2020). Novel word learning deficits in infants at family risk for dyslexia. Dyslexia, 26(1), 3–17. 10.1002/dys.1649. [ DOI ] [ PubMed ] [ Google Scholar ] Kemp N, Scott J, Bernhardt BM, Johnson CE, Siegel LS, & Werker JF (2017). Minimal pair word learning and vocabulary size: Links with later language skills. Applied Psycholinguistics, 38, 289–314. 10.1017/S0142716416000199. [ DOI ] [ Google Scholar ] Khlystova EA, Chong AJ, & Sundara M (2023). Phonetic variation in english infant-directed speech: A large-scale corpus analysis. Journal of Phonetics, 100, Article 101267. 10.1016/j.wocn.2023.101267. [ DOI ] [ Google Scholar ] Khorrami O,K, & Räsänen (2025). A model of early word acquisition based on realistic-scale audiovisual naming events. Speech Communication, 167, Article 103169. 10.1016/j.specom.2024.103169. [ DOI ] [ Google Scholar ] Kleinschmidt DF, & Jaeger TF (2015). Robust speech perception: recognize the familiar, generalize to the similar, and adapt to the novel. Psychological Review, 122(2), 148. 10.1037/a0038695. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Kuhl PK (1985). Categorization of speech by infants. In Mehler J, & Fox R (Eds.), Neonate cognition: beyond the blooming, buzzing confusion (pp. 231–262). Hillsdale, NJ: Erlbaum. [ Google Scholar ] Kuhl PK, Andruski JE, Chistovich IA, Chistovich LA, Kozhevnikova EV, Ryskina VL, .… Lacerda F (1997). Cross-language analysis of phonetic units in language addressed to infants. Science, 277(5326), 684–686. [ DOI ] [ PubMed ] [ Google Scholar ] Kuhl PK, Conboy BT, Coffey-Corina S, Padden D, Rivera-Gaxiola M, & Nelson T (2008). Phonetic learning as a pathway to language: new data and native language magnet theory expanded (NLM-eed (NLM-e. Philosophical Transactions of the Royal Society of London, B, 363, 979–1000. 10.1098/rstb.2007.2154. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Kuhl PK, Stevens E, Hayashi A, Deguchi T, Kiritani S, & Iverson P (2006). Infants show a facilitation effect for native language phonetic perception between 6 and 12 months. Developmental Science, 9, F13–F21. 10.1111/j.1467-7687.2006.00468.x. [ DOI ] [ PubMed ] [ Google Scholar ] Kuhl PK, Williams KA, Lacerda F, Stevens KN, & Lindblom B (1992). Linguistic experience alters phonetic perception in infants by 6 months of age. Science, 255, 606–608. [ DOI ] [ PubMed ] [ Google Scholar ] Ladefoged P, & Broadbent DE (1957). Information conveyed by vowels. Journal of the Acoustical Society of America, 29, 98–104. 10.1121/1.1908694. [ DOI ] [ PubMed ] [ Google Scholar ] Lahey M, & Ernestus M (2014). Pronunciation variation in infant-directed speech: Phonetic reduction of two highly frequent words. Language Learning and Development, 10(4), 308–327. [ Google Scholar ] Lavechin M, de Seyssel M, Métais M, Metze F, Mohamed A, Bredin H, .… Cristia A (2024). Modeling early phonetic acquisition from child-centered audio data. Cognition, 245, Article 105734. 10.1016/j.cognition.2024.105734. [ DOI ] [ PubMed ] [ Google Scholar ] Lindblom B (1992). Phonological units as adaptive emergents of lexical development. In Ferguson CA, Menn L, & Stoel-Gammon C (Eds.), Phonological development: models, research, implications (pp. 131–163). Timmonium, MD: York Press. [ Google Scholar ] Liu HM, Kuhl PK, & Tsao FM (2003). An association between mothers’ speech clarity and infants’ speech discrimination skills. Developmental Science, 6, F1–F10. [ Google Scholar ] Lovcevic I, Kalashnikova M, & Burnham D (2020). Acoustic features of infant-directed speech to infants with hearing loss. The Journal of the Acoustical Society of America, 148(6), 3399–3416. 10.1121/10.0002641. [ DOI ] [ PubMed ] [ Google Scholar ] Ludusan B, Mazuka R, & Dupoux E (2021). Does infant-directed speech help phonetic learning? A machine learning investigation. Cognitive Science, 45, Article e12946. 10.1111/cogs.12946. [ DOI ] [ PubMed ] [ Google Scholar ] MacWhinney B (2000). The CHILDES project: tools for analyzing talk project: tools for analyzing talk (3rd). Hillsdale, NJ: Erlbaum. [ Google Scholar ] Mann VA, & Repp BH (1980). Influence of vocalic context on perception of the [S]-[s] distinction. Perception & Psychophysics, 28(3), 213–228. [ DOI ] [ PubMed ] [ Google Scholar ] Martin A, Schatz T, Versteegh M, Miyazawa K, Mazuka R, Dupoux E, et al. (2015). Mothers speak less clearly to infants than to adults: A comprehensive test of the hyperarticulation hypothesis. Psychological Science, 26(3), 341–347. 10.1177/0956797614562453. [ DOI ] [ PubMed ] [ Google Scholar ] Mattys SL, & Jusczyk PW (2001). Do infants segment words or recurring contiguous patterns? Journal of Experimental Psychology: Human Perception and Performance, 27, 644–655. [ DOI ] [ PubMed ] [ Google Scholar ] Matusevych Y, Schatz T, Kamper H, Feldman N. m. H., & Goldwater S (2023). Infant phonetic learning as perceptual space learning: a crosslinguistic e valuation of computational models. Cognitive Science, 47(7), Article e13314. 10.1111/cogs.13314, https://onlinelibrary.wiley.com/doi/abs/10.1111/cogs.13314 . [ DOI ] [ PubMed ] [ Google Scholar ] McMurray B, Danelz A, Rigler H, & Seedorff M (2018). Speech categorization develops slowly through adolescence. Developmental Psychology, 54(1472), 10.1037/dev0000542. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Metsala JL, & Walley AC (1998). Spoken vocabulary growth and the segmental restructuring of lexical representations: precursors to phonemic awareness and early reading ability. In Metsala JL, & Ehri LC (Eds.), Word recognition in beginning literacy (pp. 89–120). Mahwah, NJ: LEA. [ Google Scholar ] Monaghan P, & Christiansen MH (2010). Words in puddles of sound: Modelling psycholinguistic effects in speech segmentation. Journal of Child Language, 37, 545–564. [ DOI ] [ PubMed ] [ Google Scholar ] Morton JB, & Trehub SE (2001). Children’s understanding of emotion in speeching of emotion in speech. Child Development, 72(3), 834–843. [ DOI ] [ PubMed ] [ Google Scholar ] Nittrouer S (2021). Speech perception by children: the structural refinement and differentiation model. In Pardo JS, Nygaard LC, Remez RE, & Pisoni DB (Eds.), The handbook of speech perceptionbook of speech perception (pp. 485–516). John Wiley and Sons, Ltd., 10.1002/9781119184096.ch18. [ DOI ] [ Google Scholar ] Pelucchi B, Hay JF, & Saffran JR (2009a). Learning in reverse: Eight-month-old infants track backward transitional probabilities. Cognition, 113, 244–247. 10.1016/j.cognition.2009.07.011. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Pelucchi B, Hay JF, & Saffran JR (2009b). Statistical learning in a natural language by 8-month-old infants. Child Development, 80, 674–685. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Perruchet P, & Vinter A (1998). Parser: A model for word segmentation. Journal of Memory and Language, 39, 246–263. [ Google Scholar ] Pitt MA, Dilley L, Johnson K, Kiesling S, Raymond W, Hume E, et al. (2007). Buckeye corpus of conversational speech (2nd release). Columbus, OH: Department of Psychology, Ohio State University. [ Google Scholar ] Pollack I (1948). Effects of high pass and low pass filtering on the intelligibility of speech in noise. The Journal of the Acoustical Society of America, 20(3), 259–266. [ Google Scholar ] Pollack I, & Pickett J (1963). The intelligibility of excerpts from conversation. Language and Speech, 6(3), 165–171. [ Google Scholar ] Pollack I, & Pickett J (1964). Frequency importance function for isolated words and for conversation of female talkers. Language and Speech, 7(2), 71–75. [ Google Scholar ] Ramon-Casas M, Swingley D, Bosch L, & Sebastián-Gallés N (2009). Vowel categorization during word recognition in bilingual toddlers. Cognitive Psychology, 59, 96–121. 10.1016/j.cogpsych.2009.02.002. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Recasens D (2018). Coarticulation. In Oxford research encyclopedia of linguistics (pp. 1–20). Oxford University Press, 10.1093/acrefore/9780199384655.013.416. [ DOI ] [ Google Scholar ] Redford MA, & Diehl RL (1999). The relative perceptual distinctiveness of initial and final consonants in CVC syllables syllables. Journal of the Acoustical Society of America, 106, 1555–1565. [ DOI ] [ PubMed ] [ Google Scholar ] Sawusch JR, & Newman RS (2000). Perceptual normalization for speaking rate: effects of signal discontinuities. Perception and Psychophysics, 62, 285–300. [ DOI ] [ PubMed ] [ Google Scholar ] Schatz T, Feldman NH, Goldwater S, Cao XN, & Dupoux E (2021). Early phonetic learning without phonetic categories: Insights from large-scale simulations on realistic input. Proceedings of the National Academy of Sciences, 118(7), Article e2001844118. 10.1073/pnas.2001844118. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Shattuck-Hufnagel S (2014). Phrase-level phonological and phonetic phenomena. In Goldrick M (Ed.), The oxford handbook of language productionbook of language production (pp. 259–274). 10.1093/oxfordhb/9780199735471.013.003. [ DOI ] [ Google Scholar ] Shockey L, & Bond ZS (1980). Phonological processes in speech addressed to children. Phonetica, 37, 267–274. [ Google Scholar ] Song JY, Demuth K, & Morgan JL (2010). Effects of the acoustic properties of infant-directed speech on infant word recognition. Journal of the Acoustical Society of America, 128, 389–400. 10.1121/1.3419786. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Stager CL, & Werker JF (1997). Infants listen for more phonetic detail in speech perception than in word-learning tasks. Nature, 388, 381–382. [ DOI ] [ PubMed ] [ Google Scholar ] Stern DN, Spieker S, Barnett RK, & MacKain K (1983). The prosody of maternal speech: infant age and context related changes. Journal of Child Language, 10, 1–15. [ DOI ] [ PubMed ] [ Google Scholar ] Sundberg U, & Lacerda F (1999). Voice onset time in speech to infants and adults. Phonetica, 56(3–4), 186–199. [ Google Scholar ] Swingley D (1999). Conditional probability and word discovery: a corpus analysis of speech to infants. In Hahn M, & Stoness SC (Eds.), Proceedings of the 21st annual conference of the cognitive science society (pp. 724–729). Mahwah, NJ: LEA. [ Google Scholar ] Swingley D (2005). Statistical clustering and the contents of the infant vocabulary. Cognitive Psychology, 50, 86–132. 10.1016/j.cogpsych.2004.06.001. [ DOI ] [ PubMed ] [ Google Scholar ] Swingley D (2009). Contributions of infant word learning to language development. Philosophical Transactions of the Royal Society B, 364, 3617–3622. 10.1098/rstb.2009.0107. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Swingley D (2016). Two-year-olds interpret novel phonological neighbors as familiar words. Developmental Psychology, 52, 1011–1023. 10.1037/dev0000114. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Swingley D (2019). Learning phonology from surface distributions, considering dutch and english vowel duration. Language Learning and Development, 15(3), 199–216. 10.1080/15475441.2018.1562927. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Swingley D, & Alarcon C (2018). Lexical learning may contribute to phonetic learning in infants: A corpus analysis of maternal spanish. Cognitive Science, 42(5), 1618–1641. 10.1111/cogs.12620. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Swingley D, & Algayres R (2024). Computational modeling of the segmentation of sentence stimuli from an infant word-finding study. Cognitive Science, 48(3), Article e13427. 10.1111/cogs.13427. [ DOI ] [ PubMed ] [ Google Scholar ] Swingley D, & Aslin RN (2007). Lexical competition in young children’s word learning. Cognitive Psychology, 54, 99–132. 10.1016/j.cogpsych.2006.05.001. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Tincoff R, Seidl A, Buckley L, Wojcik C, & Cristia A (2019). Feeling the way to words: Parents’ speech and touch cues highlight word-to-world mappings of body parts. Language Learning and Development, 15(2), 103–125. 10.1080/15475441.2019.1533472. [ DOI ] [ Google Scholar ] Treiman R, & Cassar M (1997). Can children and adults focus on sound as opposed to spelling in a phoneme counting task? Developmental Psychology, 33, 771–780. [ DOI ] [ PubMed ] [ Google Scholar ] Tsuji S, & Cristia A (2014). Perceptual attunement in vowels: A meta-analysis. Developmental Psychobiology, 56(2), 179–191. 10.1002/dev.21179. [ DOI ] [ PubMed ] [ Google Scholar ] Vihman M, & Majorano M (2017). The role of geminates in infants’ early word production and word-form recognition. Journal of Child Language, 44, 158–184. [ DOI ] [ PubMed ] [ Google Scholar ] Werker JF, & Curtin S (2005). Primir: A developmental model of speech processing. Language Learning and Development, 1, 197–234. 10.1080/15475441.2005.9684216. [ DOI ] [ Google Scholar ] Werker JF, & Tees RC (1984). Cross-language speech perception: evidence for perceptual reorganization during the first year of life. Infant Behavior and Development, 7, 49–63. [ Google Scholar ] White KS, & Morgan JL (2008). Sub-segmental detail in early lexical representations. Journal of Memory and Language, 59, 114–132. 10.1016/j.jml.2008.03.001. [ DOI ] [ Google Scholar ] Yuan J, Liberman M, et al. (2008). Speaker identification on the SCOTUS corpus corpus. Journal of the Acoustical Society of America, 123(5), 3878. [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials 1 NIHMS2163362-supplement-1.pdf (485.9KB, pdf) Data Availability Statement Trial-level data and stimuli at author’s OSF site. ACTIONS View on publisher site PDF (2.9 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top