ConceptioArchivearXiv CS
arXiv CSopen access

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

It’s Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces Blade Frisch

[email protected] Michigan Technological University Houghton, Michigan, USA

arXiv:2606.24854v1 [cs.HC] 23 Jun 2026

Michelle Kinsella

[email protected] Oregon Health & Science University Portland, Oregon, USA

Will Wade

Dylan Gaines

[email protected] Smartbox Assistive Technology Ltd Bristol, UK

[email protected] Kennesaw State University Marietta, Georgia, USA

Betts Peters

Tamara Broderick

[email protected] Oregon Health & Science University Portland, Oregon, USA

[email protected] Massachusetts Institute of Technology Cambridge, Massachusetts, USA

Keith Vertanen

[email protected] Michigan Technological University Houghton, Michigan, USA

Abstract

1

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their systems. However, evaluating AI-powered AAC interfaces can be difficult. People are intersectional beings and current evaluation metrics can struggle to capture the multifaceted and nuanced desires people may have for their AAC. We explore the complicated nature of six AAC problem spaces, explore how AI might be used in these spaces, and suggest more robust methods of evaluation that take the intersectional nuances of people into account. We also discuss broader issues that arise across these problem spaces and how they could be addressed using our proposed evaluation methods.

Metrics are essential for evaluating software performance and user experience. In order to determine how well a piece of software is performing, we must find a way to measure that performance. In machine learning and natural language processing, quantitative benchmarks are the standard for validating model efficacy. However, when these and other artificial intelligence (AI)-based approaches are integrated into augmentative and alternative communication (AAC) systems, traditional efficiency metrics often fail to capture the complexity of a user’s requirements. Evaluating AI-powered AAC requires a combination of technical performance metrics and human-centric data (e.g., self-reported usability measurements [2], interviews [47], usability testing [8]). We argue that the evaluation of AI-powered AAC interfaces must move beyond the assumption that a user’s needs and abilities are static and singular. We also argue that these evaluations should not be governed solely by technical requirements and must include the needs and desires of AAC users. While technical performance metrics are necessary for system development, any finite set of technical evaluations is inherently incomplete. We propose that a richer, pluralistic set of evaluation methods can better capture misfits between an AAC user and their communication context. Our primary contributions are:

CCS Concepts • Human-centered computing → Accessibility design and evaluation methods.

Keywords augmentative and alternative communication, AAC, artificial intelligence, metrics, evaluation, accessibility ACM Reference Format: Blade Frisch, Will Wade, Dylan Gaines, Michelle Kinsella, Betts Peters, Tamara Broderick, and Keith Vertanen. 2026. It’s Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces. In Proceedings of Speech AI for All: The What, How, and Who of Measurement Workshop at the CHI Conference on Human Factors in Computing Systems (Speech AI for All Workshop at CHI). ACM, New York, NY, USA, 8 pages. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Speech AI for All: The What, How, and Who of Measurement Workshop at the CHI Conference on Human Factors in Computing Systems, Barcelona, Spain © 2026 Copyright held by the owner/author(s).

Introduction

• Six AAC design considerations — We analyze six design considerations to take into account when evaluating AAC interfaces: speed and accuracy, mental and physical effort, agency in identity presentation, adapting to communication contexts, turn-taking, and fluctuating physical ability. • Possible AI-powered features for AAC — We propose various ways AI might be leveraged to improve the efficacy and user satisfaction with AAC systems. • Multidimensional evaluation methods — We propose evaluation methods that combine technical system metrics with human-centered design research methods to capture a broader understanding of AAC users’ needs.

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

2

Positioning

Existing research on AI in assistive technology is often bifurcated. Research often focuses on optimizing predictive models for speed [38]. While technical performance metrics are useful for evaluating a model’s performance, they may not be measuring the whole story of what matters to AAC users. If they are the sole method used to evaluate a model, they can impose artificial limitations on AAC systems solely to satisfy technical requirements rather than to ensure the system is addressing what truly matters to AAC users. Such a focus is similar to the medical model of disability, which risks pathologizing a communication disability over supporting the unique identities and desires of disabled people [33, 67]. We must move beyond imposing solutions on AAC users and instead work with them to create new approaches. It is critical to engage with AAC users in the design and intervention process [40]. This should include learning about what is important to AAC users and how to measure an AAC system’s ability to support those priorities, rather than selecting measurements because they are convenient from a technical perspective. AAC users and advocates have cautioned against viewing AI and other technological advances as “cures” for disability [48]. While recent work has explored AI bias in domains like facial recognition [15] and negative outcomes that can come from introducing AI-powered systems [58], more work is needed to explore these issues in the AAC domain. This paper bridges these perspectives. We examine the validity of standard technical performance metrics through the lens of intersectionality and argue that AAC lacks a unified evaluation strategy that accounts for the dynamic nature of communication or the diverse needs and preferences of users. We propose evaluation methods that combine AAC users’ wants and needs with technical performance metrics.

3

Related Work

AAC users are not one easily-defined group. There are many different identities people have (e.g., race, gender, societal role, views, disability status), and a person exists at the intersection of these identities [52]. This intersectionality can influence how they interact with technology. Problems arise when disability is reduced to only one identity; instead, disability can, and should, be viewed as intersectional [33, 67]. Rosemarie Garland-Thomson uses this lens to describe disability as a “misfit”: a relational encounter between an individual’s body-mind and an unaccommodating environment [25]. AI models often attempt to simplify this intersectionality by optimizing for a single target identity. As demonstrated by Buolamwini [15], such simplification can cause models to perform poorly for users at the intersection of marginalized identities [15]. In the context of AAC, this leads to technoableism [49], where systems are optimized for a hypothetical standard user and ignore that disability is intersectional and that individuals have multiple identities. Technical performance metrics can provide useful data about individual components of an interface or model, but they must be part of a larger suite of tools, methods, and measurements to better capture the intersectional nature of the people who use AAC and AI.

Frisch et al.

AAC users desire agency in the communication process, retaining control over who they interact with and how they communicate with people [21]. People change their behavior based on a myriad of social factors [27], and being disabled can lead to being stigmatized in social situations [28]. Designing AAC interfaces without supporting the agency of AAC users to control their communication and their presentation of self can impact how people use AAC and how others perceive them. When designing AAC interfaces, it may be important to consider how users value their ability to respond in a timely manner during conversations to follow social norms of communication [5]. By definition, communication involves more than just one person. AAC users will be communicating with other people, and the characteristics of these conversations may affect how AAC users and their AAC systems are perceived by others [31]. Communication partners will bring pre-existing beliefs and biases to conversations with an AAC user. The partners’ opinions may be further shaped by the AAC strategies and access methods used by the AAC user [36]. These perceptions can affect the success of communication interactions and the effective participation of AAC users in a variety of life situations [43]. Familiar communication partners may also be a valuable source of personally- or situationally-relevant information, and some AAC users may desire the ability to leverage the real-time input of these partners to supplement existing word prediction models while maintaining the autonomy to make the final decision about word or phrase selection [20, 46]. AI-powered features have the potential to influence partner perceptions of AAC users positively (e.g., by supporting the user in producing messages that conform to the partner’s expectations for the conversation) or negatively (e.g., by giving the erroneous impression that message content is being controlled by the AI and not by the user). Understanding the perspectives of communication partners could help inform best practices for training and educating those partners about what AI can and cannot do, managing expectations, and encouraging partners to presume communicative competence.

4

Exploring the Who, What, and How of AAC Interface Design

This section explores six critical domains where AI can augment AAC. However, as established in Section 3, every technical intervention risks othering the user if evaluated through a narrow lens. Our goal is to move beyond the hypothetical standard user fallacy by proposing evaluation methods that account for individual identity and environmental misfits.

4.1

Communication Speed and Accuracy

Spoken communication can be very fast at over 150 words per minute (WPM). Text-based AAC interfaces can be substantially slower (e.g., 2 WPM [38]), especially for users who require alternative access methods or interfaces due to physical or sensory impairments. This large disparity in speech production rates can be disruptive to communication [4] and even lead communication partners to form negative perceptions about AAC users [11]. The speed of text production is not the only factor of communication to consider; users also want to accurately communicate their thoughts,

It’s Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

and different users will have different preferences concerning the speed-accuracy trade-off [19]. The accuracy of AAC output can be affected not only by input errors but also by incorrect inferences made by any AI that an AAC system might use. 4.1.1 What AI Might Do. AAC interfaces often predict a user’s upcoming text based on their previous text [24]. These interfaces may make letter predictions (e.g., [44, 62]) or word predictions (e.g., [55, 61]). Non-literate users might use phoneme prediction (e.g., [53, 54]) or symbol prediction (e.g., [26, 50]). Historically, such letter, word, phoneme, or symbol predictions have leveraged statistical n-gram language models. Recent results show that large language models (LLMs) based on neural networks can improve letter and word predictions [23]. Given LLMs’ ability to transfer their knowledge to new languages despite limited training data [30], they likely can offer improved predictions for phonetic and symbolic AAC systems as well. While predicting the next letter or word can be helpful, the “holy grail” of AAC interfaces has long been to predict entire utterances to speed up communication (e.g., [18, 51]). The real promise of LLMs is that they may be able to generate “big swing” predictions that predict multiple words or entire sentences. LLMs can not only leverage substantially more of a user’s previous writing, they can also integrate other contextual clues. For example, AI might glean such clues from listening to the conversations around the user [1, 66], integrating suggestions from their communication partner [45], using a camera to infer things about a user’s surrounding environment [35], or identifying who the user is speaking with [34]. Whether AAC users or the people around them would want such contextual monitoring is a topic worth future investigation. 4.1.2 Possible Ways To Evaluate. While metrics such as words per minute are commonly used to measure communication speed, they must also be balanced with some way of measuring accuracy (i.e., it is no good producing text quickly if it is not what the user wanted to say). Accuracy might be measured in different ways depending on the situation; for example, in some contexts, specific wording may be very important to the user, while in other contexts it may only matter if AI-generated text captures the essence of what the user wanted to say. Because the correctness of an utterance is often subjective and context-dependent, we propose a multidimensional approach: • Semantic and linguistic similarity (offline) — While traditional metrics like word error rate (WER) measure verbatim accuracy, they fail to account for AI predictions that capture the meaning but not the exact phrasing. This might be possible by utilizing LLM judges and semantic embeddings that better assess intent preservation than simple text matching metrics such as WER. • User-mediated assessment (online) — In interactive user studies, accuracy should be measured through the user’s perception of how well the produced text aligns with their intent. This includes tracking correction rates, noting when a user accepts a close-enough prediction versus when they feel compelled to edit it. This can reflect the pragmatic trade-off the user may make between speed and effort.

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

• Functional success (task-based) — Evaluation can be shifted from the text itself to the outcome of the social interaction. By using task-based scenarios, researchers can measure whether the user’s communication accomplished the task, such as communicating a specific idea or requesting an action. This can validate the system’s efficacy as a tool for participation rather than just a text entry interface. Measuring the user’s intent remains a significant methodological challenge. While lab settings often provide participants with predefined communicative goals, longitudinal contextual evaluations could have users review their own logs to rate the accuracy of AI-assisted utterances and how much agency they retained when composing text. In lab user studies, it is also common to measure entry and error rate by having participants copy fixed phrases. This provides an easy way to measure error rate, but differs from how people use AAC systems in practice (i.e., converting their own thoughts into text). While it makes error rate measurement more difficult, it is possible to allow participants to freely compose messages [60]. In the case of both text copy and free composition tasks, it is important to ensure an AAC system can input whatever a user desires, including difficult-to-predict words (e.g., proper names, acronyms, passwords). One approach to measuring this aspect of an AAC system is to have participants copy phrases with rare words [59] or ask participants to compose text they suspect will be difficult for the system [22].

4.2

Physical and Mental Effort

Some people with communication disabilities also have physical disabilities that make it difficult or impossible to use standard text entry devices, such as a keyboard or touchscreen. Instead, they may use alternative access methods that allow them to compose text using other interaction methods, such as eye gaze or switches. People who use AAC, and particularly those who use alternative access methods, may experience fatigue, eye strain, or other physical symptoms as a result of system use. Some access methods, such as switches, make it possible to control a computer with minimal movement but may require more than one user action per selection. For example, in the single-switch AAC system Nomon [13, 14], the user may need to activate a switch multiple times before the system is confident enough to select a specific character or word. In such systems, it becomes crucial to consider the required physical effort to enter text and balance it with other factors such as speed and accuracy. It also takes mental effort to process partner communications, navigate and use the AAC interface, and choose between future actions sequences (e.g., how to correct any previous errors). Correcting errors can be frustrating [3] and dominate a user’s time [6]. Some access methods, such as brain-computer interfaces or eye-tracking, require constant mental attention, with limited opportunities for the user to take a break. When an AAC interface requires several taxing actions, the user can be left both physically and mentally exhausted. 4.2.1 What AI Might Do. An AI may be able to detect when a user is experiencing increased physical or mental effort. This might be done via physiological sensing (e.g., over-the-ear EEG sensors integrated into smartglasses) or by observing a user’s previous

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

interactions. This information could be used to dynamically adapt the interface to reduce effort. For example, it may be better to use a more expensive language model that offers more accurate predictions, even if the added prediction latency slows input. Another option might be to increase the number of predictions displayed. This may increase the mental effort required to search through the presented predictions, but if the predictions are accurate, it could reduce the physical effort required to input that text. However, this may depend on a user’s particular preferences, abilities, and access method (e.g., in Nomon more prediction targets can increase the physical actions required per selection). We previously discussed the idea of an AI making “big swing” predictions (i.e., predicting multiple words or sentences at once). As a user becomes fatigued, they may increasingly accept such bigger predictions and the system could increasingly present them. While we already discussed how this can increase entry rates, it can also reduce the number of input actions required by the user to enter text, especially in interfaces such as Nomon and RSVP [41] where multiple user actions are required to make a selection. If the predictions are not accurate, it can increase the number of user actions required (and therefore the effort required) to input text or correct errors made by the system. The AI could dynamically detect when the user is engaging in more corrective actions and reduce the weight of the language model to prevent additional AI-induced errors. The system could also potentially detect a user’s mental response to an erroneous selection via error-related potentials [68] sensed via EEG and engage in automatic error correction. 4.2.2 Possible Ways To Evaluate. Quantifying the physical effort required for system use can depend on the specific interface and access method. For example, in Nomon, the authors measured how many switch activations (e.g., pressing a button, blinking using eye-gaze tracking, releasing a puff of air in a sip/puff switch) were required to select a target. Another option is the CARE Efficiency Score [65] which estimates the level of effort required to activate buttons based on motor distance and visual scanning. This allows us to evaluate if AI’s big swing predictions actually reduce the motor load or if the cognitive and physical costs of correcting a major error negate the potential gains. To quantify mental effort, it can be beneficial to assess user workload with self-reported questionnaires such as the NASA Task Load Index [29] or variations adapted for human-computer interaction and assistive technology applications [9, 42]. It may also be possible to use sensor-based techniques to estimate mental effort, such as by using eye-tracking [17] or EEG [69].

4.3

Sounding Like You Want to Sound

Speech-generating AAC can often fail to correctly express voice tones, accents, cadences, and other aspects of speech that shape how people communicate [36]. Judge and Townend [32] found that users often prioritize vocal quality and regional accents as a means of maintaining social presence. The inability of an AAC system to accurately replicate the nuances of natural speech can lead to AAC users being misunderstood and misrepresented by their means of communication, and may affect communication partners’ perception of the user.

Frisch et al.

4.3.1 What AI Might Do. AI can be used to generate a voice model for a person based on banked voice data [16]. If there is no such data available, AI could also be used to synthesize a voice based on existing speech samples from other people. This could even include blending different voice samples to achieve specific attributes, such as creating a specific accent. AI could also change aspects of the voice generated by text-to-speech based on tone indicators inserted by the user and by context clues, such as what and how the communication partner is communicating. These tone indicators could instruct the text-to-speech engine on where to include pauses or how to change the pitch of a word or phrase. It could also learn the tone indicators typically selected by the user for specific phrases or frequent conversation topics and suggest them when composing text, learning their unique communication style over time. 4.3.2 Possible Ways To Evaluate. Usability testing can be used to evaluate the tone indicators and the user’s satisfaction with both the indicators and the generated speech. Diary studies (i.e., where a person will use a piece of software for a length of time and write diary entries about their experiences using the software) can then be used to see how AAC users integrate these tools into their daily communication lives. We can also measure the impact of different tone indicators by conducting usability testing on each tone indicator individually and collecting data on user satisfaction with the indicator under test. Evaluation could include rating scales where users rate the selfidentification of the voice, assessing whether the output feels like an extension of themselves or not. Additionally, gathering satisfaction data from both AAC users and their communication partners on the generated speech can help determine if the AI-generated prosody successfully conveys the user’s intended emotion or social standing to their conversation partners [36].

4.4

Code- and Context-Switching

People will change how they communicate based on the details of the situation. They may switch between languages or dialects when communicating, which is called code-switching [57]. They may change other aspects of their communication (e.g., usage of slang, level of formality) based on context: who they’re communicating with, the number of communication partners, and the content being communicated. The environmental and social context surrounding the communication also matters: conversing in a loud coffee shop and sharing information in a doctor’s office have different communication needs and present different challenges to AAC users [10, 43]. Current AAC systems offer limited support for users in code- and context-switching to better adapt to different communication partners and situations (e.g., chatting with a friend, participating in a job interview, or sharing current symptoms with a doctor), or in switching between languages for multilingual users [36]. 4.4.1 What AI Might Do. AI systems can be trained to change their predictions based on who the AAC user is communicating with and the context in which the user is communicating. For example, there could be a mode to predict more friendly and casual text for chatting with friends, where the predictions are more concerned

It’s Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

with the user sounding like themselves [37] (see Section 4.3). Another mode might focus on the user’s formal, professional tone for contexts such as job interviews. Still another mode would be helpful when communicating with doctors, where predictions should focus more on accuracy in communication and helping the AAC user say exactly what they need to communicate to their doctor. AI could also detect when utterances, or portions of an utterance, are in a different language and adjust the pronunciation of the words based on the indicated language. 4.4.2 Possible Ways To Evaluate. Evaluating code- and contextswitching first requires understanding the different communication styles and contexts AAC users have. This can be done through semi-structured interviews, contextual inquiry, ethnography, and questionnaires like the Communication Needs Questionnaire [7]. Once the styles and contexts are understood, the success of the code switch can then be measured through methods like usability testing and user satisfaction measurements. The Communicative Participation Item Bank (CPIB) [10], a self-reported measure of the difficulty an individual experiences when participating in various communication-related situations, may also be useful for evaluating the success of code-switching to adapt to different communication contexts.

4.5

Taking Part Fluidly in a Conversation

Turn-taking is one aspect of interpersonal interactions where AAC users can struggle to communicate [43]. It takes significantly more time to compose messages on AAC devices than it does to speak them, often leading to the conversation moving on or changing topics before the user can contribute [36]. Judge and Townend [32] found that users identify the ability to be fast and spontaneous as a key design requirement, yet one that is frequently unmet, leading to a restricted use of the device in dynamic social settings. Weinberg et al. [63] show that AAC users can also struggle to make use of backchanneling (e.g., adding interjections like “hmm”, “yeah”, and “uh-huh”) when interacting with both other AAC users and nonAAC users. This can also include when AAC users try to capture the attention of their communication partners in order to “take the floor” and become the active speaker [56]. 4.5.1 What AI Might Do. AI can predict the backchanneling methods that the participant would like to use in the conversation and adjust the AAC system to make use of those methods. Additionally, the AAC system could include more efficient ways to predict takethe-floor utterances, such as “Can I chime in?” or “Oh yeah, that reminds me of something!”. Scripting interactions is a strategy that some AAC users make use of to prepare for social and communitybased interactions. AI can be used to help prepare scripts for these interactions so that utterances are readily available for an AAC user to use when communicating with others. 4.5.2 Possible Ways To Evaluate. Success in turn-taking cannot be measured by text accuracy alone; it can also be measured by latency to respond and floor-taking success rates. Quantitatively, we can measure the time difference between when the user wants

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

to take the floor and when the take-the-floor signal occurs. However, because communication is a relational act, this could be complemented by interviewing communication partners. These interviews could reveal whether the AI-guided backchanneling and floor-taking make the user appear more present in the exchange, or if the AI’s timing feels uncanny or inorganic to the partner. This will help guide AAC designers in ensuring the AAC user retains their agency in self-presentation. The CPIB includes items directly related to fast-moving conversations, communication in group settings, and other contexts involving backchanneling and turn-taking, and could be used to assess the impact of AI-based features on participation in those situations. This dual approach ensures that while the AI speeds up the mechanics of turn-taking, it maintains the authenticity of the user’s social presence.

4.6

Short- and Long-Term Needs Changes

Communication needs can be dynamic and change on short- and long-term scales. For example, autistic people can have an increased need for alternative methods of communication as they become overstimulated, which decreases as the stimuli are removed. People with ALS may experience short-term changes due to fatigue or the effects of medication, but because of the degenerative nature of the condition, there can be changes in physical function over a longer period of time. These changes in physical ability can change how people interact with their AAC system, such as changing from using a touch screen to using an eye tracker or brain-computer interface. A static AAC system might not be able to support short-term changes or be usable in the long-term as an individual’s disability progresses. This can create an additional burden on the user to continuously learn and adapt to a new system when their old system is no longer able to support their needs. In addition to accommodating users’ physical, cognitive, or emotional states, AAC systems should also be able to adapt to their changing communication needs and preferences. An individual will encounter different communication contexts, partners, and topics as they move through life, and will want to express themselves in different ways. 4.6.1 What AI Might Do. AAC systems can be designed with multiple interface options based on the different communication needs a person may have. For example, an AAC system designed to support autistic people who experience shutdown could have two ways of interacting with composed utterances: an interface that provides fine-grained controls for text-to-speech (e.g., speech rate controls, speaking only portions of the utterance at a time) and a simplified interface that relies on preset controls to speak an entire utterance. AI can then learn the user’s behavior and change from the detailed interface to the simplified interface if it detects that the autistic person is entering shutdown and needs less stimulation. It can also adjust its text predictions based on the detected communication needs, such as offering larger predictions when it detects that the user needs more communication support. As discussed in Sections 4.1 and 4.2, these “big swing” predictions can help increase communication speed and reduce the physical and mental effort of composing longer utterances. The size of these predictions and how the predictions are made could be adjusted based on the user’s needs. For example, people with ALS may prefer to make

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

less use of predicted text when their physical function is relatively strong (e.g., earlier in disease progression or earlier in the day). But they may benefit from additional predictions, including entire sentences or multiple sentences, when they are more fatigued or when they are using slower alternative access methods. 4.6.2 Possible Ways To Evaluate. Evaluating changing needs requires a longitudinal approach. While it may be necessary to use technical metrics to evaluate an AI model’s ability to detect the current needs of an AAC user, these cannot wholly capture how AAC users would respond to an adaptive system. Usability testing can show the initial reactions of AAC users to a dynamic interface, but it can be difficult to detect changes in an AAC user’s communication needs on the time scale of a usability test. Diary studies are a useful tool for measuring longitudinal data, which can help measure needs changes over time. An AAC user could be given an adaptive AAC system and be asked to write diary entries whenever the system detects a needs change and adapts the interface. The AAC user could then provide information on how the adaptation impacted their communication and whether they felt the adaptation should have been made. Combining these qualitative data with technical metrics on needs detection tells a more complete story of how an AI model can adapt an AAC interface based on changes in user needs.

5

Discussion and Limitations

The individual design challenges explored in Section 4 are underpinned by broader, systemic issues. As noted by Bennett et al. [12], designing data-driven technologies for long-term conditions often risks epistemic injustice, where technical performance metrics are preferred over a user’s lived experience. In AAC, a high text entry rate can mask a user’s loss of authentic voice or autonomy. This creates an issue: an AI system may be successful by the developer’s metrics while failing the user’s sense of self. Konadl [39] highlights that current generative solutions often fail to bridge the gap between formal and informal contexts, forcing one-size-fits-all output. This creates a misfit between an AAC user’s needs and what the AAC system supports [25], as the user is coerced into normative speech patterns to satisfy the model’s requirements. This can devalue their identity in favor of technically-driven evaluation. To address this issue, AAC evaluation could be situated within the Three Domains Framework proposed by Judge and Townend [32]. While AI development often focuses solely on the “Device Design” domain (speed and reliability), successful use is equally dependent on the “Wider Picture” (environmental support) and the “Personal Context” (user identity) domains. Clinical tools, such as the Individually Prioritised Problem Assessment (IPPA) [64] and the Communicative Participation Item Bank [10], can offer a potential starting point for uncovering the requirements across these three domains. By moving beyond simple satisfaction scores, these person-centered frameworks can help determine if AI features are fostering genuine communication growth. Furthermore, as interfaces become more proactive, offering AIguided changes based on perceived context or fatigue, we face a significant UX challenge: providing granular autonomy within an interface that is already inherently difficult to access. As Judge and Townend [32] demonstrate, users identify a fundamental link

Frisch et al.

between communication speed and the dignity of the interaction. However, they also warn that simplicity of design remains a key requirement. AI interventions intended to enhance communication must not inadvertently increase the cognitive load or decrease system reliability. We propose several potential paths for how AI can be introduced into AAC interfaces in Section 4. However, these potential implementations have not yet been validated or tested with AAC users. As we discuss throughout this work, evaluating AI in AAC interfaces is complicated and should not be done with quantitative metrics alone. Any inclusion of AI in AAC interfaces should be done alongside AAC users via participatory design approaches. This will help ensure that the desires and needs of AAC users are included when evaluating an AI-powered AAC interface.

6

Conclusion

Technical performance metrics alone can struggle to collect intersectional data. It is nearly impossible to capture the many dimensions of one’s identity solely from a single, or even a set of, technical performance metrics. Instead, evaluating an AI-powered AAC interface must be done with a pluralistic set of evaluation methods guided by the wants and needs of AAC users. Both technical performance metrics and human-centered data are needed to tell the entire story of how an interface is performing. It is also critical to ensure that what is measured is guided by what AAC users decide is important. Trying to collapse people into a single identity and collect only technical performance metrics from this narrowed viewpoint is a path that leads to technoableism. Researching humans is complicated. We have explored this complexity across six AAC problem spaces and posited that people do not fit into neat boxes. Technical performance metrics for evaluating machine learning and natural language processing models can often fail to capture all of what AAC users value in their AAC systems. Just as there is great nuance in understanding these technical performance metrics and applying them correctly, there is nuance in connecting metrics to the broader user experience. Technical performance metrics must be only one star in the constellation of evaluation so that the entire story of a user’s experience with their AAC system can be told.

Acknowledgments This work was funded in part by the National Science Foundation (IIS-2402876, IIS-2402877, and IIS-2402878).

References [1] Jiban Adhikary, Robbie Watling, Crystal Fletcher, Alex Stanage, and Keith Vertanen. 2019. Investigating Speech Recognition for Improving Predictive AAC. In SLPAT ’19: Proceedings of the Workshop on Speech and Language Processing for Assistive Technologies (Minneapolis, MN). 37–43. [2] Bill Albert and Tom Tullis. 2023. Measuring the User Experience: Collecting, Analyzing, and Presenting UX Metrics (3rd ed.). Morgan Kaufmann Publishers, Cambridge, MA. [3] Ohoud Alharbi and Wolfgang Stuerzlinger. 2022. Auto-Cucumber: The Impact of Autocorrection Failures on Users’ Frustration. In Graphics Interface 2022. https: //openreview.net/forum?id=dcbsb4qTmnt [4] Norman Alm, John L. Arnott, and Alan F. Newell. 1992. Prediction and conversational momentum in an augmentative communication system. Commun. ACM 35, 5 (1992), 46–57. [5] American Speech-Language-Hearing Association. [n. d.]. Components of Social Communication. https://www.asha.org/practice-portal/clinical-topics/social-

It’s Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

communication-disorder/components-of-social-communication/ [6] Shiri Azenkot and Nicole B. Lee. 2013. Exploring the Use of Speech Input by Blind People on Mobile Devices. In Proceedings of the 15th International ACM SIGACCESS Conference on Computers and Accessibility (Bellevue, Washington) (ASSETS ’13). Association for Computing Machinery, New York, NY, USA, Article 11, 8 pages. doi:10.1145/2513383.2513440 [7] Lisa G. Bardach. 2017. Communication Needs Questionnaire. Boston Children’s Hospital Augmentative Communication Program. https: //www.childrenshospital.org/sites/default/files/2022-03/communicationneeds-questionnaire.pdf [8] Carol M. Barnum. 2021. Usability Testing Essentials: Ready, Set ...Test! (2nd ed.). Morgan Kaufmann Publishers, Cambridge, MA. [9] Richard EA Bates. 2006. Enhancing the performance of eye and head mice: a validated assessment method and an investigation into the performance of eye and head based assistive technology pointing devices. Ph. D. Dissertation. De Montfort University. [10] Carolyn Baylor, Kathryn Yorkston, Tanya Eadie, Jiseon Kim, Hyewon Chung, and Dagmar Amtmann. 2013. The Communicative Participation Item Bank (CPIB): Item bank calibration and development of a disorder-generic short form. (2013). [11] Ann Beck, Heidi Fritz, Allison Keller, and Marcia Dennis. 2000. Attitudes of school-aged children toward their peers who use augmentative and alternative communication. Augmentative and Alternative Communication 16, 1 (2000), 13– 26. [12] S J Bennett, Caroline Claisse, Ewa Luger, and Abigail C Durrant. 2023. Unpicking epistemic injustices in digital health: On the implications of designing data-driven technologies for the management of long-term conditions. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. ACM, New York, NY, USA, 322–332. [13] Nicholas Bonaker, Emli-Mari Nel, Keith Vertanen, and Tamara Broderick. 2023. A Usability Study of Nomon: A Flexible Interface for Single-Switch Users. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility (New York, NY, USA) (ASSETS ’23). Association for Computing Machinery, New York, NY, USA, Article 3, 17 pages. doi:10.1145/3597638.3608415 [14] Nicholas Ryan Bonaker, Emli-Mari Nel, Keith Vertanen, and Tamara Broderick. 2022. A Performance Evaluation of Nomon: A Flexible Interface for Noisy SingleSwitch Users. In CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 495, 17 pages. doi:10.1145/3491102.3517738 [15] Joy Buolamwini. 2024. Unmasking AI: My Mission to Protect What Is Human in a World of Machines. Random House Trade Paperbacks, New York, NY. [16] Mo Chen, Jolene Hyppa-Martin, H. Timothy Bunnell, Jason Lilley, Celestine Foo, Han Wei Tan, and Wei Shun Lim. 2023. Voice banking to support individuals who use speech-generating devices: development and evaluation of Singaporeanaccented English synthetic voices and a Singapore Colloquial English recording inventory. Augmentative and Alternative Communication 39, 4 (2023), 208–218. arXiv:https://doi.org/10.1080/07434618.2023.2181213 doi:10.1080/07434618.2023. 2181213 PMID: 36971387. [17] Siyuan Chen, Julien Epps, Natalie Ruiz, and Fang Chen. 2011. Eye activity as a measure of human mental effort in HCI. In Proceedings of the 16th International Conference on Intelligent User Interfaces (Palo Alto, CA, USA) (IUI ’11). Association for Computing Machinery, New York, NY, USA, 315–318. doi:10.1145/1943403. 1943454 [18] Patrick W Demasco and Kathleen F McCoy. 1992. Generating text from compressed input: An intelligent interface for people with severe motor impairments. Commun. ACM 35, 5 (1992), 68–78. [19] Melanie Fried-Oken, Michelle Kinsella, Ian Stevens, and Eran Klein. 2024. What stakeholders with neurodegenerative conditions value about speed and accuracy in development of BCI systems for communication. Brain-Computer Interfaces 11, 1-2 (2024), 21–32. [20] Melanie Fried-Oken, Michelle A. Kinsella, Erik Jakobs, Tom Jakobs, Aimee Mooney, Betts Peters, Rebecca Pryor, and Scott Spaulding. 2025. Smart Predict: adding partner-suggested vocabulary to increase efficiency in a dual tablet AAC typing application. Augmentative and Alternative Communication 41, 4 (2025), 395–406. arXiv:https://doi.org/10.1080/07434618.2024.2374314 doi:10.1080/07434618.2024.2374314 PMID: 39164980. [21] Blade Frisch and Keith Vertanen. 2025. Designing AAC for use in social and community contexts: a scoping review. Augmentative and Alternative Communication (Oct. 2025), 1–11. doi:10.1080/07434618.2025.2558851 _eprint: https://doi.org/10.1080/07434618.2025.2558851. [22] Dylan Gaines, Per Ola Kristensson, and Keith Vertanen. 2021. Enhancing the composition task in text entry studies: Eliciting difficult text and improving error rate calculation. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 725, 8 pages. doi:10.1145/3411764.3445199 [23] Dylan Gaines and Keith Vertanen. 2025. Adapting Large Language Models for Character-based Augmentative and Alternative Communication. In Findings of the Association for Computational Linguistics: EMNLP 2025. Association for

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

Computational Linguistics, Suzhou, China, 15273–15291. doi:10.18653/v1/2025. findings-emnlp.826 [24] Nestor Garay-Vitoria and Julio Abascal. 2006. Text prediction systems: a survey. Universal Access in the Information Society 4, 3 (2006), 188–203. Issue 3. [25] Rosemarie Garland-Thomson. 2011. Misfits: A Feminist Materialist Disability Concept. Hypatia 26 (06 2011), 591 – 609. doi:10.1111/j.1527-2001.2011.01206.x [26] Nicola Gatti and Matteo Matteucci. 2006. CABA2L a Bliss predictive composition assistant for AAC communication software. In Enterprise Information Systems VI. Springer, 277–284. [27] Erving. Goffman. 1959. The Presentation of Self in Everyday Life. Anchor Books, New York, NY, USA. [28] Erving. Goffman. 1963. Stigma: Notes on the Management of Spoiled Identity. Simon & Schuster Inc, New York, New York. [29] Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. Vol. 52. Elsevier, 139–183. [30] Hansi Hettiarachchi, Tharindu Ranasinghe, Paul Rayson, Ruslan Mitkov, Mohamed Gaber, Damith Premasiri, Fiona Anting Tan, and Lasitha Uyangodage (Eds.). 2025. Proceedings of the First Workshop on Language Models for Low-Resource Languages. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates. https://aclanthology.org/2025.loreslm-1.0/ [31] Linda A. Hoag, Jan L. Bedrosian, Kathleen F. MCcoy, and Dallas E. Johnson. 2008. Hierarchy of Conversational Rule Violations Involving Utterance-Based Augmentative and Alternative Communication Systems. Augmentative and Alternative Communication 24, 2 (Jan. 2008), 149–161. doi:10.1080/07434610802038288 _eprint: https://doi.org/10.1080/07434610802038288. [32] Simon Judge and Gillian Townend. 2013. Perceptions of the design of voice output communication aids. Int. J. Lang. Commun. Disord. 48, 4 (April 2013), 366–381. [33] Alison Kafer. 2013. Feminist, Queer, Crip. Indiana University Press, Bloomington. https://research.ebsco.com/linkprocessor/plink?id=16af4027-107a-3aedb257-fad97d457a08 [34] Shaun K. Kane, Barbara Linam-Church, Kyle Althoff, and Denise McCall. 2012. What we talk about: designing a context-aware communication tool for people with aphasia. In Proceedings of the 14th International ACM SIGACCESS Conference on Computers and Accessibility (Boulder, Colorado, USA) (ASSETS ’12). Association for Computing Machinery, New York, NY, USA, 49–56. doi:10.1145/2384916. 2384926 [35] Shaun K Kane and Meredith Ringel Morris. 2017. Let’s Talk about X: Combining image recognition and eye gaze to support conversation for people with ALS. In Proceedings of the 2017 Conference on Designing Interactive Systems. 129–134. [36] Shaun K. Kane, Meredith Ringel Morris, Ann Paradiso, and Jon Campbell. 2017. "At times avuncular and cantankerous, with the reflexes of a mongoose": Understanding Self-Expression through Augmentative and Alternative Communication Devices. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing (CSCW ’17). Association for Computing Machinery, New York, NY, USA, 1166–1179. doi:10.1145/2998181.2998284 [37] Eran Klein, Michelle Kinsella, Ian Stevens, and Melanie Fried-Oken. 2024. Ethical issues raised by incorporating personalized language models into brain-computer interface communication technologies: a qualitative study of individuals with neurological disease. Disability and Rehabilitation: Assistive Technology 19, 3 (2024), 1041–1051. arXiv:https://doi.org/10.1080/17483107.2022.2146217 doi:10. 1080/17483107.2022.2146217 PMID: 36403143. [38] Heidi Horstmann Koester and Sajay Arthanat. 2018. Text entry rate of access interfaces used by people with physical disabilities: A systematic review. Assistive Technology 30, 3 (2018), 151–163. [39] Daniel Konadl. 2024. A Generative AI-based approach to support automated utterance generation for different conversational contexts within AAC systems. [40] Janice Light and David McNaughton. 2013. Putting People First: ReThinking the Role of Technology in Augmentative and Alternative Communication Intervention. Augmentative and Alternative Communication 29, 4 (Dec. 2013), 299–309. doi:10.3109/07434618.2013.848935 _eprint: https://doi.org/10.3109/07434618.2013.848935. [41] Umut Orhan, Kenneth E. Hild, Deniz Erdogmus, Brian Roark, Barry Oken, and Melanie Fried-Oken. 2012. RSVP Keyboard: An EEG Based Typing Interface. Proceedings of the 2012 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) (2012), 10.1109/ICASSP.2012.6287966. doi:10.1109/ ICASSP.2012.6287966 [42] Betts Peters, Aimee Mooney, Barry Oken, and Melanie Fried-Oken. 2016. Soliciting BCI user experience feedback from people with severe speech and physical impairments. Brain-Computer Interfaces 3, 1 (2016), 47–58. [43] Betts Peters, Jack Wiedrick, and Carolyn Baylor. 2023. Effects of aided communication on communicative participation for people with amyotrophic lateral sclerosis. American journal of speech-language pathology 32, 4 (2023), 1450–1465. [44] Brian Roark, Jacques De Villiers, Christopher Gibbons, and Melanie Fried-Oken. 2010. Scanning methods and language modeling for binary switch typing. In Proceedings of the NAACL HLT 2010 Workshop on Speech and Language Processing for Assistive Technologies. 28–36.

Speech AI for All Workshop at CHI, April 16, 2026, Barcelona, Spain

[45] Brian Roark, Andrew Fowler, Richard Sproat, Christopher Gibbons, and Melanie Fried-Oken. 2011. Towards technology-assisted co-construction with communication partners. In Proceedings of the second workshop on speech and language processing for assistive technologies. 22–31. Huff[46] Brian Roark, Melanie Fried-Oken, and Chris Gibbons. 2015. man and Linear Scanning Methods with Statistical Language Models. Augmentative and Alternative Communication 31, 1 (2015), 37–50. arXiv:https://doi.org/10.3109/07434618.2014.997890 doi:10.3109/07434618.2014. 997890 PMID: 25672825. [47] Irving Seidman. 2019. Interviewing as Qualitative Research: A Guide for Researchers in Education and the Social Sciences (5th ed.). Teachers College Press, New York, NY. [48] Darryl Sellwood, Lateef McLeod, Kevin Williams, Katie Brown, and Graham Pullin. 2024. Imagining alternative futures with augmentative and alternative communication: a manifesto. Medical Humanities 50, 4 (2024), 620–623. [49] Ashley Shew. 2023. Against Technoableism: Rethinking Who Needs Improvement. W. W. Norton, New York, New York. [50] Hugh Stewart and Ann Wilcock. 2000. Improving the communication rate for symbol based, scanning voice output device users. Technology and Disability 13, 3 (2000), 141–150. [51] John Todman, Norman Alm, Jeff Higginbotham, and Portia File. 2008. Whole utterance approaches in AAC. Augmentative and alternative communication 24, 3 (2008), 235–254. [52] Sarah J Tracy. 2020. Qualitative Research Methods: Collecting Evidence, Crafting Analysis, Communicating Impact (2nd ed.). John Wiley & Sons, Hoboken, NJ. [53] Ha Trinh, Annalu Waller, Rolf Black, and Ehud Reiter. 2010. Further Development of the PhonicStick: The application of phonic-based acceleration methods to the speaking joystick. In 14th Biennial Conference of the International Society of Augmentative and Alternative Communication: Communicating Worlds. [54] Ha Trinh, Annalu Waller, Keith Vertanen, Per Ola Kristensson, and Vicki L. Hanson. 2012. iSCAN: A Phoneme-based Predictive Communication Aid for Nonspeaking Individuals. In ASSETS ’12: Proceedings of the ACM SIGACCESS Conference on Computers and Accessibility. 57–64. [55] Keith Trnka, John McCaw, Debra Yarrington, Kathleen F. McCoy, and Christopher Pennington. 2009. User Interaction with Word Prediction: The Effects of Prediction Quality. ACM Transactions on Accessible Computing 1, 17:1–17:34. Issue 3. [56] Stephanie Valencia, Mark Steidl, Michael Rivera, Cynthia Bennett, Jeffrey Bigham, and Henny Admoni. 2021. Aided Nonverbal Communication through Physical Expressive Objects. In Proceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’21). Association for Computing Machinery, New York, NY, USA, 1–11. doi:10.1145/3441852.3471228 [57] Gerard Van Herk. 2018. What is Sociolinguistics? (2nd ed.). John Wiley & Sons, Inc, Hoboken, NJ. [58] Krishna Venkatasubramanian, Haven Hardie, and Tina-Marie Ranalli. 2025. Toward a taxonomy of negative outcomes from the use of AI-driven systems for

Frisch et al.

people with disabilities. In Proceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’25). Association for Computing Machinery, New York, NY, USA, 1–18. doi:10.1145/3663547.3746359 [59] Keith Vertanen, Dylan Gaines, Crystal Fletcher, Alex M. Stanage, Robbie Watling, and Per Ola Kristensson. 2019. VelociWatch: Designing and Evaluating a Virtual Keyboard for the Input of Challenging Text. In CHI ’19: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Glasgow, Scotland). 1–14. doi:10.1145/3290605.3300821 [60] Keith Vertanen and Per Ola Kristensson. 2014. Complementing Text Entry Evaluations with a Composition Task. ACM Transactions on Computer-Human Interaction 21, 2, Article 8 (February 2014), 33 pages. doi:10.1145/2555691 [61] Tonio Wandmacher, Jean-Yves Antoine, Franck Poirier, and Jean-Paul Départe. 2008. SIBYLLE, An Assistive Communication System Adapting to the Context and Its User. ACM Transactions on Accessible Computing 1, Article 6, 6:1–6:30 pages. Issue 1. [62] David J Ward, Alan F Blackwell, and David JC MacKay. 2000. Dasher—a data entry interface using continuous gestures and language models. In Proceedings of the 13th annual ACM symposium on User interface software and technology (San Diego, CA, United States). ACM Press, 129–137. doi:10.1145/354401.354427 [63] Tobias M Weinberg, Claire O’Connor, Ricardo E. Gonzalez Penuela, Stephanie Valencia, and Thijs Roumen. 2025. One Does Not Simply ‘Mm-hmm’: Exploring Backchanneling in the AAC Micro-Culture. In Proceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’25). Association for Computing Machinery, New York, NY, USA, 1–14. doi:10.1145/3663547.3746381 [64] Roelof Wessels, Jan Persson, Øivind Lorentsen, Renzo Andrich, Maurizio Ferrario, Wija Oortwijn, Thijs Van Beekum, and Luc De Witte. 2002. IPPA: Individually Prioritised Problem Assessment. Technology and Disability 14, 3 (2002), 141–145. [65] Brian Whitmer. 2026. AAC Effort Algorithms (Algorithm Version 0.2). Google Document, Publicly Available Online. https://docs.google.com/document/d/ 1ZJAt1JkpXcHgazEkWMFxxD_l117eD21p1uEFLMqjrjA/edit [66] Bruce Wisenburn and D Jeffery Higginbotham. 2008. An AAC application using speaking partner speech recognition to automatically produce contextually relevant utterances: Objective results. Augmentative and alternative communication 24, 2 (2008), 100–109. [67] A. J. Withers. 2012. Disability Politics & Theory. Fernwood Publishing, Halifax, Nova Scotia, Canada. [68] Mine Yasemin, Aniana Cruz, Urbano J Nunes, and Gabriel Pires. 2023. Single trial detection of error-related potentials in brain–machine interfaces: a survey and comparison of methods. Journal of Neural Engineering 20, 1 (Jan. 2023), 016015. doi:10.1088/1741-2552/acabe9 [69] Yanmei Zhu, Qian Wang, and Li Zhang. 2021. Study of EEG characteristics while solving scientific problems with different mental effort. Scientific Reports 11, 1 (Dec. 2021), 23783. doi:10.1038/s41598-021-03321-9

Record · ID 303241 · SHA-256 51c2796cd180751d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.