Received: Added at production
Revised: Added at production
Accepted: Added at production
DOI: xxx/xxxx
ARTICLE TYPE
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
arXiv:2609.21600v1 [cs.AI] 18 Sep 2026
Andy Gray 1
Jake Hobbs
School of Design, Bath Spa University, Bath, UK
Abstract Access to academic support is a key determinant of student success, yet students do not experience that access equally. While some students readily seek assistance from lecturers or tutors, others hesitate due to anxiety, fear of judgement, uncertainty about expectations, or low confidence in their own understanding. These help‐seeking disparities may be particularly evident in computing education, where programming tasks are cumulative and cognitively demanding. Although stu‐ dents increasingly turn to general‐purpose generative AI tools for assistance, such systems can produce responses that are inaccurate, insufficiently contextualised, or misaligned with module expectations. This study presents and evaluates Beacon, a course‐specific Retrieval‐Augmented Generation (RAG) system designed to provide private, immediate, and module‐aligned academic support. Grounding responses exclusively in approved teaching materials, Beacon was designed to complement existing educational support by lowering barriers to help‐seeking while encourag‐ ing independent learning. Using a Design‐based research approach the study followed an iterative design process to develop Beacon with analysis conducted through a mixed‐methods evaluation combining questionnaires with semi‐structured interviews with both students and staff at a Higher Education institution. Students consistently described Beacon’s responses as closely aligned with module content and more trustworthy than unrestricted generative AI tools, valuing its use of pseudocode and scaffolded explanations over direct solutions. Although participants remained ap‐ propriately cautious about trusting AI‐generated responses without verification, they consistently viewed the system as a valuable first point of support before consulting lecturers or official learn‐ ing resources. The findings suggest that carefully designed course‐specific AI systems may reduce barriers to academic support by occupying an intermediary space between independent study and formal academic support. Rather than replacing educators, educational AI may be most valuable when it broadens access to academically appropriate guidance while preserving the pedagogical role of lecturers. KEYWORDS
Artificial Intelligence; Large‐Language Models; Retrieval‐Augmented Generation; Education; Stu‐ dent Support
CONTEXT AND IMPLICATIONS Rationale for this study Students do not experience access to academic support equally. While universities provide a range of formal support mechanisms, many stu‐ dents remain reluctant to seek assistance due to anxiety, fear of judge‐ ment, low confidence, or uncertainty about where to ask for help. As
2026;:1–27
wileyonlinelibrary.com/journal/
© 2026 Copyright Holder Name
1
2
a result, students increasingly turn to general‐purpose large language
ways that promote equitable access to learning support. For educa‐
models (LLMs), such as ChatGPT, as an accessible source of immediate
tors, the study suggests that course‐specific AI systems can comple‐
academic support. Although these systems are readily available, they
ment existing teaching practices by providing students with immediate,
frequently generate responses that are detached from module‐specific
module‐aligned guidance outside scheduled teaching sessions. Rather
learning outcomes, institutional expectations, and assessment require‐
than replacing lecturers or tutors, such systems may act as an accessible
ments, potentially creating confusion, reinforcing misconceptions, or
first point of support, encouraging students to build confidence before
encouraging inappropriate use.
engaging with formal academic assistance.
This study therefore investigates whether a course‐specific Retrieval‐
For higher education institutions, the findings demonstrate the im‐
Augmented Generation (RAG) system can provide a more educationally
portance of designing AI systems around pedagogical objectives rather
appropriate form of AI‐supported learning. By grounding responses in
than technological capability alone. Grounding responses in approved
approved teaching materials, the system aims to provide timely, trust‐
teaching materials, maintaining transparency regarding source informa‐
worthy, and context‐aware academic support that complements exist‐
tion, and aligning AI outputs with module learning outcomes offer one
ing teaching practices while reducing barriers to help‐seeking. In doing
approach to supporting responsible AI adoption while preserving trust
so, the approach also has the potential to promote academic integrity by
in institutional teaching practices. Although course‐specific AI is not a
encouraging students to engage with module‐aligned guidance rather
substitute for effective teaching or student support services, it may help
than relying on unrestricted external AI systems.
reduce barriers to help‐seeking for students who are less likely to access conventional forms of academic support.
Why the new findings matter
More broadly, this study contributes to the growing body of research on educational AI by illustrating how generative AI can be designed to address educational challenges rather than simply automate exist‐
The findings demonstrate that barriers to academic support extend be‐
ing practices. The findings encourage future research to investigate
yond the simple availability of support services. Students’ willingness
how similar approaches perform across different disciplines, student
to seek help is shaped by factors such as confidence, anxiety, fear of
populations, and institutional contexts, and whether such systems can
judgement, and perceptions of whether their questions are appropriate
contribute to longer‐term improvements in student engagement, confi‐
to ask in staff‐mediated or public learning environments. By providing
dence, and educational equity.
private, immediate, and course‐specific guidance, Beacon offered stu‐ dents a low‐friction means of accessing support before engaging with lecturers or tutors. Rather than replacing existing support structures, the system occupied an intermediary space between independent study and formal academic support, lowering the threshold for help‐seeking while preserving the central role of educators. The study also contributes to wider debates surrounding the respon‐ sible and equitable integration of generative AI in higher education. Although unrestricted AI systems can increase access to academic assis‐ tance, they may also reinforce educational inequalities when students receive responses that are inaccurate, insufficiently contextualised, or beyond the intended level of study. The findings suggest that course‐ specific Retrieval‐Augmented Generation (RAG) offers one possible approach to addressing these challenges by grounding responses in ap‐ proved teaching materials, promoting transparency, and encouraging scaffolded learning rather than answer substitution. More broadly, this work demonstrates how educational AI can be designed not simply to improve efficiency, but to widen equitable access to academically appropriate learning support.
1
INTRODUCTION
Access to academic support is not experienced equally across higher education (HE). Concerns about judgement, uncertainty about where to ask for help, and limited access to academic staff can discourage students from seeking assistance. These challenges are particularly pro‐ nounced in large classes where opportunities for individual support may be limited. In this paper, help‐seeking disparity refers to unequal patterns of access to academic support that arise not only from the availability of support services, but also from students’ confidence, prior experi‐ ence, anxiety, fear of judgement, and ability to navigate institutional systems. Within computing education, these disparities may be partic‐ ularly pronounced because programming difficulties are often highly individual, cumulative, and difficult for students to articulate. As a result, students who are less confident or more anxious may delay seeking help, rely on unsuitable external resources, or disengage from learning activities. These confidence and anxiety‐related barriers to
Implications The findings have implications for educators, institutions, and re‐ searchers seeking to integrate generative AI into higher education in
help‐seeking are not evenly distributed across the student population. Prior research associates elevated anxiety and reduced confidence in help‐seeking with mature‐student (Chapman 2017), first‐generation or widening‐participation status (Koh et al. 2022) and gender and cultural background (Bornschlegl et al. 2020, Ruihua et al. 2025). Help‐seeking disparity, as examined in this study, may therefore intersect with, rather
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
3
than sit apart from the demographic and identity‐based disparities that
The remainder of the paper is structured as follows. Section 2 reviews
generative AI (GenAI) in education is increasingly expected to address
the literature concerning help‐seeking, generative AI, RAG, and trust in
(James et al. 2024, Ni et al. 2026).
educational AI systems. Section 3 describes the design and technical
In recent years, students have increasingly turned to GenAI tools
implementation of Beacon. Section 4 outlines the research methodol‐
such as ChatGPT to assist with their learning. While these systems can
ogy including findings from an initial exploratory pilot study. Section 5
provide rapid responses to questions, they often generate explanations
then presents and discusses the main study findings. Sections 6 and 7
that are misaligned with course‐specific materials or that assume levels
consider the study’s limitations and directions for future work, before
of prior knowledge beyond those of novice learners. This can lead to
Section 8 concludes the paper.
confusion or reinforce misunderstandings. Furthermore, they may intro‐ duce new disparities by advantaging students who are already able to evaluate AI‐generated responses critically, while disadvantaging those
2
BACKGROUND AND LITERATURE
who are less confident or less able to identify inaccuracies. Retrieval‐Augmented Generation (RAG) offers a potential solution to
Access to academic support is a fundamental component of student
this challenge by grounding large language model (LLM) responses in cu‐
success in HE. However, access to support is not experienced equally.
rated sources of information. By retrieving relevant content from course
While universities provide a range of formal and informal support
materials before generating a response, a RAG system can provide
mechanisms, including lectures, tutorials, office hours, peer learning,
answers that are more closely aligned with the curriculum.
and online resources, students differ considerably in their willingness
This study explores the development and evaluation of a RAG‐based
and ability to engage with these opportunities. Factors such as confi‐
academic support system designed specifically for computing students
dence, anxiety, fear of judgement, prior educational experiences, and
in programming modules. The system integrates an open‐source LLM
perceptions of belonging can all influence whether students seek help
with a retrieval layer built from university‐curated teaching materials.
when encountering academic difficulties. Consequently, disparities in
The aim of this research is to examine whether such a system can pro‐
help‐seeking behaviour may contribute to wider inequalities in learn‐
vide accurate, relevant, and reassuring academic support while reducing
ing experiences and academic outcomes, even when equivalent support
hesitation in seeking help.
provision exists.
This study makes three contributions to research on educational AI
Recent advances in GenAI have introduced new possibilities for ad‐
and academic help‐seeking. First, it provides a mixed‐methods eval‐
dressing these challenges. Students are increasingly turning to LLMs as
uation of a course‐specific RAG system incorporating both student
an accessible source of immediate academic support, yet unrestricted AI
and academic perspectives, examining not only usability and perceived
systems frequently produce responses that are disconnected from insti‐
response quality but also trust, learning support, and help‐seeking. Sec‐
tutional curricula, pedagogical intentions, or assessment expectations.
ond, it examines course‐specific RAG as a potential intermediary layer
This creates a tension between expanding access to learning support
between independent study and formal academic support, exploring
and ensuring that such support remains educationally appropriate, trust‐
whether private, immediate, and module‐aligned assistance can lower
worthy, and equitable. This literature review is therefore structured
perceived barriers to initiating help‐seeking. Third, it identifies design
as follows. First, the study context is established by exploring the
tensions of course‐specific AI support, including the balance between
difficulties in learning programming which can heighten help‐seeking
accessibility and cognitive effort, curriculum alignment and conversa‐
disparities. Second, factors influencing help‐seeking behaviour in HE
tional flexibility, and pedagogical scaffolding and learner autonomy.
are examined before the growing role of AI in supporting learning are
These findings provide insights into how course‐specific generative
considered. The review concludes by proposing that course‐specific
AI might complement, rather than replace, existing academic support
Retrieval‐Augmented Generation (RAG) systems offer a promising ap‐
structures.
proach to reducing barriers to academic support while maintaining
The study is guided by the following research questions:
• RQ1: How do students perceive the usefulness and relevance of a course‐specific RAG system for academic support?
alignment with institutional teaching practices.
2.1
Learning Programming
• RQ2: To what extent do students perceive the system as reducing barriers associated with academic help‐seeking?
• RQ3: How do students evaluate the trustworthiness of responses generated by a system grounded in course materials?
The focal context for this study is students learning computer program‐ ming, a subject considered difficult to learn and where understanding can take years to develop (Robins et al. 2003, Sentance and Csizmadia
• RQ4: How do academic staff perceive the system’s value in reducing
2017). The abstract nature of the subject means novice programmers
barriers to student help‐seeking, and what tensions does this raise
often struggle, particularly when combining concepts or applying knowl‐
for maintaining pedagogical safeguards and scaffolded learning?
edge to new problems (Robins et al. 2003, Saeli et al. 2011). This can create vicious cycles in knowledge development, especially when a typi‐ cal course introduces new principles each session in an upward learning
4
curve. However, a weak understanding of foundational concepts can in‐
whilst for unconfident learners help‐seeking is positively associated
hibit understanding of more advanced principles (Duke et al. 2000, Saeli
with academic success (Broadbent and Howe 2023). This creates a con‐
et al. 2011).
tradiction where those who need help the most are the ones who
These issues are increased by diverse cohorts where some students
seek it the least despite the positive benefits it will offer (Ryan et al.
enter HE with existing programming experience, while others do not.
1998, Karabenick and Knapp 1991, Micari and Calkins 2021). Again,
Sentance and Csizmadia 2017 argue skills disparities among students
this leads to educational disparities, with student confidence or self‐
are one of the biggest challenges in teaching programming. Faced
efficacy influenced by prior experience (Hill et al. 2022) and personal
with peers they perceive as more competent, students can experience
or environmental factors (Miao et al. 2025) including race, age and gen‐
anxiety and reduced confidence, which impacts engagement (Morales‐
der (Huang 2013, Brown et al. 2021, Koh et al. 2022, Micari and Calkins
Navarro et al. 2024, Rosenstein et al. 2020, Baker and Smith 2019).
2021). These issues are observed first hand by the authors, with some
Furthermore, research shows that novice programmers struggle to seek
students being reluctant to seek clarification or help, even when strug‐
help effectively and will either avoid help even when needed, or abuse
gling. This is then further evident in module feedback and assessment
help and bypass real learning (Marwan et al. 2020).
reflections where students comment that they should have made more contact with their tutor. Cultivating a supportive culture that normalises help‐seeking is es‐
2.2 Help‐Seeking Behaviour in Higher Ed‐ ucation
sential to address these issues and involves fostering belonging and
In HE more broadly students are expected to show independence in
have the potential to support academic learning by summarising infor‐
trust (Sithaldeen et al. 2022), enhancing tutor training (Mcfarlane 2016), and promoting peer support (Krisi and Nagar 2021). Additionally, LLMs
academic work, particularly with assessment tasks. To develop indepen‐
mation and providing feedback (Laato et al. 2023, Radford et al. 2019).
dence and achieve academic success, help‐seeking behaviour is argued
While these tools are beneficial, they can sometimes generate content
to play an important role (Fong et al. 2023). However, students often
that misaligns with course goals, leading students off‐track (Eager and
avoid seeking support due to anxiety, fear of mistakes, and a lack of
Brunton 2023). Carefully integrating LLMs, possibly through RAG frame‐
confidence (Salim 2022). In addition, a fear of seeming inadequate, the
works, ensures course relevance by drawing from specific databases
desire to appear competent, or being unaware of the support avail‐
(Lewis et al. 2020). Therefore, supporting students in overcoming help‐
able can further inhibt help‐seeking (Karabenick and Knapp 1991). This
seeking reluctance requires institutional efforts to normalise seeking
body of work has been synthesised in a systematic review of 55 stud‐
assistance with value in exploring the possibilities enabled by artificial
ies spanning a decade of research, which identifies confidence, fear of
intelligence tools in supporting such efforts.
judgement, and limited awareness of available support as recurring pre‐ dictors of help‐seeking avoidance across the higher education literature (Li et al. 2023). HE culture can inadvertently reinforce this avoidance by implying stu‐
2.3 cation
AI‐Supported Learning in Higher Edu‐
dents should already possess skills to manage complex tasks, making help‐seeking seem like a weakness (Sithaldeen et al. 2022, Robiullah
Artificial intelligence (AI) has become an increasingly prominent feature
et al. 2026). Within competitive environments, which may be caused
of HE, supporting activities including personalised learning, formative
by comparing oneself to peers, students may also internalise confusion,
feedback, assessment, academic writing, and student support (Holmes
leading to heightened anxiety and poorer performance (Newton and
et al. 2019, Luckin and Holmes 2016, Selwyn 2019, Gray et al. 2024
Miah 2017). Furthermore, large class sizes and limited access to aca‐
2025). Rather than functioning solely as administrative technologies, AI
demic staff can discourage students from asking for guidance (Tinto
systems are increasingly positioned as learning companions capable of
2012), while lecturers, who are often seen as a source of support,
providing immediate explanations, answering questions, and supporting
may themselves be uncertain about their role in student assistance
students outside scheduled teaching sessions.
(Grayson et al. 1998), or lack confidence in their ability to support diverse students (Mcfarlane 2016).
The emergence of LLMs and GenAI has accelerated this shift. Stu‐ dents now routinely use systems such as ChatGPT to explain concepts,
Willingness to seek help is also influenced by students gender, age,
debug code, summarise readings, and clarify assessment requirements
cultural background and personality (Bornschlegl et al. 2020, Ruihua
(Laato et al. 2023, Eager and Brunton 2023). LLMs also help with lan‐
et al. 2025). In particular, research has shown that mature students can
guage tasks, offering feedback on grammar and structure, as well as
face issues with imposter syndrome and confidence (Chapman 2017).
personalising learning content based on student needs (Radford et al.
Therefore, as the diversity of the student population expands, so do the
2019). The availability of these systems provides students with immedi‐
disparities regarding help‐seeking behaviour.
ate access to support regardless of time or location, potentially reducing
Confidence influences help‐seeking, with confident learners more
reliance on traditional forms of academic assistance (Laato et al. 2023,
likely to seek support (Broadbent and Howe 2023). However, confident
Yang 2025). This flexibility may be particularly valuable for students who
learners are also found to benefit the least from help‐seeking strategies,
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
are reluctant to seek help from lecturers or peers due to anxiety, low confidence, or fear of judgement.
5
This approach offers particular promise for educational applications because it allows responses to be grounded in course materials, institu‐
While these systems offer new opportunities to support learning
tional resources, or other curated sources. As a result, students receive
and teaching, they also introduce a number of challenges that affect
answers that are contextually relevant and aligned with the curriculum.
their reliability in educational settings. Concerns have been raised about
RAG improves LLMs by allowing them to access external databases,
the accuracy of generated outputs, the presence of biases in training
ensuring they provide up‐to‐date accurate information. Traditional mod‐
data, and the difficulty users may face when evaluating the quality of
els like GPT rely only on the information they were trained on, which
responses (Bommasani et al. 2021, Bender et al. 2021).
can become outdated (Brown 2020). RAG overcomes this by combin‐
One widely discussed limitations of GenAI is their tendency to pro‐
ing the ability of the model to generate text with real‐time information
duce incorrect or fabricated information, sometimes referred to as “hal‐
retrieval, leading to more accurate and relevant responses (Lewis et al.
lucinations”. Because these systems generate text by predicting likely
2020).
word sequences rather than retrieving verified facts, they may produce
In HE, RAG can help students retrieve relevant academic papers,
outputs that appear convincing but contain inaccuracies. Alternatively,
articles, and reliable sources, making research easier. It can also gen‐
they may assume inappropriate levels of prior knowledge, or recom‐
erate summaries or explanations, helping students engage with large
mend approaches that differ from those expected within a particular
amounts of material for assignments and exams (Izacard and Grave
module or assessment (Eager and Brunton 2023, Caines et al. 2023,
2020). Furthermore, RAG can personalise learning by providing infor‐
Yang 2025). This poses particular risks in educational contexts where
mation tailored to a student’s needs, ensuring they receive up‐to‐date
students may rely on generated responses as authoritative sources of
content (Borgeaud et al. 2022).
information (Ji et al. 2023). A recent systematic review of GenAI in
The application of RAG specifically to course‐ and institution‐specific
computer science education also found that hallucinated or misleading
student support has grown rapidly. A recent survey of 47 studies on RAG
outputs can increase cognitive load during programming and debugging
chatbots in education maps this emerging sub‐field across application
tasks. Consequently, this may disproportionately disadvantage students
type, knowledge domain, and evaluation approach (Swacha et al. 2025).
with limited prior programming experience or from under‐resourced
Within computing education specifically, several course‐specific imple‐
learning contexts, thereby widening rather than narrowing existing
mentations have been reported: Neumann et al. (2025) developed and
educational disparities (Adejumo et al. 2026).
evaluated an LLM‐driven RAG chatbot for a databases and information
Furthermore, GenAI poses challenges related to transparency and
systems course, reporting an 88% accuracy rate against course content;
accountability. The scale and complexity of modern language models
Lang et al. (2025) evaluated a RAG chatbot in an online programming
make it difficult for users to understand how responses are produced or
course, finding that students with greater prior knowledge engaged
which sources influenced the output. As a result, educators and learn‐
more with advanced queries; Alsafari et al. (2024) compared intent‐
ers may find it difficult to assess the reliability of generated information
based and RAG‐based teaching assistants for a data mining course;
or determine when AI tools should be trusted (Bommasani et al. 2021).
Németh et al. (2025) piloted a RAG‐based tutor across four courses at
To address these concerns, students must critically engage with LLM‐
two universities, reporting expert‐coded accuracy rates alongside stu‐
generated content, ensuring they use these tools to complement, rather
dent and lecturer perceptions; and Tran Huu Van et al. (2026) reported
than replace, their academic efforts (Caines et al. 2023, Yan et al.
a course‐specific agentic RAG chatbot for IT student support with a
2023). Without appropriate guidance, learners may struggle to critically
comparable architecture and privacy‐aware design. Beacon extends this
evaluate AI‐generated content.
emerging body of work by combining both student and academic‐staff
These opportunities and challenges suggest that educational AI
evaluation within a single study, and by considering not only usability
should not simply aim to maximise automation or information access,
and response accuracy but also the impact on trust, learning support,
but instead support meaningful learning within the context of insti‐
and help‐seeking.
tutional teaching practices. This has led increasing attention towards AI systems that can provide contextually grounded, course‐specific guidance while remaining aligned with approved educational resources.
2.4 Retrieval‐Augmented Generation for Educational Support
2.5 Trust, Transparency Human‐Centred Educational AI
and
Trust is widely recognised as a prerequisite for the effective adoption of AI in education (Department for Education 2023, Quality Assurance Agency for Higher Education 2023, Jisc 2023). Students must feel con‐
RAG combines information retrieval with LLM generation. Rather than
fident that AI‐generated guidance is reliable, relevant, and appropriate
rely solely on the internal knowledge of a LLM, RAG systems retrieve
to their learning context, while educators require assurance that AI sys‐
relevant documents from a predefined knowledge base and use them to
tems support rather than undermine established pedagogical practices
inform the generated response (see Figure 1 for a visual representation).
(Gonsalves and Lin 2025, Gray et al. 2026, Holmes et al. 2019, Nazaret‐ sky et al. 2022). For educational AI systems intended to support learning,
6
F I G U R E 1 A visual demonstration of how RAG can help inform the LLMs with relevant material stored within a vector database (Greer 2024). The user will ask a question, which, in turn, the smart retriever, a vector similarity searcher, will query the Vector database to retrieve the relevant information to present to the LLM as a generator. The LLM then produces an output representing the user’s question and the additional information retrieved from the vector database. In our context, lesson content and assessment details are stored in the vector database to aid the students.
trust influences not only whether students choose to use the system,
with end users. Such approaches recognise that technically accurate re‐
but also whether they engage critically with the guidance provided and
sponses are insufficient if learners find systems difficult to use, cannot
integrate it appropriately into their learning.
determine when responses should be trusted, or are unable to relate
Research suggests that trust is shaped by more than technical per‐
AI‐generated guidance to their existing learning resources.
formance. Transparency, explainability, and opportunities for human
For course‐specific AI systems, these principles suggest that success‐
oversight contribute to how users perceive the reliability and credibility
ful educational support depends not only on retrieving relevant infor‐
of AI‐generated responses (Glikson and Woolley 2020, Khosravi et al.
mation, but also on presenting that information in ways that encourage
2022). Systems that provide little insight into how responses are gener‐
confidence, independent learning, and appropriate help‐seeking. In‐
ated may be perceived as opaque ”black boxes”, making it difficult for
terfaces should therefore minimise barriers to use, communicate the
students and educators to judge the validity or appropriateness of the
provenance of generated responses, and enable students to verify ex‐
information provided. Explainable AI (xAI) therefore seeks to improve
planations against approved teaching materials before applying them to
transparency by enabling users to understand the basis of AI‐generated
their own work. Such design choices encourage students to remain ac‐
outputs through meaningful explanations or supporting evidence (Adadi
tive participants in the learning process rather than passive recipients
and Berrada 2018). Within educational contexts, this transparency can
of AI‐generated answers (Jin et al. 2023, Lan and Zhou 2025).
help students verify AI‐generated guidance against trusted learning re‐
Taken together, the literature suggests that trustworthy educational
sources while enabling educators to evaluate whether responses remain
AI emerges from the combination of technically grounded responses,
aligned with module learning outcomes and assessment expectations.
transparent presentation of supporting evidence, and human‐centred
However, transparency alone does not automatically establish trust.
design. Rather than viewing AI solely as a means of increasing efficiency,
Students must also perceive AI systems as intuitive, accessible, and re‐
these principles position educational AI as a mechanism for widening
sponsive to their learning needs. Consequently, human‐centred design
equitable access to academically appropriate learning support. This per‐
has become an increasingly important principle within educational AI,
spective is particularly relevant for students who may be reluctant to
emphasising that systems should be designed around the needs, expe‐
seek help through conventional channels due to anxiety, low confi‐
riences, and educational contexts of their intended users rather than
dence, or fear of judgement. These principles therefore informed both
technological capability alone (Dix 2016, Shneiderman 2022, Becker
the design of Beacon and the methodological approach adopted in its
2020). Human‐centred approaches advocate iterative design, user in‐
evaluation, providing the foundation for a course‐specific AI system
volvement, accessibility, and continuous refinement through evaluation
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
intended to complement existing teaching while reducing barriers to academic support.
7
and assessment expectations rather than relying solely on the broad knowledge of a general‐purpose LLM. 2. Support independent learning: Beacon should encourage students to understand concepts and develop problem‐solving skills through
3
SYSTEM DESIGN
explanations and guidance rather than simply providing complete solutions.
The student support tool, referred to throughout this paper as Beacon,†
3. Promote transparency and trust: Students should be able to under‐
was designed to provide immediate, course‐specific academic guidance
stand where information originates enabling them to verify expla‐
that complemented existing teaching while reducing barriers to help‐
nations against official module resources. For example, by stating
seeking. Rather than relying on unrestricted LLMs, which may generate
the material was from session presentation x and slide y, or from
responses beyond the scope of a module, Beacon grounds responses
handout z.
exclusively in approved teaching materials. This section describes the
4. Complement existing teaching practices: Beacon should enhance,
educational design objectives, the RAG architecture adopted to achieve
rather than replace, interactions with lecturers and tutors by pro‐
them, the construction of the course‐specific knowledge base, and the
viding timely support between teaching sessions and encouraging
user interface developed through iterative evaluation.
students to seek further assistance when appropriate.
3.1
Design Objectives
These objectives informed each stage of Beacon’s development, from the construction of the course‐specific knowledge base and re‐ trieval pipeline to the design of the user interface. Rather than viewing
The primary objective of Beacon was to provide students with imme‐
the LLM as the primary source of knowledge, Beacon treats approved
diate, course‐specific academic guidance that complemented existing
teaching materials as the authoritative foundation for all responses.
teaching while reducing barriers to help‐seeking. Rather than replacing
RAG was therefore adopted not simply as a technical solution, but as
lecturers, tutorials, or existing learning resources, Beacon was designed
an educational design choice that enables AI‐generated support to re‐
to act as an accessible first point of support when students encoun‐
main transparent, contextually relevant, and aligned with institutional
tered difficulties outside scheduled teaching. This was motivated by
teaching practices.
the recognition that many students are reluctant to seek academic help due to factors such as anxiety, fear of judgement, uncertainty regarding module expectations, or lack of confidence in their own understanding.
3.2
System Architecture
Thus, a core aim for Beacon was to lower the threshold for accessing support while encouraging students to engage more confidently with
The student support tool adopts a modular architecture that combines
formal teaching provision when required.
a course‐specific knowledge base with a LLM to provide contextu‐
Furthermore, the authors observed a sharp‐rise in academic mis‐
ally grounded responses. Rather than relying solely on the pretrained
conduct cases involving generative AI in recent years stemming from
knowledge of the LLM, every response is generated using relevant
students consulting unrestricted tools such as ChatGPT for program‐
teaching materials retrieved from the module knowledge base. This
ming assistance and receiving responses that draw on techniques, li‐
ensures that generated guidance remains aligned with the learning out‐
braries, or approaches beyond the scope of the module or assessment.
comes, teaching content, and assessment expectations of the module
When such content is incorporated into submitted work, the resulting
while preserving the conversational interaction expected of modern AI
mismatch with taught material and student understanding can prompt
assistants.
closer scrutiny from module staff, and in some cases leads to formal aca‐
The overall workflow is illustrated in Figure 1. Teaching materials are
demic misconduct proceedings. Because Beacon’s responses would be
first processed and indexed within a vector database to create a search‐
grounded exclusively in approved teaching materials, the aim is to re‐
able knowledge base. When a student submits a question, the query
duce this particular pathway to misconduct by keeping guidance within
is converted into a semantic representation and matched against the
the scope of the module.
indexed course materials to identify the most relevant content. The re‐
To achieve these aims Beacon was developed around four educa‐ tional design objectives:
trieved passages are then incorporated into the prompt supplied to the LLM, allowing responses to be generated using information drawn di‐ rectly from approved module resources rather than relying exclusively
1. Provide course‐specific academic support: Responses should re‐ main aligned with the module learning outcomes, teaching materials,
on the model’s internal knowledge. This architectural approach was selected to support the educational design objectives outlined in the previous section. Grounding responses in module‐specific resources helps ensure that explanations remain con‐
†
Beacon was developed and evaluated under the working name DebuggyDuck during its initial pilot development. It was renamed prior to the main study reported here; both names refer to the same underlying system.
sistent with the material taught during class, reducing the likelihood that students receive guidance beyond the intended scope of the module.
8
Furthermore, because responses are explicitly linked to retrieved teach‐
student’s query was compared with the stored document embeddings.
ing materials, students are encouraged to verify explanations against
The TOP‐7 highest‐ranked chunks were returned, with no fixed similar‐
the original resources, supporting transparency, independent learning,
ity threshold, and were subsequently incorporated into the generation
and trust in the system.
prompt.
The architecture comprises four principal components:
The knowledge base was indexed before deployment and updated whenever approved teaching materials were added or revised. Only ma‐
1. a curated knowledge base containing approved teaching resources;
terials selected by the module teaching team were included, ensuring
2. a semantic retrieval component responsible for identifying relevant
that the retrieved content remained course‐specific and pedagogically
learning materials;
appropriate.
3. a LLM responsible for generating conversational responses; and 4. a web‐based user interface through which students interact with the system.
3.2.2 Generation Pipeline
Retrieval‐Augmentation
Therefore, ultimately the LLM functions as the conversational in‐ terface, whereas approved teaching materials remain the authoritative
The retrieval pipeline (Figure 2) distinguished between the initial user
source of academic content.
message and subsequent conversational turns. When a student sub‐
The following sections describe each of these components in greater
mitted the initial question, the message was converted directly into an embedding and used to perform a semantic similarity search over the
detail.
indexed course materials.
3.2.1
Knowledge Base Construction
For each subsequent message, the system generated a self‐contained retrieval query using the latest user message and the preceding con‐ versation history. This step was introduced because conversational
The knowledge base was constructed from teaching materials approved
follow‐ups may be semantically underspecified when considered in iso‐
for use within the participating computing modules. These materi‐
lation. For example, a request such as “Can you explain more?” does not
als included lecture slides, laboratory exercises, formative assessment
identify the concept to which it refers and would therefore be unlikely
activities, module handbooks, assessment guidance, and supplemen‐
to retrieve relevant material if embedded directly. Rewriting the mes‐
tary learning resources. Including multiple resources enabled Beacon
sage using the preceding context produced a more informative search
to retrieve both conceptual explanations and practical guidance while
query while preserving the student’s intended meaning.
remaining aligned with the content and expectations communicated through formal teaching.
The initial or rewritten query was embedded and compared with the knowledge‐base embeddings using cosine similarity. The TOP‐7
To provide a consistent representation across different document
highest‐ranked document chunks were retrieved and incorporated into
types, the teaching materials were prepared in Markdown. Presenta‐
the final generation prompt. This prompt comprised the system instruc‐
tion materials were authored using MARP, while supporting documents
tions, the student’s original message, the retrieved course materials, and
were stored as standard Markdown files. This format preserved mean‐
the relevant conversation history.
ingful structural features, including headings, lists, code blocks, and links,
The rewritten query was used for retrieval purposes only. The stu‐
while removing visual and presentation‐specific formatting that was not
dent’s original message was retained within the final prompt so that the
required for semantic retrieval.
generated response remained natural and appropriate to the conversa‐
Before indexing, each document was segmented into smaller textual
tion. This process enabled multi‐turn interactions to remain grounded
chunks. This approach was intended to retain sufficient contextual in‐
in relevant course content, even where the student’s latest message
formation within each segment while preventing large documents from
contained little contextual information.
reducing retrieval precision. Each chunk was stored together with metadata describing its orig‐ inal source. Depending on the resource type, this metadata included
3.2.3
Prompt Construction
the module or session identifier, document title, resource type, topic, slide number, and the original text. Retaining this information enabled
The retrieved document segments were incorporated into a structured
retrieved passages to be traced back to their source materials and
prompt together with the student’s question before being supplied
displayed alongside generated responses. This supported transparency
to the language model. The prompt instructed the model to answer
by allowing students to verify AI‐generated explanations against the
using only the retrieved teaching materials, avoid speculation when
official module resources.
insufficient evidence was available, and encourage conceptual under‐
A vector embedding was generated for each chunk using an embed‐ ding model. The embeddings were stored in a PostgreSQL database using the pgvector extension. At retrieval time, the embedding of the
standing rather than directly completing assessed work. Where relevant,
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
9
Student Question
Initial Question?
Yes
No
Rewrite Query using Conversation Context
Generate Query Embedding Semantic Similarity Search Vector Database
Retrieve Top‐k Document Chunks
Construct LLM Prompt System Prompt + User Question + Retrieved Content + Conversation History
Large Language Model
Grounded Response Returned to Student
F I G U R E 2 Retrieval pipeline used to generate grounded responses. Initial questions are embedded directly for semantic retrieval, whereas follow‐up questions are first rewritten into self‐contained queries using the preceding conversation before retrieval is performed. The final prompt supplied to the language model combines the system prompt, user query, retrieved teaching materials, and conversation history.
the generated response was accompanied by references to the origi‐
minimalist design, consisting of a single text input field and submission
nal teaching resources, allowing students to verify explanations against
button. After a question was submitted, the generated response was
official module content.
streamed beneath the input area along with the relevant supporting con‐ tent (Figure 3). This lightweight interface prioritised rapid development
3.3
User Interface
and evaluation over visual sophistication, allowing the focus of the pilot study to remain on the effectiveness of the underlying RAG pipeline. Following the pilot study, the application was reimplemented us‐
Students interacted with Beacon through a conversational interface
ing Flask to provide a more scalable architecture suitable for web
that enabled them to submit questions in natural language and re‐
hosting and wider deployment. Feedback obtained during the pilot
ceive contextually grounded responses. The application is designed to
(see Section 4.2) also informed several interface refinements based
resemble familiar chat interfaces to reduce barriers to adoption and
around a more conventional chat interface that more closely reflected
encourage student engagement. Responses were presented in a sup‐
commercial conversational AI systems. This provided a more intuitive
portive, conversational tone to encourage exploration and independent
interaction model while preserving transparency by continuing to dis‐
learning.
play the supporting source material alongside each response (Figure 4
Although versions of Beacon developed for the exploratory pilot and
). In addition, the underlying language model was upgraded from “GPT‐
main study share the same underlying RAG architecture, the implemen‐
OSS‐20B” to “GPT‐OSS‐120B” in Version 2. Consequently, the pri‐
tation evolved throughout the project. The pilot application (Version 1)
mary differences between the two versions were improvements to the
was developed using Streamlit to enable rapid prototyping and itera‐
user interface, deployment architecture, and language model capability,
tive testing with students. The pilot interface adopted a deliberately
rather than changes to the core RAG workflow.
10
embedded using Nomic Embed Text and indexed within the vector database prior to deployment. The complete application was developed using open‐source tech‐ nologies, allowing Beacon to be deployed locally or within institutional infrastructure without dependence on commercial AI platforms. This supports institutional control over teaching resources, model selec‐ tion, and student data while facilitating future adaptation to different modules or programmes.
4
METHODOLOGY
4.1
Research Design
This study adopted a Design‐based research (DBR) approach, which aims to improve educational practice through iterative design in authen‐ tic settings (Wang and Hannafin 2005). A typical DBR process follows five steps: 1. Identify a problem (established in Section 1 and Section 2) 2. Design an intervention (outlined in Section 3) 3. Implement and evaluate (see Section 4.2 and Section 5) 4. Refine the design (as illustrated by the changes between the pilot and main study) 5. Develop theory and/or design principles (see Section 5 and Section 8) Using the DBR approach the research followed an iterative design process comprising two stages. An initial pilot study evaluated a first prototype, focusing on the technical feasibility of the RAG approach and the usability of the interface. Findings from this pilot informed the development of a second version of Beacon, incorporating refinements to both the user interface and underlying infrastructure. This revised system was subsequently evaluated through a larger user study. This iterative cycle aligns with the DBR approach where successive refine‐ ment based on user feedback enables systems to better address learner needs (Hoadley and Campos 2022). DBR uses various approaches to evaluate the effectiveness of the designed intervention. This combination of methods is argued to in‐ F I G U R E 3 The Version 1 web application interface. The pilot imple‐
crease the validity and applicability of the research (Wang and Hannafin
mentation displays the generated response together with the support‐
2005). As such, data collection combined both qualitative and quan‐
ing source material used to produce the answer.
titative data through structured questionnaires and semi‐structured interviews. Questionnaires provided quantitative measures of students’ perceptions of Beacon, while interviews explored participants’ experi‐
3.4
Implementation Details
ences in greater depth, allowing themes to emerge regarding usability, trust, learning behaviours, and opportunities for improvement. Along‐
The application was implemented using Python’s Flask library and de‐
side student feedback, interviews were also conducted with lecturers
ployed as a web‐based application hosted on Python Anywhere. The
to gain their perspectives of Beacon. The combination of methods and
knowledge base was stored using PostgreSQL with the pgvector ex‐
gathering of multiple perspectives enabled the triangulation of find‐
tension to support semantic similarity search. Teaching resources were
ings, providing both breadth and depth when evaluating the educational value of Beacon (Cohen et al. 2009).
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
11
F I G U R E 4 The Version 2 web application interface. The redesigned chat interface presents generated responses alongside the supporting source material used to produce each answer.
4.2
Exploratory Pilot Study
In line with the iterative nature of the DBR approach, an exploratory pilot study was used to assess the feasibility and proof‐of‐concept validity of Beacon. During the pilot, Beacon was not deployed for in‐ dependent student use; instead, it was presented through a structured demonstration that illustrated its functionality and intended use. The demonstration included examples of typical student queries, such as requests for clarification of programming concepts and interpretation of assignment requirements. This approach enabled participants to ob‐ serve how the system retrieved relevant information and generated responses grounded in course materials. The pilot prioritised demonstrating core functionality and pedagog‐ ical alignment rather than full technical maturity. This allowed for an initial evaluation of how students perceived the system before further development and refinement.
4.2.1
Pilot Study Procedure
The pilot study was conducted with undergraduate computing students who attended scheduled teaching sessions where Beacon was demon‐ strated. Multiple groups of students participated in these sessions, providing a cross‐section of the cohort within the module context.
During each session, Beacon was introduced and demonstrated by the researcher. Students were shown how queries could be submitted and how the system generated responses using module‐specific mate‐ rials. The demonstration included a series of representative examples designed to reflect realistic academic scenarios, allowing students to understand how the system might be used to support their learning. Following the demonstration, students were invited to complete a structured questionnaire designed to capture their perceptions of the system. The questionnaire focused on students’ impressions of the sys‐ tem’s clarity, usefulness, relevance to course content, and perceived trustworthiness. It also explored how students anticipated using such a system in relation to their existing study practices. Participation in the study was voluntary, and all responses were col‐ lected anonymously. The aim of this procedure was to capture students’ immediate perceptions following exposure to the system, rather than to evaluate long‐term patterns of use. A total of nine valid responses were collected and used for subsequent analysis.
4.2.2
Pilot Study Evaluation
Findings from the pilot indicated a generally positive student response to the system, alongside identifiable areas for improvement. Regarding clarity, 44% of respondents reported that the system was very clear, with a further 44% indicating it was somewhat clear. A small
12
proportion remained neutral, suggesting that additional explanation or
(Faulkner 2003, Nielsen 2000). Participants represented a range of
demonstration may improve accessibility. Confidence in response qual‐
computing disciplines, including Computing, Cyber Security, Games De‐
ity was also high, with 56% of participants reporting that outputs were
velopment, and Creative Computing to ensure the evaluation reflected
consistently logical and the remainder indicating that responses were
diverse programming backgrounds, varying levels of confidence, and dif‐
mostly reliable.
fering degrees of familiarity with generative AI technologies. While the
Perceived relevance was particularly strong, with 89% of participants
majority of participants were first‐year students, second‐ and third‐year
rating the system’s responses as highly aligned with course materials.
students were also represented, enabling perspectives to be gathered
This suggests that the retrieval mechanism successfully grounded re‐
from learners at different stages of study.
sponses in module‐specific content, addressing a common limitation of general‐purpose AI tools. Students also reported that the system was useful in supporting assessment understanding, with all participants indicating that it was
Before interacting with Beacon all participants completed a pre‐ experiment questionnaire to establish baseline information relating to programming confidence, help‐seeking behaviours, use of existing mod‐ ule resources, and prior engagement with generative AI systems.
either definitely or somewhat helpful. This reinforces the system’s
Interaction with Beacon involved completing six task scenarios, a
potential role as a first point of support when students encounter
method recommended in usability testing for ensuring users meaning‐
difficulty.
fully engage with the task (McCloskey 2014). These scenarios were
Trust in the system was more cautious. While the majority of partic‐
design to reflect real‐world usage and included questions and issues
ipants expressed some level of trust, only a small proportion indicated
commonly experienced programming lessons (See Appendix C). Nine
complete confidence in the responses. This suggests that, although the
students completed the tasks during a two‐hour in person workshop
system is perceived as useful, students do not treat it as an authorita‐
with the remaining six completing the tasks remotely.
tive source. Instead, it is positioned as a supplementary tool, with many indicating that they would still verify important information.
After completing the task scenarios participating students completed a post‐experiment questionnaire evaluating the system’s usability, clar‐
Participants also recognised potential risks associated with the sys‐
ity, relevance, trustworthiness, learning support, and potential influence
tem. Concerns were raised regarding the possibility of incorrect informa‐
on future help‐seeking behaviours. Although responses were collected
tion and the risk of over‐reliance, particularly if students use the system
anonymously and no personally identifiable information was recorded,
as a substitute for independent thinking. These concerns highlight the
each participant used a unique anonymous identifier. This enabled pre‐
importance of transparency and appropriate integration into teaching
and post‐experiment questionnaire responses to be linked at the indi‐
practice.
vidual level while preserving participant anonymity. Five of the fifteen
Despite these limitations, many participants indicated that the sys‐
students subsequently took part in semi‐structured interviews to ex‐
tem could reduce hesitation in seeking help, particularly by enabling pri‐
plore their experiences in greater depth (receiving a further £10 voucher
vate interaction. This suggests that AI‐based support systems may play
for their time).
a role in addressing barriers associated with traditional help‐seeking behaviours.
To complement the student evaluation, four academic staff indepen‐ dently reviewed Beacon from a pedagogical perspective. All reviewers
Overall, the findings suggest that a course‐specific RAG system can
had experience teaching on the module associated with the study and
provide relevant and accessible academic support, while also introduc‐
represented a range of disciplinary backgrounds, including Computing,
ing important considerations relating to trust, transparency, and student
Cyber Security, Creative Computing, and Games Development. Collect‐
dependency. The feedback from the pilot study resulted in enhance‐
ing perspectives from both students and academics enabled Beacon to
ments to the user interface, deployment architecture, and language
be evaluated from both learner and educator viewpoints.
model capability as described in Section 3.
Ethical approval for the study was obtained through the University’s institutional ethics process, and no personally identifiable information
4.3
Participants
Participants in the main study were undergraduate students enrolled
was collected throughout the study.
4.4
Data Collection
on computing‐related degrees at a British University. Students were invited to participate voluntarily in the evaluation of Beacon as an AI‐
4.4.1
Questionnaires
supported learning tool. Before participation, students were informed of the purpose of the study and advised responses would be collected
Two structured questionnaires were used during the evaluation. The
anonymously and used solely for research purposes. Participating stu‐
pre‐experiment questionnaire collected demographic information to‐
dents received a £10 gift voucher, redeemable at a retailer of their
gether with measures relating to programming confidence, existing sup‐
choice, as compensation for their time.
port strategies, and prior experience of generative AI technologies. The
A total of fifteen students participated in the main study, a figure
post‐experiment questionnaire evaluated students’ perceptions of Bea‐
widely cited as sufficient to identify all critical usability problems
con following use, including measures of usability, usefulness, clarity,
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
13
trustworthiness, relevance, transparency, and perceived support for in‐
thematic analysis described by (Naeem et al. 2025). The use of AI in the‐
dependent learning. The questionnaires included five‐point agreement
matic analysis is argued to be more time‐efficient, reduce human bias
items, frequency‐rating items, comparative‐rating items, supplemented
and uncover insights that might otherwise be overlooked. As a result AI
by optional open‐ended responses to capture additional comments.
can enhance and improve the efficiency of the analysis (Naeem et al. 2025). To ensure accuracy, outputs generated by Copilot were reviewed
4.4.2
Semi‐Structured Interviews
Semi‐structured interviews were used to gain greater insight into partic‐ ipant experiences of using Beacon. Semi‐structured interviews occupy a methodological position between structured questionaries and open ethnographic observation, allowing researchers to investigate prede‐
and verified by the researchers at each stage of the analysis. Consistent with recommendations for qualitative research, careful documentation of the six‐step process, including all prompts used to generate findings, have been retained to enhance the credibility and trustworthiness of the findings (Blandford and Rugg 2002, Kamsin et al. 2012, Furniss et al. 2011).
fined topics while retaining the flexibility to explore unexpected issues raised by participants (Blandford 2013, Blandford et al. 2008). Thus, this approach sought to understand not only participants’ opinions of Beacon but also the reasoning behind those opinions. A total of nine interviews were conducted. Five with students and four with academic staff. Collecting perspectives from both students and academics enabled Beacon to be evaluated from both learner and educator viewpoints. Students were randomly selected for the interviews, with no se‐ lection criteria applied other than their indication of willingness to participate on the study consent form. Interview questions focused on perceptions of usability, trust, learning support, transparency, and potential improvements. Follow‐up questions enabled participants to elaborate on their experiences and provide examples of how Beacon influenced their learning. Interviews with academic staff focused on the educational appropri‐ ateness of the generated responses, alignment with module materials, and the potential integration of Beacon into existing teaching and student support practices.
5
RESULTS & DISCUSSION
5.1 5.1.1
Student Views Pre‐Experiment Questionnaire
The pre‐experiment questionnaire explored students’ existing help‐ seeking behaviours, confidence levels, and use of AI‐based support tools prior to interacting with Beacon. Overall, students reported moder‐ ate levels of programming confidence, with a mean self‐rated program‐ ming ability of approximately 6.25 out of 10. The grades achieved by students in the foundational programming module ranged from third‐ class to first‐class honours, indicating a diverse range of academic attainment among participants. The findings revealed an interesting tension between students’ will‐ ingness to seek support and their preference for independent problem‐ solving. Although participants generally reported feeling comfortable asking lecturers for help, many also demonstrated signs of hesitation and avoidance when encountering difficulties. The majority of students
4.5
Data Analysis
(62.5%) indicated they sometimes avoided asking for help even when they needed it, while a greater proportion reported anxiety when they
Questionnaire responses to quantitative measures were analysed to
did not understand a topic (75%). At the same time, there was a pref‐
calculate response frequencies and percentages, enabling comparison
erence for autonomy, with most students agreeing that they preferred
of student perceptions towards Beacon and its potential impact on
to attempt solving problems independently before seeking support
learning and help‐seeking behaviour. Responses to qualitative free‐text
(87.6%). These findings suggest that, while institutional support struc‐
questions were analysed analysed using a thematic approach to identify
tures are available, many students still attempt to manage uncertainty
recurring insights and emergent themes. Data analysis was supported
privately before engaging with formal sources of assistance.
by Microsoft Copilot, which was used to assist in summarising quantita‐
Student responses also highlighted widespread use of GenAI tools
tive results and identifying potential themes within the qualitative data.
within existing study practices. The majority of participants reported
All analyses and AI‐assisted outputs were reviewed and validated by the
using GenAI either sometimes or often when working on programming‐
authors to ensure accuracy, consistency, and appropriate interpretation
related tasks, while only a smaller subset (25%) indicated that they never
of the findings.
used such systems. AI tools were used across a broad range of activi‐
Interviews were conducted via Microsoft Teams, with transcripts gen‐
ties, including debugging, concept explanation, assignment clarification,
erated automatically. Immediately after the interview the researcher
code interpretation, study planning, and academic writing support. Sev‐
reviewed the transcripts to verify accuracy, correct errors and remove
eral students described AI systems as particularly valuable because they
identifying information. The final transcripts were analysed with the
offered rapid responses, adaptive explanations, and personalised sup‐
support of Microsoft Copilot using the six‐step process for AI‐assisted
port that could be tailored to individual levels of understanding. Others highlighted that AI systems allowed them to ask follow‐up questions
14
without feeling judged or pressured, making them easier to engage with
as a form of accessible, low‐friction support. These findings there‐
than traditional support channels in some situations.
fore reinforce the rationale for the development of a course‐specific Retrieval‐Augmented Generation system capable of providing contex‐
”The use of AI tools allows for personalised responses and follow up
tualised, module‐aligned assistance while reducing barriers associated
questions, which is much more flexible and efficient than the module‐
with conventional help‐seeking behaviours
provided information and third party websites.” Participant melo‐f6 Despite this widespread use of AI, students did not position these tools as entirely replacing traditional module resources. Participants continued to identify module‐provided content, such as lecture slides, recordings, GitHub repositories, and workshop materials, as the most helpful forms of support because these resources were directly aligned with course expectations and assessment requirements. Students em‐ phasised the importance of contextual relevance, noting that general online resources or unrestricted AI systems could sometimes provide ex‐ planations that were too advanced, insufficiently tailored to their level of study, or disconnected from the approaches used within the module itself. ”Everything there is more direct to the level of coding that is expected of us, AI and even youtube can often give you things we haven’t yet been taught or don’t understand.” Participant bano‐k7 The qualitative responses further suggested that many students en‐ gage with AI critically and iteratively, rather than relying passively on generated outputs. Several participants described refining prompts, requesting simplified explanations, verifying responses independently, and using AI systems primarily as a mechanism for clarification rather than direct answer generation. ”AI tools allow clarification using online materials. In the event I do not understand what generative AI outputs, I can ask it specific rele‐ vant questions such as what a specific sentence meant in the context,
5.1.2
Post‐Experiment Questionnaire
Post‐experiment questionnaire responses demonstrated generally pos‐ itive perceptions of Beacon across usability, relevance, learning sup‐ port, and accessibility‐related measures. Students frequently described Beacon as intuitive (73.3%) and easy to use (73.35%), with many par‐ ticipants agreeing that they felt confident navigating the interface and interacting with Beacon (66.7%). Respondents indicated that they would use Beacon again for module support (33% neutral; 46.6% agree) and would recommend it to other students (33% neutral; 40% agree). However, the findings also identified recurring usability concerns, par‐ ticularly regarding response speed, information density, and interface layout. Multiple students described Beacon’s outputs as overly long or visually overwhelming, suggesting that future versions should offer clearer formatting, collapsible sections, summarised responses, or visual distinctions between explanation categories. Despite these criticisms, the overall tone of responses suggested that students viewed Beacon as both promising and practically useful within a learning context with the majority having a positive experience using Beacon (33% neutral; 60% agree). Key to positive perceptions was Beacon’s use of module‐specific materials. Many participants identified this as Beacon’s most valuable characteristic, emphasising that responses felt grounded in the actual content, terminology, and expectations of their modules. When com‐ pared with general‐purpose AI systems such as ChatGPT, several stu‐ dents noted that unrestricted systems often generated solutions that ex‐
where it can clarify in detail for me.” Participant feno‐j5
ceeded their level of study or provided complete code solutions without supporting understanding. In contrast, Beacon was praised for produc‐
However, concerns surrounding trust and reliability were also evi‐
ing pseudocode, scaffolded explanations, and guided problem‐solving
dent throughout the dataset. Students acknowledged that AI systems
that encouraged students to engage more critically with problems rather
could sometimes provide incorrect or misleading information and recog‐
than simply copying solutions. Numerous students highlighted that Bea‐
nised the risks associated with over‐reliance on generated responses.
con “forced” them to think through solutions independently while still
In particular, some participants noted that AI explanations occasion‐
providing useful clarification and direction. When asked about the im‐
ally exceeded their current level of understanding, which could further
pact of Beacon on their learning, 53.35% of students agreed Beacon
contribute to confusion rather than resolving it.
improved their understanding of programming concepts, 53.3% felt more confident solving programming problems and 60% felt more confi‐
”I found ChatGPT hard to use as it was giving me solutions above my
dent tackling difficult topics. Consequently 66.7% indicated that Beacon
study level at the beginning.” Participant zelo‐d6
supported rather than replaced learning, while 80% agreed that Beacon helped them learn independently. These findings therefore support the
Overall, the pre‐experiment findings demonstrate that students al‐
intended pedagogical design of Beacon. Rather than being a solution
ready occupy a complex relationship with both formal academic support
generator Beacon promotes independent learning and the development
and generative AI systems. While they value independence and auton‐
of learner confidence and autonomy.
omy in their learning practices, they also experience anxiety and hesi‐ tation when seeking help through traditional channels. Simultaneously,
”traditional generative AI (such as ChatGPT or Gemini) rush to give you
students are already integrating AI tools into their learning workflows
a full answer. Resultably, I dont learn a thing from it, just the result.
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
15
Beacon however forces me to consider facts, but work out the answer
the same question depending on prompt phrasing. Additionally, a mi‐
myself.” Participant feno‐j5
nority felt responses were overcomplicated, or failed to improve their understanding compared to asking a tutor directly.
Perceptions of trust and relevance were generally positive, although participants demonstrated cautious rather than unconditional confi‐
”small problems should have small answers. Especially when the target
dence in Beacon. Many agreed that Beacon’s answers were relevant
audience is people who are too shy to ask a teacher or somewhere in
(73.3%), and aligned with module materials (93.3%), while the inclusion
that ballpark. Having a small issue and getting a massive response can
of referenced source material significantly increased trust in the gen‐
be overwhelming and make them feel further behind than they are.”
erated outputs (86.7%). Students frequently reported that grounding
Participant taro‐x5
responses in module‐specific content made Beacon feel more aca‐ demically reliable than unrestricted GenAI tools. Nevertheless, trust
Although many students believed that Beacon promoted learning
remained conditional, with most participants indicating that they would
more effectively than unrestricted generative AI tools, some partici‐
still verify important information with lecturers (40% neutral; 46.7%
pants worried that students could still become dependent on Beacon
agree) or official course resources. Several students also reported occa‐
or bypass independent critical thinking altogether. These responses
sional inaccuracies, hallucinations, or confusing responses, particularly
indicate that the effectiveness of Beacon may vary according to in‐
when Beacon lacked sufficient contextual information within the stored
dividual learning preferences, confidence levels, and prior experience.
teaching materials.
Additionally, these concerns align with broader debates surrounding the educational risks of AI‐supported learning technologies and reinforce
”It is very rigid in the sense of the second you of out of scope or ask
the need for careful pedagogical integration and transparent system
something somewhat unrelated it becomes slightly useless.” Partici‐
design.
pant raku‐n3 When asked about Beacons influence on help‐seeking behaviours many students agreed that Beacon made it easier to seek help privately (86.7%), reduced anxiety when unsure about a topic (53.3%), and could be particularly beneficial for students with low confidence (93.3%) or concerns about judgement (80%). Several participants specifically noted that Beacon reduced barriers associated with embarrassment or reluc‐ tance to ask questions publicly. Students frequently highlighted the value of immediate support outside normal staff availability, position‐ ing Beacon as a low‐friction support mechanism that students could access without fear of negative evaluation. Importantly, while many participants still viewed tutors as the most effective source of help for complex or nuanced issues, Beacon was often described as a valu‐ able “first step” before approaching teaching staff. This reinforces the broader argument that AI‐supported systems may complement rather than replace traditional academic support structures. ”Beacon was helpful because it provided immediate support and made it easier to ask questions that I might otherwise feel hesitant or embarrassed to ask in class.” Participant noku‐h7 Despite broadly positive feedback, positive experiences of Beacon were not universal. Some participants identified several important lim‐ itations and risks associated with Beacon. One of the most common criticisms concerned response speed, with many students comparing Beacon unfavourably to commercial AI systems that produce near‐ instant outputs. In some cases, students described Beacon struggling with very small code snippets, generating excessively detailed explana‐ tions for relatively simple problems, or producing different solutions to
5.1.3
Student Interviews
Analysis of the student interviews identified four key themes that compliment the findings from the pre‐ and post‐study questionnaires. Firstly, course‐grounded guidance enhances the tools learning rele‐ vance. Students valued Beacon as it restricted responses to the scope of the module. This course‐specifc grounding could prevent overly complex or irrelevant explanations that can be produced by a general‐ purpose AI. In contrast, Beacon was perceived being able to better explain concepts at an appropriate level. ”With Beacon, I really found it useful that it was actually based on the exact knowledge we were taught in class.” Student Interview 3 In turn this reduced the need for students to navigate explanations that assumed greater prior knowledge or introduced unfamiliar con‐ cepts. Its ability to direct students to relevant slides, sessions, and exam‐ ples also reduced the time and knowledge required to locate appropriate academic resources via a virtual learning environment. Furthermore, the 24/7 availability of Beacon was acknowledged as a benefit for stu‐ dents who lacked peer‐support networks or needed assistance outside scheduled teaching hours. Secondly, the psychological safety afforded by Beacon reduces help‐ seeking hesitation. In the interviews several students described feeling embarrassed or nervous about asking questions and acknowledged be‐ ing anxious about revealing that they may be struggling with problems perceived as basic or fundamental. ”I think for me, it was very useful because as you know, I never coded before. And some of the questions, like what I was thinking of asking in
16
the first year, I felt it was embarrassing. Like, oh my God, I don’t even
that can just give you the answer, simply because there’s little to no
know this, you know.” Student Interview 3
friction to using it in comparison to Beacon where you actually have to work things out.” Student Interview 4
Beacon was viewed as being able to provide a private and non‐ judgemental first point of support, which could make help‐seeking more
This friction is compounded by user‐interface and user‐experience
accessible to students who lacked confidence or were reluctant to
issues students identified when using Beacon. The responses given by
approach teaching staff. However, rather than merely replace tutors,
Beacon are highly templated which results in lengthy responses ‐ both
students suggested Beacon acted as a bridge, allowing them first to de‐
in terms of time and word count ‐ even if the question is short. Fur‐
velop confidence before seeking tutor support. This therefore suggests
thermore, the templated nature of the responses restricts Beacon’s
that Beacon can support help‐seeking behaviour rather than eliminate
ability to have conversational style interactions which are the norm in
it.
mainstream AI tools. These issues risked responses having unnecessary complexity which may in turn limit Beacons ability to support those in ”it’s kind of encouraging more questions, I think, which would even‐
most need of assistance.
tually, I think, would enhance knowledge and marks because people would be more free to ask something like Beacon rather than go to
”it would come back with a with like the most full‐on response it possi‐
you.” Student Interview 3
bly can with like all of these different steps to the answer. But it’s like, just give me a quick answer. If I want more information, let me ask for
Thirdly, trust emerged through course‐grounding and transparency. Students perceived Beacon’s responses as trustworthy because they
more information. Don’t give it all to me at once.” Student Interview 1
could see the source of information. This trust was further reinforced by the course‐grounding of the materials. As Beacon stayed within
These issues represent an ongoing challenge in the design of similar
the module boundaries students trusted responses reflected what was
tools. A clear tension exists between providing sufficient information
taught in‐class and therefore felt confident that the information pro‐
to demonstrate the value and encourage engagement in the learning
vided was consistent with module expectations and could support them
process, versus withholding too much information so students circum‐
in meeting assessment requirements.
vent the intervention in favour of mainstream AI tools that provide the answer rather than support learning.
”Beacon would know roughly more of the scope and how we’ve worked.
In summary, across the five interviews Beacon’s was viewed posi‐
And I think that would be really good for like creating the ideas [...]
tively for providing private, timely, appropriately scoped, and pedagogi‐
Because when you give a brief to an AI, it doesn’t. It’s not always the
cally constrained assistance. However, its capacity to address disparities
best. I would hope that Beacon, because it has more of the resources,
depended on careful calibration. Excessive guidance could enable task
would understand, oh, okay, we’re going to do this and it’s going to use
completion without learning, while excessive restriction, lengthy re‐
this.” Student Interview 5
sponses, and inflexible presentation could introduce new barriers. The findings therefore suggest that equitable course‐specific RAG requires
Finally, across the interviews a pedagogical tension between scaf‐ folded learning and information convenience emerged. Students found value in Beacon’s attempt to encourage learning, rather than provide direct answers.
a balance between accessibility, productive cognitive effort, conversa‐ tional adaptability, and continued access to human and peer support. Overall, the findings of the post‐experiment questionnaires and in‐ terviews suggest that Beacon successfully occupied a middle ground between independent learning and formal academic support. Students
”So something like ChatGPT would just go, here’s the answer. Whereas
valued Beacon’s ability to provide context‐specific, module‐aligned as‐
Beacon basically makes you do a bit more work and gives you some
sistance while still encouraging active engagement with problems. The
of the steps towards reaching that answer rather than just being like,
findings also indicate that students perceived Beacon as more peda‐
here you go.” Student Interview 4
gogically supportive than unrestricted generative AI tools because it constrained outputs within the boundaries of module expectations and
However, students also cited this as a point of frustration, especially
promoted scaffolded learning rather than answer substitution. How‐
when Beacon would be overly restrictive or refuse to give answers be‐
ever, the results demonstrate that trust remained conditional, students
yond the scope of the module. Students suggested that this creates a
continued to value human tutors for deeper support, and significant
risk where students who are time‐pressured, or not motivated to en‐
usability refinements are necessary before broader deployment. Taken
gage with learning the material, would circumvent Beacon and use less
together, these findings provide evidence that course‐specific Retrieval‐
restrictive tools.
Augmented Generation systems may help reduce barriers to academic support while reinforcing independent study practices within HE com‐
”Students who will use the quick route, which happen regularly any‐ way, would typically just use one of the mainstream language models
puting contexts.
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
5.2
Academics’ Views
Academic staff, like the students, saw value in Beacon as a scaffolded learning companion. In particular, the use of pseudocode, simple start‐ ing structures and links to relevant module materials were highlighted for providing structured learning guidance, rather than direct answers. This approach was considered appropriate for novice programmers be‐ cause it encouraged students to translate guidance into working syntax and learn through implementation. I think the main benefit to Beacon is definitely the explanation, the way that it can clearly explain step by step what everything means compared to other sort of chat bots or AI tools that will simply kind of give you the answer [...] and sometimes the AI expects you to know certain knowledge, but this is where sort of beacon is. It has the step up to it, it can actually explain core fundamentals of certain concepts.” Staff Interview 1 Scaffolding was further aided by Beacons grounding in module mate‐ rials, which academic staff viewed as being a key mechanism in building trust in the responses. Consequently, staff felt Beacon could provide a more focused and trustworthy alternative to online searches and general‐purpose AI, which may expose novice students to irrelevant, overly advanced, or pedagogically inappropriate information. Further‐ more, this curriculum grounding was considered valuable in ensuring Beacon reflected what tutors intended students to learn and was able to communicate the practices and perspectives valued within the course rather than reproducing the average or dominant position found across internet sources. it’s focused on code lab resources, so that you know, it’s not going to lead you astray into a bunch of really confusing stuff, that’s great, right? [...] Like there’s just the way in which when you’re trying to learn something, it’s really helpful to not be trying to deal with a synthesis of everything that’s ever been said by anyone anywhere on the internet.” Staff Interview 2 Regarding help‐seeking, academic staff also viewed Beacon as an accessible gateway to academic support. Beacon was deemed benefi‐ cial for students who may have missed class, need concepts repeated, lacked confidence or require support outside normal teaching hours. ”I think this is probably most useful for students who you know, didn’t get it the first time or weren’t there the first time maybe, and so kind of need a bit of repeat explanation to sort of work through it.” Staff Interview 2 However, rather than replace the tutor, Beacon was viewed as be‐ ing complementary to tutor support and could allow students to seek initial guidance independently before escalating unresolved questions
17
to teaching staff. This compliments the students views of Beacon pro‐ viding a safe private space to develop confidence before seeking tutor support. ”I think this is one really good step for students to be able to go, okay, let’s try and learn it through this tool, see how far I can get with it with the clear explanations. And then if they don’t get it, they can go one step higher.” Staff Interview 1 Although staff saw the positive benefits of Beacon, they also re‐ mained cautious and had concerns it may allow students to bypass the learning process. Staff noted unease with the ability of existing AI tools to give direct answers and the need for Beacon to support knowledge development rather than replace it by generating solutions. ”this is a good one because I think ChatGPT does a lot for you. I think people are skipping that development stage.” Staff Interview 3 To this end staff valued Beacon because it resisted giving answers, but similar to student sentiments felt the nature of the responses were sometimes too rigid. The templated nature of responses meant staff noted the answers sometimes felt unnatural, lacked concision, or were unable to provide a natural conversational style experience typical of other available tools. Additionally, an ability to manipulate responses was noted, meaning it was on occasion possible to elicit a more direct answer. ”when kind of probing it to give some actual code, I think at first it gave the pseudo code and then I asked it to adapt it in a certain way and then it did give the C++ code” Staff Interview 1 Staff also noted that whilst it knows the context of the module, it does not know the students context or level of knowledge (e.g. strug‐ gling with the fundamentals, or already beyond module scope). Staff emphasised that access to accurate information does not necessar‐ ily constitute accessible or educationally appropriate support. Lengthy and technically dense responses could increase the cognitive demands placed on students who were already confused. Conversely, inconsis‐ tent guardrails could allow students to obtain complete or assessment‐ relevant code without undertaking the intended learning. Responses therefore needed to be sensitive to the student’s current knowledge, the type of question being asked and the relevant stage of the curricu‐ lum. ”this is the thing that I think people struggle with. How do you give explanations that make sense to people who don’t already get it? And I think I think the LLM suffer a little bit from [...] You know, they always want to give you the Wikipedia style. Here is this thing explained by people who know what it is.” Staff Interview 1 Thus, Beacon in its current form was noted as being best suited for those in the middle range regarding skill and understanding
18
”I think it would be generally useful for the students kind of in the mid‐ dle, maybe a little bit on the earlier side, to just kind of get some more familiarity and just kind of be more immersed in some of the language as well.” Staff Interview 4 To make Beacon more adaptable staff suggested adding options for personalisation, either via on‐screen settings or onboarding questions so Beacon could gain an understanding of the individual learner with responses then tailored accordingly. ”if someone can just choose on a toggle, like, you know, explain like I’m five mode or I get it mode, that’s interesting [...] you know, pitch the response. Maybe there’s something even to say that as time goes on, like literally as the module is progressing, the prompt changes, right?.” Staff Interview 2 Finally, staff also highlighted an inherent tension between educa‐ tional intent and student expectations (or desires). While restricting Beacon to taught content can preserve learning outcomes and prevent students from using advanced techniques prematurely, this same restric‐ tion could cause frustration among students who want a quick answer and thus abandon the system in favour of unrestricted generative AI. I think putting on several different student hats, I can definitely see this as a good way to learn sort of step by step in getting the understanding of a concept for sure. I think putting on another student hat, I can see how it’s frustrating that it just doesn’t give the answer straight away.” Staff Interview 1 Overall, staff interviews, like the students, positioned Beacon as an intermediate layer in the help‐seeking process, providing timely access to trusted course resources before students escalated difficulties to teaching staff. Academics identified particular value for students with limited prior knowledge, those who had missed teaching, and those re‐ luctant to ask questions in class. Its curriculum grounding could reduce the navigational and evaluative burden associated with online resources while protecting novice learners from overly advanced, irrelevant or poor‐quality information. However, overall educational value depends on adaptive scaffolding, persistent guardrails, cognitively accessible pre‐ sentation, transparent curriculum boundaries and clear routes to human assistance. Beacon should therefore supplement rather than replace academic and peer support.
5.3.1 RQ1: How do students perceive the usefulness and relevance of a course‐specific RAG system for academic support? Students particularly valued Beacon’s ability to explain concepts at the level of the module and direct them towards relevant slides, sessions, and examples. This contrasted with search engines, online tutorials, and general‐purpose generative AI, which could assume greater prior knowl‐ edge, introduce concepts beyond the curriculum, or provide complete solutions without supporting understanding. However, while Beacon’s attempt to preserve active learning was seen as useful, the findings revealed inconsistency in how effectively this was achieved. Some par‐ ticipants believed that Beacon introduced productive friction by requir‐ ing students to work towards a solution. Others found that it could be prompted to provide substantial or complete code, potentially enabling task completion without meaningful engagement. Persistent guardrails are therefore required across the full conversation, rather than only within the initial response. Usability was also a consistent concern. Lengthy and templated re‐ sponses sometimes made accurate information difficult to locate or understand. This is especially significant because students seeking sup‐ port may already be confused or have limited confidence in the subject. Participants recommended concise initial answers, clearer presentation of code and pseudocode, direct links to source materials, and optional expansion of explanations. More conversational follow‐up responses and greater student control over response length, format, and level of technical detail could improve cognitive accessibility.
5.3.2 RQ2: To what extent do students perceive Beacon as reducing barriers associ‐ ated with academic help‐seeking? Beacon was perceived as a useful first step in the help‐seeking process with findings suggesting that Beacon is best understood not simply as an AI tutor, but as an intermediary form of academic support between inde‐ pendent study and formal help‐seeking. Across the questionnaires and interviews, participants valued the ability to access private, immediate, and module‐aligned guidance while continuing to regard lecturers and tutors as important sources of support for more complex or nuanced dif‐ ficulties. The value of Beacon therefore appeared to lie not in replacing human support, but in lowering the threshold for initiating help‐seeking and enabling students to develop sufficient understanding either to con‐
5.3
Summary of Findings
Taken together, the findings provide answers to the studies four guiding research questions.
tinue independently or to seek more focused assistance from teaching staff. Beacon therefore extends the support ecosystem by providing a private and module‐aligned first point of assistance. This is important be‐ cause disparities in academic support are not only produced by whether support exists, but also by whether students feel able to access it. Although this study did not collect demographic data linking help‐ seeking behaviour to specific learner characteristics (Section 6), the mechanism identified here has plausible relevance to the demographic
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
19
disparities that generative AI in education is often expected to ad‐
Overall, the findings suggest that course‐specific RAG has the poten‐
dress. Confidence and anxiety driven avoidance of help‐seeking has
tial to reduce help‐seeking disparities associated with prior knowledge,
been associated with gender (Huang 2013, Brown et al. 2021), first‐
confidence, attendance, resource navigation, and the availability of
generation and widening‐participation status (Koh et al. 2022, Li et al.
teaching staff. However, equitable access depends on more than making
2023), mature‐student status (Chapman 2017), and disability and neuro‐
assistance continuously available. The support must also be cognitively
divergence (Anselimus 2026), among other factors. Generative AI tools
accessible, responsive to individual learning needs, appropriately scaf‐
more broadly carry both the potential to narrow and the risk of widen‐
folded, and integrated with existing academic and peer‐support relation‐
ing these disparities, depending on how they are designed and deployed
ships. Beacon’s potential therefore lies in balancing four requirements:
(Ni et al. 2026, James et al. 2024, Adejumo et al. 2026). Beacon’s de‐
lowering the threshold for initiating help‐seeking, maintaining produc‐
sign ‐ offering private, low‐friction, course‐aligned support outside the
tive cognitive effort, adapting assistance to the learner and question,
scrutiny of staff‐mediated or public learning environments ‐ targets pre‐
and providing clear routes to human support when automated guidance
cisely the anxiety‐ and confidence‐related mechanisms most frequently
is insufficient.
implicated in these disparities. This suggests a plausible, though as yet
The findings therefore offer some practical insight for the design
untested, pathway by which course‐specific RAG systems could nar‐
of course‐specific RAG systems. Success depends less on retrieval
row rather than widen educational disparities for the groups this special
accuracy and more on how effectively they balance curriculum ground‐
issue is concerned with; establishing whether this pathway holds in
ing, trust, scaffolded learning and psychological support. Through
practice is the central task of the demographic‐stratified evaluation
curriculum‐grounding responses become more useful and relevant to
proposed in Section 7.
the students context and more accurately reflect the tutor intent and intended learning outcomes. Consequently, this leads to greater trust
5.3.3 RQ3: How do students evaluate the trustworthiness of responses generated by a system grounded in course materials? The grounding of Beacon in lecturer‐created materials was deemed cen‐ tral to enhancing the trustworthiness of responses. This trust emerged from transparency. As information sources were clearly attributed to lecture slides and leture material was often directly cited students felt the response were accurate and trusted explanations were consistent with the module’s intended level and learning outcomes.
in the responses, especially when information sources are transparently attributed to provide a platform to find out more and review existing study materials. Furthermore, design should prioritise guidance over di‐ rect answers to scaffold learning, though careful balance is required to avoid too much friction and meet student expectations for convenience. Finally, by providing a private space to seek help, these systems offer psychological safety that can influence help‐seeking behaviour. Thus, the design of these systems should augment not replace human‐support by including psychological safeguards for adaptive guidance to meet dif‐ ferent learner skills and encourage human escalation for persistent or complex problems.
5.3.4 RQ4: How do academic staff per‐ LIMITATIONS ceive Beacon’s value in reducing barriers to stu‐ 6 dent help‐seeking, and what tensions does this raise for maintaining pedagogical safeguards This study has several limitations. The evaluation was conducted within a single course context at a single institution, and involved a rela‐ and scaffolded learning tively small number of participants. Consequently, the findings may not
Staff shared student views regarding Beacon’s ability to reduce help‐ seeking barriers. Benefits were perceived for students who lacked prior programming experience, had missed teaching, struggled to locate rele‐ vant resources, were working outside scheduled contact hours, or felt embarrassed about asking questions they perceived as basic. Scaffolded learning, curriculum grounding and resistance to providing direct an‐ swers were seen as key to this value, yet concerns were raised about the amount of friction these safeguards imposed. Students seeking rapid solutions might bypass Beacon in favour of unrestricted generative AI, while excessive restriction could prevent the system from providing sufficiently useful assistance. Therefore an ongoing challenge is main‐ taining the balance between pedagogical safeguards alongside student expectations from fast low‐friction AI support.
generalise to other disciplines, institutions, or student populations. Beacon was also evaluated as a prototype rather than as part of a fully integrated learning management environment. Although partici‐ pants were provided task scenarios designed to reflect real‐world usage, it is acknowledged that these do not fully capture the range of ques‐ tions students may ask. Furthermore, these tasks were conducted in a single sitting by each participant, thus usage does not reflect on‐going interaction that would be expected if deployed in a live module setting. Consequently, the perceived usefulness and relevance of the tool may not accurately reflect its actual impact on module outcomes or its ability to address students’ real‐time learning needs. A related limitation concerns the verification of response accuracy. Evidence of Beacon’s reliability in this study relied on self‐reported student and staff perceptions rather than independent verification of
20
Beacon’s outputs. Although participants reported specific instances of
change as students become more confident, and whether Beacon en‐
hallucination and guardrail circumvention, these were identified anec‐
courages greater engagement with module materials. Such research
dotally rather than systematically quantified. Future evaluations would
would also help determine whether AI‐supported assistance comple‐
benefit from combining perception‐based measures with an indepen‐
ments existing academic support structures or whether there is a risk
dent assessment of response accuracy.
of students becoming overly reliant on automated guidance. Examining
The study also did not examine differences in system use or per‐
long‐term outcomes would therefore provide a clearer understanding
ceived benefit across specific learner groups, such as students with
of the pedagogical value of course‐specific RAG systems and their role
disabilities, neuro‐divergent students, students from different socioeco‐
within HE learning support.
nomic backgrounds, or students with varying language backgrounds. As
Future work should also examine whether Beacon supports different
a result, while the findings suggest that Beacon may reduce some barri‐
student groups equitably, including students with lower confidence, stu‐
ers to help‐seeking, further research is required to determine whether
dents with disabilities, neuro‐divergent students, students from widen‐
such systems reduce educational disparities in practice and for whom
ing participation backgrounds, and students with different levels of
they are most effective.
prior programming experience. This would allow future research to
Finally, this study did not evaluate Beacon’s potential to reduce aca‐
move beyond general perceptions of usefulness and examine whether
demic misconduct arising from unrestricted use of general‐purpose
course‐specific AI support can measurably reduce disparities in access
generative AI, despite this being one of the original motivations for de‐
to academic help.
velopment. Beacon was evaluated with students after their module had
Finally, future work should examine whether course‐specific RAG
already been completed, rather than during live, assessed coursework,
systems such as Beacon can measurably reduce academic misconduct
so there was no opportunity to observe whether its use influenced
arising from unrestricted generative AI use, given the potential pathway
submitted work or the incidence of misconduct cases. This remains an
from out‐of‐scope AI‐generated content to misconduct proceedings
anticipated rather than demonstrated benefit of the system.
discussed in Section 3. This would require deployment during a live, assessed module, comparing misconduct referral rates, or the propor‐ tion of flagged submissions attributable to out‐of‐scope AI‐generated
7
FUTURE WORK
content, before and after the system’s introduction.
Future work will explore the scalability of Beacon across multiple modules and investigate how such tools might be integrated into in‐
8
CONCLUSION
stitutional learning platforms. While this study focused on a specific computing context, a key next step is to examine whether the same RAG
The principal contribution of this study extends beyond the technical
approach can be applied across different modules, levels of study, and
evaluation of a course‐specific RAG system. Rather than viewing gener‐
subject areas. This would require consideration of how module mate‐
ative AI solely as a productivity tool or a potential source of academic
rials are collected, structured, updated, and indexed so that responses
misconduct, this study demonstrates how AI can be designed to ad‐
remain accurate and aligned with current teaching content. Scaling
dress educational disparities by improving access to timely, contextually
would also involve evaluating whether Beacon can maintain response
relevant academic support. Students who may otherwise hesitate to
quality when drawing from larger and more diverse knowledge bases,
ask questions due to anxiety, fear of judgement, lack of confidence, or
particularly where modules differ in terminology, assessment design,
limited staff availability can instead access immediate guidance that re‐
and expected student prior knowledge. Integration into institutional
mains aligned with module expectations. In this way, Beacon provides
learning platforms would allow students to access support alongside
the potential for a more equitable model of learning while preserving op‐
lecture slides, workshop tasks, assessment briefs, and other approved
portunities for independent thinking and meaningful engagement with
resources. This could reduce barriers to use by embedding AI‐supported
course materials.
assistance within the existing learning ecosystem rather than requiring students to access a separate tool.
As GenAI becomes increasingly embedded within HE, the question is no longer whether students will use AI to support their learning, but how
Additional research is also required to examine the long‐term impact
institutions can shape that use responsibly and equitably. This study ar‐
of AI‐supported academic assistance on student learning outcomes. The
gues that course‐specific Retrieval‐Augmented Generation offers one
present study captured students’ immediate perceptions of Beacon fol‐
promising approach. By grounding AI responses within approved edu‐
lowing use, but it does not establish whether sustained engagement
cational resources and positioning AI as a complement rather than a
leads to measurable improvements in confidence, conceptual under‐
replacement for educators, institutions may be able to reduce barriers
standing, assessment performance, or help‐seeking behaviour over time.
to help‐seeking while maintaining trust, academic integrity, and peda‐
Future longitudinal studies could investigate how students use the sys‐
gogical quality. Ultimately, the greatest opportunity for educational AI
tem across an entire module or academic year, whether usage patterns
may lie not in replacing existing support structures, but in making them more accessible to the students who need them most.
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
AUTHOR CONTRIBUTIONS Andy Gray: Writing – review & editing, Writing – original draft, Con‐ ceptualisation, Software, Project administration, Methodology, Investi‐ gation, Formal analysis, Data curation, Funding acquisition. Jake Hobbs: Writing – review & editing, Writing – original draft, Data curation, Funding acquisition, Methodology, Investigation, Formal analysis, Project administration.
ACKNOWLEDGEMENTS For the purpose of Open Access, the author has applied a CC‐BY public copyright licence to any Author Accepted Manuscript (AAM) ver‐ sion arising from this submission. All underlying data to support the conclusions are provided within this paper.
FUNDING INFORMATION This project was conducted as a result of acquiring research funding from Bath Spa Universities Creativity & Curiosity Projects Fund.
CONFLICT OF INTEREST STATEMENT The authors declare no potential conflict of interest.
ETHICS STATEMENT Ethics approval was obtained from the School of Design ethics com‐ mittee at Bath Spa University (Research Ethics Approval Number: 240226JH).
CONSENT TO PARTICIPATE AND PUBLICATION All participants completed a consent from before data collection took place. Participants consented to anonymised quotes being published.
DATA AVAILABILITY STATEMENT All the data presented in the manuscript were obtained with the par‐ ticipants’ consent to be published. Full data relating to this study are available on request from the author(s).
AI USE DECLARATION STATEMENT During the preparation of this work author(s) used Claude (Anthropic; Sonnet 5) to assist with drafting, summarising, proofreading, refining academic content and to help identify potentially relevant literature. In addition Microsoft Co‐Pilot was used to support data analysis as de‐ scribed in the methodology section. After using these tools the author(s) reviewed and edited content as needed and assume full responsibility for the content of the manuscript.
REFERENCES
Adadi, A. & Berrada, M. (2018) Peeking inside the black‐box: A survey on explainable artificial intelligence (xai). IEEE Access, 6, 52138–52160. doi:10.1109/ACCESS.2018.2870052. Adejumo, A.A. et al. (2026) A systematic review of the impact of genai on learning performance, ai hallucinations, and problem‐ solving in computer science education. Computers and Education: Artificial Intelligence,. Alsafari, B. et al. (2024) Towards effective teaching assistants: From intent‐based chatbots to llm‐powered teaching assistants. Natural Language Processing Journal,.
21
Anselimus, S. (2026) Generative ai for university students with disabilities: challenges and opportunities. Educational Research,. Baker, R.S. & Smith, L. (2019) Artificial intelligence in education: Promise and implications for teaching and learning. OECD Educa‐ tion Working Papers,(218). Becker, C.R. (2020) Learn Human‐Computer Interaction: Solve human problems and focus on rapid prototyping and validating solutions through user testing. : Packt Publishing Ltd. Bender, E.M., Gebru, T., McMillan‐Major, A. & Shmitchell, S. (2021) On the dangers of stochastic parrots: Can lan‐ guage models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency„ 610–623.doi:10.1145/3442188.3445922. Blandford, A. (2013) Eliciting people’s conceptual models of activi‐ ties and systems. International Journal of Conceptual Structures and Smart Applications (IJCSSA), 1(1), 1–17. Blandford, A., Green, T.R., Furniss, D. & Makri, S. (2008) Evaluating system utility and conceptual fit using cassm. International Journal of Human‐Computer Studies, 66(6), 393–409. Blandford, A. & Rugg, G. (2002) A case study on integrating contex‐ tual information with analytical usability evaluation. International journal of human‐computer studies, 57(1), 75–99. Bommasani, R., Hudson, D.A., Adeli, E. et al. (2021) On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258,. Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Mil‐ lican, K. et al. Improving language models by retrieving from trillions of tokens. In: International conference on machine learning. PMLR, 2022, pp. 2206–2240. Bornschlegl, M., Meldrum, K. & Caltabiano, N.J. (2020) Variables related to academic help‐seeking behaviour in higher education– findings from a multidisciplinary perspective. Review of Education, 8(2), 486–522. Broadbent, J. & Howe, W.D.W. (2023) Help‐seeking matters for on‐ line learners who are unconfident. Distance Education, 44, 106 – 119. doi:10.1080/01587919.2022.2155616. Brown, D., Barry, J.A. & Todd, B.K. (2021) Barriers to academic help‐ seeking: The relationship with gender‐typed attitudes. Journal of Further and Higher Education, 45(3), 401–416. Brown, T.B. (2020) Language models are few‐shot learners. arXiv preprint arXiv:2005.14165,. Caines, A., Benedetto, L., Taslimipoor, S., Davis, C., Gao, Y., Ander‐ sen, O. et al. (2023) On the application of large language models for language teaching and assessment technology. arXiv preprint arXiv:2307.08393„ 173–197.doi:10.48550/arXiv.2307.08393. Chapman, A. (2017) Using the assessment process to overcome imposter syndrome in mature students. Jour‐ nal of Further and Higher Education, 41, 112 – 119. doi:10.1080/0309877X.2015.1062851. Cohen, L., Manion, L., Morrison, K. et al. (2009) Research methods in education. Vol. 6. : routledge London. Department for Education (2023) Generative artificial intelligence (ai) in education. Department for Education, Last updated 12 August 2025. Policy paper. https://www.gov.uk/government/publications/ URL generative‐artificial‐intelligence‐in‐education/ generative‐artificial‐intelligence‐ai‐in‐education Dix, A. (2016) What is human‐computer interaction (hci). Interaction Design Foundation. https://acortar. link/mETsan,. Duke, R., Salzman, E., Burmeister, J., Poon, J. & Murray, L. Teaching programming to beginners‐choosing the language is just the first step. In: Proceedings of the Australasian conference on Computing education, 2000, pp. 79–86. Eager, B. & Brunton, R. (2023) Prompting higher education towards ai‐augmented teaching and learning practice. Journal of University Teaching and Learning Practice,.doi:10.53761/1.20.5.02.
22
Faulkner, L. (2003) Beyond the five‐user assumption: Bene‐ Jisc (2023) Generative ai: implications and considerations for education. fits of increased sample sizes in usability testing. Behavior Jisc. Report. Research Methods, Instruments, & Computers, 35(3), 379–383. URL https://www.jisc.ac.uk/reports/generative‐ai Kamsin, A., Blandford, A. & Cox, A.L. (2012) Personal task manage‐ doi:10.3758/BF03195514. Fong, C.J., Gonzales, C., Hill‐Troglin Cox, C. & Shinn, H.B. (2023) Aca‐ ment: my tools fall apart when i’m very busy! In: CHI’12 Extended demic help‐seeking and achievement of postsecondary students: Abstracts on Human Factors in Computing Systemspp. 1369–1374. A meta‐analytic investigation. Journal of Educational Psychology, Karabenick, S.A. & Knapp, J.R. (1991) Relationship of academic help 115(1), 1. seeking to the use of learning strategies and other instrumental Furniss, D., Blandford, A. & Curzon, P. Confessions from a grounded achievement behavior in college students. Journal of educational theory phd: experiences and lessons learnt. In: Proceedings of the psychology, 83(2), 221. SIGCHI Conference on Human Factors in Computing Systems, 2011, Khosravi, H., Shum, S.B., Chen, G., Conati, C., Tsai, Y.S., Kay, pp. 113–122. J. et al. (2022) Explainable artificial intelligence in educa‐ Glikson, E. & Woolley, A.W. (2020) The human–machine trust rela‐ tion. Computers and Education: Artificial Intelligence, 3, 100074. doi:10.1016/j.caeai.2022.100074. tionship: Going beyond trust in automation. Computers in Human Behavior, 111, 106446. doi:10.1016/j.chb.2020.106446. Koh, J., Farruggia, S.P., Back, L.T. & Han, C.w. (2022) Self‐efficacy Gonsalves, C. & Lin, Z. (2025) Clear in advance to whom? explor‐ and academic success among diverse first‐generation college stu‐ ing ‘transparency’of assessment practices in uk higher education dents: The mediating role of self‐regulation. Social Psychology of institution assessment policy. Studies in Higher Education, 50(7), Education, 25(5), 1071–1092. 1454–1470. Krisi, M. & Nagar, R. (2021) The effect of peer mentoring on men‐ Gray, A., Lindsay, S., Pearson, J., Crick, T. & Rahat, A. (2026) tors themselves: A case study of college students. International Rendering transparency to ranking in educational assessment Journal of Disability, Development and Education, 70, 803 – 815. via bayesian comparative judgement. Review of Education, 14(2), doi:10.1080/1034912X.2021.1910934. e70149. Laato, S., Morschheuser, B., Hamari, J. & Björne, J. (2023) Gray, A., Rahat, A., Crick, T. & Lindsay, S. (2024) A bayesian ac‐ Ai‐assisted learning with chatgpt and large language mod‐ tive learning approach to comparative judgement within educa‐ els: Implications for higher education. 2023 IEEE International tion assessment. Computers and Education: Artificial Intelligence, 6, Conference on Advanced Learning Technologies (ICALT)„ 226– 100245. 230.doi:10.1109/ICALT58122.2023.00072. Gray, A., Rahat, A., Crick, T. & Lindsay, S. (2025) Bayesian active Lan, M. & Zhou, X. (2025) A qualitative systematic review on ai em‐ learning for multi‐criteria comparative judgement in educational powered self‐regulated learning in higher education. npj Science assessment. arXiv preprint arXiv:2503.00479,. of Learning, 10(1), 21. Grayson, A., Clarke, D. & Miller, H. (1998) Help‐seeking Lang, G. et al. (2025) Ai‐powered learning support: A study of among students: Are lecturers seen as a potential retrieval‐augmented generation (rag) chatbot effectiveness in an source of help? Studies in Higher Education, 23, 143–155. online course. Information Systems Education Journal,. doi:10.1080/03075079812331380354. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, Greer, M. (2024) Illustration of retrieval‐augmented N. et al. (2020) Retrieval‐augmented generation for knowledge‐ generation (rag). https://community.intel. intensive nlp tasks. Advances in Neural Information Processing Sys‐ com/t5/Blogs/Thought‐Leadership/Big‐Ideas/ tems, 33, 9459–9474. Li, R. et al. (2023) College student’s academic help‐seeking behavior: RAG‐Infusing‐Realism‐and‐Domain‐Expertise‐into‐LLMs‐but‐could/ A systematic literature review. Behavioral Sciences,. post/1574493, figure from Intel Community blog post, accessed 12 Mar 2026. Luckin, R. & Holmes, W. (2016) Intelligence unleashed: An argument Hill, H.M.M., Zwahr, J. & Gonzalez III, A. (2022) Evaluating research for ai in education,. self‐efficacy in undergraduate students: experience matters. Jour‐ Marwan, S., Dombe, A. & Price, T.W. Unproductive help‐seeking in programming: What it is and how to address it. In: Proceedings of nal of the Scholarship of Teaching and Learning, 22(1). Hoadley, C. & Campos, F.C. (2022) Design‐based research: What the 2020 ACM conference on innovation and technology in computer it is and why it matters to studying online learning. Educational science education, 2020, pp. 54–60. Psychologist, 57(3), 207–220. McCloskey, M. (2014) Turn user goals into task scenarios for usability Holmes, W., Bialik, M. & Fadel, C. (2019) Artificial Intelligence in Edu‐ testing. cation: Promises and Implications for Teaching and Learning. : Center URL https://www.nngroup.com/articles/ for Curriculum Redesign. task‐scenarios‐usability‐testing/ Huang, C. (2013) Gender differences in academic self‐efficacy: A Mcfarlane, K. (2016) Tutoring the tutors: Supporting effective per‐ meta‐analysis. European journal of psychology of education, 28(1), sonal tutoring. Active Learning in Higher Education, 17, 77 – 88. 1–35. doi:10.1177/1469787415616720. Izacard, G. & Grave, E. (2020) Leveraging passage retrieval with Miao, H., Guo, R. & Li, M. (2025) The influence of research self‐ generative models for open domain question answering. arXiv efficacy and learning engagement on ed. d students’ academic preprint arXiv:2007.01282,. achievement. Frontiers in Psychology, 16, 1562354. James, T. et al. (2024) Levelling the playing field through genai: Har‐ Micari, M. & Calkins, S. (2021) Is it ok to ask? the impact of instruc‐ nessing artificial intelligence to bridge educational gaps for equity tor openness to questions on student help‐seeking and academic and disadvantaged students. Widening Participation and Lifelong outcomes. Active Learning in Higher Education, 22(2), 143–157. Learning,. Morales‐Navarro, L., Giang, M.T., Fields, D.A. & Kafai, Y.B. (2024) Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y. et al. (2023) Survey Connecting beliefs, mindsets, anxiety and self‐efficacy in com‐ of hallucination in natural language generation. ACM Computing puter science learning: An instrument for capturing secondary Surveys, 55(12), 1–38. doi:10.1145/3571730. school students’ self‐beliefs. Computer Science Education, 34(3), Jin, S.H., Im, K., Yoo, M., Roll, I. & Seo, K. (2023) Supporting students’ 387–413. self‐regulated learning in online learning using artificial intelli‐ Naeem, M., Smith, T. & Thomas, L. (2025) Thematic analysis and gence applications. International Journal of Educational Technology artificial intelligence: A step‐by‐step process for using chatgpt in in Higher Education, 20(1), 37. thematic analysis. International Journal of Qualitative Methods, 24, 16094069251333886.
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
Nazaretsky, T., Ariely, M., Cukurova, M. & Alexandron, G. (2022) Teachers’ trust in ai‐powered educational technology and a pro‐ fessional development program to improve it. British Journal of Educational Technology, 53(4), 914–931. doi:10.1111/bjet.13179. Németh, R. et al. (2025) Exploring the use of retrieval‐augmented generation models in higher education: A pilot study on artificial intelligence‐based tutoring. Social Sciences & Humanities Open,. Neumann, A. et al. (2025) An llm‐driven chatbot in higher educa‐ tion for databases and information systems. IEEE Transactions on Education,. Newton, P.M. & Miah, M. (2017) Evidence‐based higher education– is the learning styles ‘myth’important? Frontiers in psychology, 8, 241866. Ni, W. et al. (2026) Mapping the impact of generative ai in higher ed‐ ucation: a scoping review of psychological and equity dimensions. Frontiers in Psychology,. Nielsen, J. (2000) Why you only need to test with 5 users. URL https://www.nngroup.com/articles/ why‐you‐only‐need‐to‐test‐with‐5‐users/ Quality Assurance Agency for Higher Education (2023) The future of assessment: Artificial intelligence in education. Quality Assurance Agency for Higher Education. Guidance. URL https://www.qaa.ac.uk/docs/qaa/guidance/ artificial‐intelligence‐in‐education.pdf Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I. et al. (2019) Language models are unsupervised multitask learn‐ ers. OpenAI blog, 1(8), 9. Robins, A., Rountree, J. & Rountree, N. (2003) Learning and teaching programming: A review and discussion. Computer science educa‐ tion, 13(2), 137–172. Robiullah, A., Quadrelli, L., Remache, L., Akolgo, D.R., Ramirez, G., Covarrubias, R. et al. (2026) Navigating the hidden curriculum: A study of resource‐based and stories‐based interventions in higher education. Behavioral Sciences, 16(2), 273. Rosenstein, A., Raghu, A. & Porter, L. Identifying the prevalence of the impostor phenomenon among computer science students. In: Proceedings of the 51st ACM technical symposium on computer science education, 2020, pp. 30–36. Ruihua, L., Che Hassan, N. & Saharuddin, N. (2025) Understanding academic help‐seeking among first‐generation college students: a phenomenological approach. Humanities and Social Sciences Com‐ munications, 12(1), 56. Ryan, A.M., Gheen, M.H. & Midgley, C. (1998) Why do some stu‐ dents avoid asking for help? an examination of the interplay among students’ academic efficacy, teachers’ social–emotional role, and the classroom goal structure. Journal of educational psy‐ chology, 90(3), 528. Saeli, M., Perrenet, J., Jochems, W.M. & Zwaneveld, B. (2011) Teach‐ ing programming in secondary school: A pedagogical content knowledge perspective. Informatics in education, 10(1), 73–88. Salim, H.N. (2022) The causes of student’s reluctance to ask ques‐ tion’s when attending lectures. Teaching English and Language Learning English Journal (TELLE),.doi:10.36085/telle.v2i3.4719. Selwyn, N. (2019) Should robots replace teachers?: AI and the future of education. : John Wiley & Sons. Sentance, S. & Csizmadia, A. (2017) Computing in the curriculum: Challenges and strategies from a teacher’s perspective. Education and information technologies, 22(2), 469–495. Shneiderman, B. (2022) Human‐Centered AI. : Oxford University Press. Sithaldeen, R., Phetlhu, O., Kokolo, B. & August, L. (2022) Student sense of belonging and its impacts on help‐seeking behaviour. South African Journal of Higher Education, 36(6), 67–87. Swacha, J. et al. (2025) Retrieval‐augmented generation (rag) chat‐ bots for education: A survey of applications. Applied Sciences,. Tinto, V. (2012) Leaving college: Rethinking the causes and cures of student attrition. : University of Chicago press.
23
Tran Huu Van et al. (2026) A course‐specific agentic rag chatbot for it student support: Architecture, local deployment, and pre‐ liminary evaluation at hai phong university. In: Next‐Generation Computing Systems and Technologies. Wang, F. & Hannafin, M.J. (2005) Design‐based research and technology‐enhanced learning environments. Educational technol‐ ogy research and development, 53(4), 5–23. Yan, L., Sha, L., Zhao, L., Li, Y., Martínez‐Maldonado, R., Chen, G. et al. (2023) Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology,.doi:10.1111/bjet.13370. Yang, F. (2025) A strengths, weaknesses, opportunities, and threats (swot) analysis of chatgpt in academic help‐seeking. International Journal of Teaching and Learning in Higher Education, 36(1), 10.
How to cite this article: Gray A., Hobbs J., Reducing Barriers to Aca‐ demic Support: Evaluating a Course‐Specific RAG System for Address‐ ing Help‐Seeking Disparities in Higher Education. Preprint. Not yet peer reviewed. 2026;00(00):1–29.
APPENDIX
A EXPLORATORY PILOT STUDY QUESTION‐ NAIRE Understanding the System 1. How clearly did you understand how the RAG system works based on the demonstration?*
• • • • •
Very clearly Somewhat clearly Neutral Somewhat unclearly Not at all
2. Did the system’s responses make sense in relation to the questions asked?*
• • • • •
Always Most of the time Sometimes Rarely Never
Perceived Usefulness 1. How relevant were the system’s responses to the course materials?*
• Very relevant • Somewhat relevant • Neutral
24
• Somewhat irrelevant • Very irrelevant
Participant and Educational Background 1. ID Number
2. Do you think this system could help clarify assessment require‐ ments?*
• • • • •
2. Year of study 3. Course 4. How would you rate your current programming ability?
Yes, definitely
5. What grade did you achieve in CodeLab I?
Yes, somewhat Neutral
Existing Learning Resources and AI Use
Not really No, not at all
1. Which online tools do you use to aid your understanding of the 3. Would you trust the system’s responses when seeking academic support?*
• • • • •
programming concepts introduced in CodeLab? 2. Which module resources do you use to aid your understanding of CodeLab module content?
Yes, completely
3. How do you use AI tools within CodeLab?
Yes, somewhat
4. What types of questions or prompts do you ask when using AI tools?
Neutral
5. Which resources do you find most helpful in aiding your understand‐
Not really
ing of programming concepts?
No, not at all
6. Please provide a reason for your response to the previous question.
Impact on Help‐Seeking & Learning
Help‐Seeking and Learning Confidence Please indicate your agreement with the following statements using a
1. If this system were available to you, how likely would you be to use
five‐point Likert response format:
it before asking a lecturer?*
• • • • •
1 = Strongly Disagree, Very likely
2 = Disagree,
3 = Neutral,
4 = Agree,
5 = Strongly Agree
Somewhat likely Neutral
1. I feel comfortable asking lecturers for help.
Somewhat unlikely
2. I often hesitate to ask questions in class.
Very unlikely
3. I worry about being judged when asking for help. 4. I prefer to work through problems on my own before asking for help.
2. Do you think the system could reduce hesitation in seeking help for assessments?*
• • • • •
5. I sometimes avoid asking for help even when I need it. 6. I feel confident when studying difficult topics. 7. I feel anxious when I do not understand a topic.
Yes, significantly
8. I feel comfortable admitting when I do not understand something.
Yes, somewhat Neutral
Frequency of AI Use and Help‐Seeking Be‐ haviour
No impact It would make me less likely to seek help
B MAIN STUDY: QUESTIONNAIRE
STUDENT
PRE‐
Participants were asked to indicate how often they engaged in the following behaviours: 1. Use generative AI tools (e.g., ChatGPT or similar) when working on
Prior to using Beacon, participants completed the following question‐ naire to establish their educational background, existing use of learning resources and generative AI tools, and attitudes towards help‐seeking and learning programming.
programming tasks for this module. 2. Avoid asking a tutor for help, even when you feel unsure about something. 3. Use generative AI tools instead of asking a tutor for help.
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
C MAIN STUDY: TASK SCENARIOS
25
You did not feel comfortable asking them to explain again.
Task 1
#include <iostream> using namespace std;
Given the code below you may feel you should already understand what happens here. This is the kind of question students often avoid asking
void askUser() {
tutors because it feels basic. There is no expectation that you already
int choice;
know the answer.
cout << "Continue? (1 = Yes, 2 = No): "; cin >> choice;
int count; if (choice == 1) {
cout << count;
cout << "Continuing...\n"; } else if (choice == 2) { Use Beacon to explore what this code does and why it behaves this way.
cout << "Exiting...\n"; } else {
In your own words, write an explanation below of:
cout << "Invalid input. Try again.\n"; askUser();
• What output you would expect • Why that output may occur
} } int main() {
Task 2
askUser(); }
You encounter the following error message when compiling your code. error: expected ';' before 'return'
Compiler messages are a common source of stress and frustration. Use Beacon to help you understand:
• What this error message is trying to tell you • What kind of mistake might have caused it
Task 3 Your tutor asks you to explain a while loop to another student. You are still learning this yourself, and you are given the following vague
Use Beacon to help you:
• Understand what your tutor likely meant by their feedback • Explore why using a loop might be more appropriate than recursion in this case Provide:
• A rewritten version of askUser() or Pseudocode • A short written explanation of why this approach is clearer or safer
Task 5
explanation to improve:
You’ve been given the following task to complete by your tutor.
“A while loop runs while something is true.”
Playing with Strings
Use Beacon to help you develop a clearer, more helpful explanation of a while loop.
Task 4 Your tutor gave feedback on your code in class. They mentioned that your askUser() function was poorly designed due to a bad use of recur‐ sion and that a loop should be used instead. You did not fully understand what they meant or how to change the code.
Write a program that asks for a user’s first name and last name separately. The program should pass these strings to a function which returns the users full name as a single string. Next create another function that replaces every a, e, i , o, u with the letter z and returns the converted string Create a final function that reverses the user’s name and returns the re‐ versed string.
26
This task is intentionally larger than previous ones and may feel over‐ whelming at first. You may use Beacon to support your thinking and
Perceived Answer Quality and Trust
development and provide below:
Please indicate your agreement with the following statements:
• Your current solution or partial solution • A short summary of how you used Beacon to develop your solution
1. Beacon’s answers were clear and easy to understand. 2. Beacon’s answers were accurate and relevant to my query. 3. Beacon made good use of module‐specific materials. 4. I trusted the information provided by Beacon.
Task 6
5. I felt confident using the answers in my learning.
Use Beacon and ask it the kinds of questions you would normally ask another AI system (for example, ChatGPT or similar tools) when work‐ ing on programming tasks.
6. I understood where Beacon’s answers came from. 7. Seeing source material increased my trust in the responses. 8. I would still verify important information with lecturers or course materials. 9. I was concerned that the system might provide incorrect informa‐
Ask as many or as few questions as you like.
D MAIN STUDY: QUESTIONNAIRE
tion.
STUDENT
POST‐ Errors and Hallucinations 1. How often did you receive incorrect or unhelpful responses?
Participants completed the following questionnaire after using the Bea‐ con AI support system.
Response Scale Unless otherwise stated, the following items were measured using a five‐ point Likert scale:
• • • • •
Never Rarely Sometimes Often Very Often
2. If you encountered incorrect or confusing responses, please de‐ scribe an example.
1 = Strongly Disagree,
2 = Disagree,
3 = Neutral,
4 = Agree,
5 = Strongly Agree
Impact on Learning
Usability and User Experience
Please indicate your agreement with the following statements:
Please indicate your agreement with the following statements:
1. Using Beacon improved my understanding of programming con‐
1. Beacon was easy to use.
2. I feel more confident in solving programming problems after using
cepts. 2. The interface felt intuitive and straightforward. 3. I felt confident navigating and interacting with Beacon. 4. Beacon responded in a timely manner.
Beacon. 3. Beacon helped me learn independently without needing staff sup‐ port.
5. I had a positive experience using Beacon.
4. Using Beacon increased my confidence in tackling difficult topics.
6. I would like to use Beacon again for module support.
5. Beacon supported my learning rather than replacing it.
7. I would recommend Beacon to other students. Open‐ended question: Open‐ended questions: 6. In what ways did Beacon support (or not support) your learning? 8. What features would you like to see added or improved in future versions of Beacon? 9. Please provide any additional comments that would help improve
Accessibility and Psychological Safety
Beacon. Please indicate your agreement with the following statements:
Reducing Barriers to Academic Support: Evaluating a Course‐Specific RAG System for Addressing Help‐Seeking Disparities in Higher Education
1. I felt more comfortable asking Beacon than asking a tutor.
27
10. Would you use Beacon again in future modules? Why or why not?
2. Beacon made it easier to seek help privately. 3. Beacon reduced my anxiety when I was unsure about a topic. 4. Beacon made academic support feel more accessible. 5. Beacon provides useful support outside of normal staff availability. 6. Beacon could be especially helpful for students with low confidence.
F POST PILOT ACADEMIC SEMI‐ STRUCTURED INTERVIEW QUESTIONS
7. Beacon could help students who feel judged when asking questions.
1. What were your expectations of Beacon before you started using it?
8. Some students may rely on the system without thinking indepen‐
2. How would you describe your overall experience interacting with
dently.
Beacon? 3. Were there any moments where using Beacon felt particularly
Comparison with Alternative Support Sources Compared to other sources of support, please indicate your evaluation of Beacon:
4. Did you encounter any difficulties or frustrations while using Bea‐ con? 5. How would you assess the accuracy and usefulness of Beacon’s responses?
1. Compared to online resources (e.g., Stack Overflow, YouTube), Bea‐ con was: 2. Compared to generative AI tools (e.g., ChatGPT, Microsoft Copilot), Beacon was:
Response scale: |
Worse
|
About the Same
6. Do you see Beacon being useful for students? How / Why? 7. How did Beacon compare to other support sources you typically use? 8. Is there anything you would change about Beacon? 9. Would you use Beacon again in future modules? Why or why not?
3. Compared to asking a tutor, Beacon was:
Much Worse
smooth or helpful?
|
Better
|
Much Better Open‐ended question: 4. Please provide your reasoning for the responses given in the previ‐ ous questions.
E POST PILOT STUDENT SEMI‐ STRUCTURED INTERVIEW QUESTIONS 1. What were your expectations of Beacon before you started using it? 2. How would you describe your overall experience interacting with Beacon? 3. Were there any moments where using Beacon felt particularly smooth or helpful? 4. Did you encounter any difficulties or frustrations while using Bea‐ con? 5. How would you assess the accuracy and usefulness of Beacon’s responses? 6. Did using Beacon change how you approached your programming work or problem‐solving? 7. In what ways, if any, did Beacon support your learning during the pilot? 8. How did Beacon compare to other support sources you typically use? 9. Is there anything you would change about Beacon?