Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders Preprint, compiled May 21, 2026 Zoe Pfister
1∗, Clemens Sauerwein
1 , Benedikt Dornauer
1 , Tina Mersch
Michael Vierhauser
2 , Christian Wolf
2 , Ruth Breu1 , and
1
1 University of Innsbruck, Department of Computer Science, Technikerstraße 21a, 6020 Innsbruck, Austria 2 EKS InTec, Danziger Straße 3, 88250 Weingarten, Germany
arXiv:2605.21269v1 [cs.SE] 20 May 2026
Abstract The transition toward Industry 5.0 is reshaping industrial work environments with an emphasis on humancentricity, enabling close collaboration between humans and machines to enhance productivity and flexibility. However, such systems typically require monitoring of human workers and operators, often involving sensitive data, raising significant privacy concerns. As a result, affected workers and unions frequently reject humanmachine collaboration features due to a lack of transparency regarding privacy threats and implemented mitigation strategies. To enable early stakeholder involvement, establish trust, and support informed decision-making, privacy implications must be communicated in a way understandable to non-technical stakeholders. Yet, current Requirements Engineering (RE) practices provide limited methodological support for making privacy threats and mitigations accessible to non-technical stakeholders (e.g., individual workers or their representative unions). In this RE@Next paper, we propose a conceptual framework that guides software design from human monitoring-related use cases and requirements to informed decision-making guidance focusing on non-technical stakeholders. Building on principles such as Privacy by Design, the framework leverages Large Language Models (LLMs) to transform technical artifacts into accessible privacy reports. We share initial insights from two industry use cases, evaluate the quality of the generated reports, and outline future research directions toward integrating privacy transparency into RE processes for human-centric industrial systems.
1
Introduction
Germany [11]. These and many other examples highlight how privacy concerns can hinder the acceptance of monitoringThe transition toward Industry 5.0 emphasizes human- enabled systems and fuel opposition from workers and unions. centricity [1] and promotes close collaboration between humans and machines. In particular, these interaction points require From a requirements perspective, system requirements that instakeholder-centered Cyber-Physical Systems (CPS) engineering volve monitoring humans often face significant criticism and and operation [2]. The overarching objective is to enhance pro- resistance from worker unions, employee representatives, and ductivity while simultaneously increasing flexibility [3]. To en- privacy advocates. Practices are frequently regarded as intrusure correct and safe operation, as well as enable self-adaptation sive and ethically problematic, raising fundamental concerns during collaborative tasks, systems must implement monitors about employee privacy, trust, and dignity. Research in this to collect and analyse diverse runtime information [4, 5, 6]. In area indicates that workers and their representatives often view scenarios involving human collaboration, monitored data is not pervasive monitoring technologies with suspicion, particularly solely limited to pure machine-related data, but also includes due to concerns related to job quality, privacy violations, and information about human operators and workers [7, 8]. For reduced autonomy [12, 13]. In many cases, collective bodies and example, consider a mobile robot carrying heavy tooling while unions frequently argue that such systems should not be implefollowing a worker [9]. To operate safely, the robot must track mented without transparent safeguards, informed consent, clear both the operator and nearby humans to prevent collisions. We documentation of purpose, and effective mitigation strategies refer to requirements that involve the collection and processing to address privacy risks [14, 15]. In addition, compliance with of data about human stakeholder activities as human monitoring regulatory frameworks, such as the GDPR, introduces further constraints during system design. While several researchers requirements. have investigated how the GDPR and similar regulations can However, monitoring humans at runtime inevitably raises the be integrated and aligned with Requirements Engineering (RE) issue of collecting potentially sensitive information, posing processes or how privacy-related requirements can be extracted significant privacy concerns, potentially hampering, or even [16, 17, 18, 19], current RE practices still lack structured propreventing the acceptance and implementation of collaborative cesses that bridge privacy design artifacts and informed decisionfeatures altogether. In the past, several cases of extensive making by non-technical stakeholders. As a result, privacy workplace surveillance and misuse have resulted in pushback, threats, trade-offs, and mitigation strategies must be communipublic backlash, and regulatory scrutiny. For example, Amazon cated in a way accessible to non-technical stakeholders, including France Logistique was fined 32 million euros in 2023 for intrusive workers and their representatives, to increase transparency and employee monitoring that violated GDPR principles [10]. In support technology acceptance. a similar case in 2020, H&M was fined 35 million euros for monitoring several hundred employees at a service centre in In this RE@Next paper, we address this challenge by integrating privacy transparency considerations into Software EngineerAccepted for publication at the 34th IEEE International Requirements Engineering Conference (RE@Next’26). correspondence: [email protected]
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders
2
ing (SE) and RE processes. We develop a first iteration of a framework that guides software design from human monitoring requirements to informed decision-making guidance by workers and their representatives. Specifically, we employ established security and privacy analysis techniques, such as Data Flow Diagrams (DFD) and the STRIDE framework [20], and provide these artifacts as input to Large Language Models (LLMs). The agents transform the technical artifacts – together with the use cases and requirements – into accessible privacy reports. With this approach, we intend to increase involvement and trust of non-technical stakeholders.
how data moves through a system by modelling processes, data stores, external entities, data flows, and trust boundaries [22]. DFDs also form the basis for STRIDE-based threat modelling. STRIDE is a widely used approach for systematically identifying security threats during system design [23, 24]. It categorises threats into six classes: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, and Elevation of Privilege. STRIDE analyses each element of a DFD for potential threats within these categories. This structured approach helps developers identify security risks early in the design phase and derive appropriate mitigations before implementation. Since In this early research, we aim to pursue the following research STRIDE explicitly addresses information disclosure threats, which are particularly relevant in privacy-sensitive systems [23], objectives (ROs): DFDs combined with STRIDE provide a link between GDPRRO1: Develop a structured RE process that supports privacy oriented privacy analysis and structured threat identification. analysis and communication of privacy threats and miti- LLMs/Agents supporting RE Activities: LLMs are transgation tactics related to human monitoring requirements, former models [25] trained on large text corpora and designed thereby facilitating broader understanding and technology to generate plausible continuations of textual prompts [26]. One acceptance among stakeholders. key aspect is their ability to perform “in-context learning”, i.e., RO2: Investigate how LLM-assisted transformations of tech- adopting new tasks at runtime based on natural language innical privacy artifacts can improve the accessibility and structions (prompts) [26]. This flexibility has led to increasing understandability of privacy threats and mitigation strate- interest in using LLMs to support various activities in the SE lifecycle [27, 28, 29]. However, LLMs often act as black boxes, gies for non-technical stakeholders. lacking transparency (e.g., through hallucinations) and controlOur contributions in this paper are threefold: First, we present a lability, especially in complex tasks [30]. This limitation is structured process that links monitoring requirements, privacy particularly relevant when generating privacy-related artifacts design artifacts, and use cases to decision-making processes in for stakeholders, where each artifact must be rigorously validated worker unions. Second, we propose an approach to increase and, if necessary, edited by a human to ensure accuracy and technology acceptance in Industry 5.0 by making privacy threats correctness. To address this problem, Wu et al. [26] propose and their mitigations more accessible through LLM-assisted Chaining LLM steps together, where a complex problem is reporting. This includes generating human-readable reports decomposed into multiple sub-problems executed sequentially. that explain the privacy implications of human monitoring Each step receives inputs from previous steps of the chain. Naturequirements and planned mitigations. Third, we report on an rally, humans can analyse the intermediate outputs and edit them initial validation, gaining insights through a questionnaire and should they detect incorrect statements, such as misalignment. feedback session with our industry collaborator, and outline Another important factor influencing output quality is prompt directions for future research areas. formulation. For example, Brown et al. [31] have shown that few-shot prompting, where the prompt includes examples of the task to be completed, can significantly improve model performance compared to zero-shot prompts. A more recent approach is chain-of-thought prompting [32], where the input prompt contains chain-of-thought examples of the task to solve. This, for example, increased performance on arithmetic and complex symbolic reasoning tasks, among others. Additional work has also explored template-based prompting approaches for structuring Data availability: We provide all supplementary ma- LLM interactions in specific domains [33]. terial and additional exemplar descriptions as part To better understand what monitoring data is collected, and of a repository: https://github.com/Ethical-Human- how privacy may be affected, we enrich traditional use case Machine-Interaction/PrivacyArtifacts2Report. templates (e.g., goals, scenarios) [34] with monitoring aspects.
The remainder of this paper is laid out as follows: Section 2 presents background, existing research, and our motivating example. In Section 3, we introduce our framework and process steps. We apply the framework using the motivating example in Section 4. Section 5 presents a proof-of-concept validation using two industry use cases. Finally, Sections 6 and 7, outline our research roadmap, ongoing work, and conclusions.
The “monitoring use cases” describe what data is collected during monitoring, what equipment is used for collection, and who the stakeholders being monitored are. These monitoring Privacy in industrial CPS concerns how workers’ personal data use cases can then be further enriched with information for a is collected, processed, and shared. In monitoring-enabled GDPR Data Privacy Impact Assessment (DPIA) [21, 35]. environments, such data may include behavioural, positional, or performance-related information about workers. Regulatory Motivating Example: Our motivating example is drawn from frameworks such as the GDPR treat privacy as a design concern a use case defined by our industry collaborators at EKS-InTec by requiring data protection measures and transparent, accessible (InLine Control of Product Assembly). The goal is to improve product quality during assembly. To achieve this, the system information for data subjects [21]. uses camera-based tracking to monitor the assembly process and Security and Privacy Analysis Techniques: In Software Engi- identify deviations or faulty components. When such deviations neering, Data Flow Diagrams are commonly used to visualize
2
Background & Related Work
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders Technical View (Requirements)
§
System Requirements
Feedback Cycle
Agreement
Affected Stakeholders
3
Technical View (Security/Privacy Analysis)
identify relevant Requirements
Stakeholder Report
Monitoring Req. & Privacy-related HumanMonitoring Use Cases
model data flow A and data criticality
D
Combiner
Non-Technical View
EasyReq STRIDEHandler
Data Flow Diagrams
B
STRIDE Analysis
C
DFD2 Mermaid Explaination and Transformation Layer
Figure 1: Overview of our proposed framework, transforming human monitoring requirements and privacy analysis artifacts into stakeholderoriented privacy reports.
are detected, the system notifies shop-floor workers in real time, temporarily pausing the assembly process until the error is corrected. While the primary functionality of the use case revolves around quality assurance in the assembly process, it inherently necessitates continuous monitoring of workers. This introduces several privacy concerns related to the collection of video footage and behavioural information of shop-floor workers. Without systematic and rigorous privacy modelling and the translation of these technical specifications into transparent privacy reports, the system risks both non-compliance with GDPR and significant opposition from worker unions due to the sensitive nature of the captured video and behavioural data.
3
Towards a Privacy Transparency Framework for Human Monitoring Use Cases
Our framework (cf. Fig. 1) builds upon established software engineering artifacts. System engineers identify use cases and requirements that involve monitoring of human stakeholders. These A human monitoring requirements and use cases serve as the starting point for subsequent privacy analysis and stakeholder communication activities. In the following, we describe the framework and its core components. 3.1
Framework Overview
process of accepting or rejecting the proposed privacy-related requirements (e.g., shop-floor workers or worker unions). The goal is to facilitate informed stakeholder discussions and support consensus-building regarding the acceptability of the proposed solutions. Based on feedback, the process either results in formal agreement on privacy requirements or triggers an iterative feedback cycle in which the use case and requirements can be revised and adapted. 3.2
Explanation and Transformation Layer
The core component of the framework is the Explanation and Transformation Layer, which bridges the gap between technical artifacts and the perspectives of non-technical stakeholders. It receives the monitoring use case, its associated requirements, DFD, and the STRIDE analysis, and transforms them into a structured privacy report. This transformation is performed by a workflow of four specialized LLM agents: First, the DFD2Mermaid Agent converts a DFD image into a machine-readable Mermaid2 diagram and generates a textual summary describing its components, data flows, and trust boundaries. This serves two purposes: reducing the number of tokens required to process the graph, and providing a human-readable representation of the model, allowing a human reviewer to validate the LLM’s interpretation of the data flow and correct errors early in the process.
The conversion of technical artifacts into stakeholder-oriented explanations is handled by the EasyReq Agent and the STRIDE-Handler Agent. The latter receives the DFD Mermaid diagram and summary of the DFD2Mermaid Agent, along with the original requirement and the STRIDE analysis, as input. It iterates over the STRIDE entries, summarizing each identified threat in language comprehensible to non-technical stakeholders. The agent also enriches the summary with concrete privacy threat examples relevant to the given use case and explains how each identified threat is intended to be mitigated. In parallel, the EasyReq Agent rewrites the original use case and requirements into a simplified version suitable for non-technical stakeholders. In addition, it generates a short rationale explaining the purpose and intended benefit of each requirement and why The use case, associated requirements, Data Flow Diagram, and the requirement is important, as well as the motivation for its STRIDE analysis then serve as input to the LLM Agents within implementation (cf. Fig. 3). the Explanation and Transformation Layer. Within this layer, Both agents follow explicit guidelines for outputting informathe complex technical artifacts are transformed into a D privacy tion, as specified in their respective input prompts. They are report tailored to non-technical audiences. The resulting reports 2 See https://mermaid.ai. are presented to stakeholders involved in the decision-making
After identifying relevant monitoring use cases and requirements, the framework applies established security and privacy analysis techniques. Specifically, this involves constructing B a Data Flow Diagram and C performing a STRIDE analysis. These artifacts support the systematic identification of potential privacy risks affecting stakeholders and the development of planned mitigation strategies to address them. Together with the human monitoring requirements and use cases, these artifacts can also support the preparation of a GDPR-mandated Data Protection Impact Assessment (DPIA) [21]. For example, the artifacts align well with the DPIA Data Wheel [35], which maps questions about data processing activities to GDPR principles, facilitating regulatory compliance.
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders explicitly instructed to avoid technical terminology and instead explain the concept using simple terms, while keeping the output concise. Further, the EasyReq Agent is instructed to focus on the what and why of the use case rather than the how, while the STRIDE-Handler Agent illustrates threats through concrete examples and addresses privacy and surveillance-related concerns. Finally, the Combiner Agent combines the outputs of all preceding agents into a structured HTML privacy report.
You are a security engineer with expertise in translating technical security information for non-technical stakeholders. Your task is to take a technical software security requirement and its use case, then rewrite the requirement in clear, accessible language that non-technical stakeholders, especially worker unions, can understand. You should also explain why this security requirement is important and planned.
4
Your goal is to:
4
[input-documents]
Prototype Implementation and Usage Example
We implemented an initial research prototype of the Explanation and Transformation Layer workflow using n8n [36], a workflow automation platform for orchestrating agents. In future iterations, we plan to further extend and implement the workflow using LangChain3 to allow for more flexible agent coordination and integration with additional LLMs. The DFD2Mermaid Agent uses Gemini 2.5 Pro, which consistently produced valid conversions from DFD images into Mermaid representations during preliminary experiments, in contrast to other models. The remaining agents use Claude Sonnet 4.5. Extended reasoning (“Thinking Mode”) is only enabled for the STRIDE-Handler Agent, where more complex reasoning is required to interpret and summarize STRIDE analysis results. Prompts were initially designed manually using Claude Console following the RICE template (Role, Instructions, Context, Constraints, Examples) [28]. We then iteratively refined these prompts using Claude Sonnet 4.5 to conform to the Claude Prompting Best Practices [37, 38], e.g., structuring inputs via XML tags. Additionally, for both the EasyReq Agent and DFD2Mermaid Agent, we included a scratchpad section to the prompt, outlining a list of steps to follow during text generation (cf. Figure 2), facilitating the model’s chain-of-thought reasoning capabilities [32].
1. Rewrite the requirement in plain, non-technical language that avoids jargon, acronyms, and technical terminology 2. Explain the value and rationale for why this requirement is planned [guidelines for non-technical communication] Before writing your final output, use the scratchpad to: 1. Identify the core purpose of the technical requirement 2. Note any technical terms that need to be simplified or explained 3. Think about what business risks this requirement mitigates 4. Consider what stakeholders care about most
<scratchpad>[Your analysis here]</scratchpad> [examples]
Figure 2: Single-Shot Prompt Template with Chain-of-Thought Scratchpad used for Simplifying Requirements and Use Cases for Non-technical Stakeholders. Used by the EasyReq Agent.
threat categories for privacy, namely Information Disclosure, Spoofing, and Tampering.
Based on these artifacts, the workflow transforms the data for nontechnical stakeholders. The DFD is first converted to Mermaid syntax and summarized in natural language. The resulting output, together with the STRIDE analysis, the use case, and the As this implementation serves as an initial end-to-end proof of requirements, is then provided as input to the STRIDE-Handler concept, we intend to compare the current solution with other Agent , which generates structured threat explanations. An LLM providers, such as ChatGPT4 , or Le Chat5 as part of future example is shown on the bottom-right of Fig. 3, where the work (cf. Section 6). information disclosure threat is converted into structured prose: Using the Framework: We demonstrate our framework using the motivating example InLine Control of Product Assembly 1. Plain-language threat description (what): The threat is described in accessible language, supported by illustrative introduced in Section 2. The workflow begins when a system examples for better clarification. In the example, the agent exengineer uploads design artifacts: A the use case (including plains that unauthorized individuals could access raw camera goals, scenarios, and monitoring properties) and associated footage. requirements; B a Data Flow Diagram; and C the results of a STRIDE analysis. A partial example is shown on the 2. Explanation of stakeholder impact (why): The potential left side of Fig. 3. In this example, the DFD models how consequences are described from the perspective of the stakecamera data flows from the sensor to an Edge Device Data holders (e.g., shop-floor workers in the motivating example). Processor, which analyses the footage for assembly errors and The explanation highlights that unauthorized access to camera forwards detected violations to a Cloud Platform (annotated footage is a clear privacy violation and could be misused for as (1) in Fig. 3). Applying STRIDE to this data flow reveals surveillance beyond the intended purpose of quality control. a potential information disclosure threat, where unauthorized 3. Mitigation explanation (how): The proposed mitigation stratindividuals could gain access to raw camera data if the data egy is discussed. In the example, encryption is described flow is insufficiently protected. A possible mitigation in this as making the data unreadable to anyone without proper scenario is encryption. Since the final report focuses on privacy authorization. aspects, we scoped the STRIDE analysis to the most relevant Each STRIDE threat is transformed following this structure. 3 See https://www.langchain.com. 4 See https://chatgpt.com. 5 See https://chat.mistral.ai/chat.
Next, the EasyReq Agent simplifies the original use case and requirements, producing a natural-language version and a rationale
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders Generated Outputs
Technical Artifacts
Use case & Requirements
Req-1: When a worker completes each assembly step, the system shall monitor and verify that the action performed matches the predefined correct action for that step via camera-based tracking.
Data flow diagram & STRIDE Analysis Shop-Floor
Threat Type
Camera Camera Data
1
Alerts
Edge Device Data Processor
Shop-Floor Worker
2 Assembly Violations
Assembly Events
1 Information Disclosure
Read Write Audit
2 Information Disclosure
Audit Engine
Description
Mitigation Unauthorized All data in transit individuals could will be encrypted. gain access to raw camera data. A compromised data processor could leak sensitive worker data.
Data encryption at rest, data retention minimiziation.
Translation Layer
Name: InLine Control of Product Assembly Description The system can detect whether the worker made a mistake during assembly. Goal: Increase the quality of assembled products. Reduce faults. [...]
Cloud Platform
5
System Description: The company is building a system that acts as a quality assistant for workers on the assembly line. Using cameras, the system watches each step of the assembly process and compares it to the correct procedure. If a worker performs a step incorrectly, the system immediately shows a message [...] System Purpose: The primary goal of this system is to help workers produce higher quality products and catch mistakes immediately [...] Unauthorized Access to Raw Camera Footage What This Means: The cameras record raw video footage of the assembly area, which includes images of workers. Unauthorized people could potentially gain access to this footage—like someone obtaining security camera recordings without permission. Why Mitigation Is Important: Raw camera footage shows workers’ faces, movements, and activities throughout their shifts. Unauthorized access to this footage is a serious privacy violation and could be misused for surveillance or monitoring beyond the stated purpose of quality control. How Privacy Will Be Protected: All data leaving the camera will be encrypted, making it unreadable to anyone without proper authorization.
Figure 3: Example of Inputs and Generated Outputs of the Workflow based on the example (UC1) introduced in Section 2.
explaining the intended purpose/benefit of each requirement (cf. on established document quality criteria. Since we are not top-right side of Fig. 3). aware of existing frameworks to evaluate non-technical privacy Finally, the Combiner Agent integrates both outputs, generating reports, we base our validation criteria on those proposed by the final report D in HTML format. The report begins with an Krishna et al. [27], who evaluate LLM-generated Software executive summary, followed by a simplified system description Requirements Specifications (SRS). Specifically, we collected the following metrics: (1) internal consistency, (2) redundancy, (3) and purpose statement produced by the EasyReq Agent. Next, completeness with respect to technical artifacts, (4) conciseness, the report lists the privacy threat explanations, each containing (5) correctness, and (6) understandability on 5-point Likert the three-part structure described above. This ordering ensures scales. These criteria capture general aspects of document that a non-technical reader first understands what the system does and why it is being built before being introduced to potential quality and provide a suitable baseline for evaluating structured privacy reports. After completing the questionnaire, the expert privacy risks and mitigation strategies. participated in a semi-structured interview to provide qualitative feedback on report structure and understandability.
5
Preliminary Validation
The goals of our approach are to develop a structured RE process that supports privacy analysis and communication to non-technical stakeholders (RO1) and to automatically generate an easily understandable privacy report based on a set of (technical) requirements and relevant standards and guidelines (RO2). Through this preliminary validation, we aim to obtain initial insights into the quality of our generated reports and the feasibility of our proposed process and workflow. We, therefore, focus on two aspects: First, we demonstrate the feasibility of our proposed workflow using our end-to-end prototype applied to two real-world use cases (cf. Fig. 3) from our industry collaboration (cf. Section 5.2). Second, we assess the quality of the generated reports using a combination of Software Requirements Specification (SRS) metrics and qualitative expert feedback through a survey and semi-structured interview with a domain expert (cf. Section 5.3). The full set of interview questions and results can be found in the supplementary material. 5.1
Validation Setup
We were provided with two use cases by our industry collaborators: (UC1) InLine Control of Product Assembly, monitoring workers to detect assembly errors and sequence violations, and (UC2) Cycle Time Monitoring for Lean Production, tracking worker assembly durations to identify process inefficiencies. In both cases, requirements include camera-based monitoring of shop-floor workers, making them suitable candidates for privacy threat analysis and subsequent stakeholder-oriented reporting. To assess report quality, we designed a questionnaire based
The domain expert was provided with two reports: one manually created report – created by one researcher, and cross-checked by two additional researchers – and one automatically generated using our prototype implementation. The reports were presented in a blinded manner, i.e., the expert was not informed which report was generated by the framework. After answering the questions concerning the 6 quality metrics, the domain expert participated in a follow-up feedback session, where we conducted a semi-structured interview to gain more information on how to improve the report’s structure and understandability in future work. 5.2
End-To-End Report Generation
For each of the two use cases and their requirements, we manually created a Data Flow Diagram and conducted a STRIDE analysis, focusing on the privacy-relevant threat categories of Information Disclosure, Spoofing, and Tampering. Based on these artifacts, we generated reports using our workflow and manually created baseline reports for comparison. Repeated executions of our workflow revealed minor inconsistencies in the generated output, confirming the non-deterministic nature of LLM-based workflows. For example, section titles for privacy violations were sometimes derived directly from threat titles of the STRIDE analysis, which might be unclear to non-technical stakeholders. We also observed minor formatting variations (bullet points vs. prose) across several workflow runs. To reduce visual bias, both manual and generated reports were formatted uniformly before being shared with participants.
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders 5.3
Expert Feedback
6
Finally, the interview revealed that inconsistencies within the input artifacts themselves propagate into the generated reports. For example, one artifact proposed encryption as a mitigation against threats where the adversary potentially holds the decryption keys. Since the LLM cannot produce internally consistent output from contradictory inputs, such issues carry over into the final reports. This highlights a fundamental limitation of the approach: the quality of the generated privacy report is bound by the quality and consistency of the initial technical artifacts, making rigorous validation of input documents a prerequisite for meaningful report generation.
Quantitative Results: Initial feedback from our expert indicates that the generated reports matched or outperformed the manually created baseline reports across all evaluated quality dimensions. The largest differences were observed in the completeness, conciseness, and understandability measures for UC1, where the generated report scored two points higher (on a 5-point scale). The internal consistency varied between the individual UCs, with the generated report scoring a 3 for UC1 and 5 for UC2. Related to understandability, the expert perceived the generated reports as more accessible from a language standpoint compared to the manually written ones. However, the generated reports 6 Discussion & Research Roadmap introduced other issues, such as undefined abbreviations (e.g., LTPE for Lean Time Processing Engine), highlighting the need Our initial validation, preliminary results, and expert feedback for human validation of generated outputs. indicate that we are able to partially achieve our research obQualitative Results: The follow-up interview provided ad- jectives: our early version of the framework demonstrates how ditional insights into how the reports’ understandability and privacy analysis can be systematically integrated into early restructure could be further improved with respect to RO2. Cur- quirements engineering processes (RO1), and how structured, rently, each STRIDE threat is presented as a separate section in stakeholder-oriented privacy reports can be generated from techthe report. However, the expert pointed out that this structure can nical artifacts (RO2). Building on these results, we identify five result in perceived redundancy, making it difficult to understand research directions for our future work. that the sections actually describe different parts of the data Integration of diverse Privacy Analysis Frameworks: In flow. For example, multiple threats in UC1 used encryption as our initial prototype, we employed the STRIDE framework [20] their primary mitigation tactic (e.g., camera to data processor to for systematically collecting and documenting privacy-relevant cloud provider). Grouping such related threats and explaining threats and corresponding mitigations. However, several alternathem along the data flow could improve readability and reduce tive frameworks exist, and security and privacy researchers have redundancy, making the report more concise. The expert indi- developed various approaches for specifically analysing privacy cated that restructuring reports to reflect data flow dependencies risks, such as LINDDUN [39], the DPIA Data Wheel frameand shared mitigation strategies complemented by a (simplified) work [35], or STPA-Priv [40]. Given our modular architecture, visualization of the DFD (cf. Fig. 3), could further increase these can be easily integrated as new agents. As part of RO1 both conciseness and stakeholder comprehension. However, (cf. Section 1) – extending our structured process – we will analthey highlighted that there is a trade-off, and care must be taken, yse other assessment frameworks, investigate the strengths and ensuring that the report does not become too technical. These limitations, and evaluate their suitability in early requirements findings provide initial evidence for RO2, while also highlighting elicitation processes to improve technology acceptance. Furtherareas for improvement in report structuring and presentation. more, we plan to develop an ontology, relating concepts across The expert further suggested enhancing the report with a more frameworks (threats, assets, controls, privacy goals). From this, comprehensive introduction, describing the system in greater we aim to derive recommendations and guidelines for selecting detail, as well as a concluding section summarizing key insights, a framework based on project characteristics and the specific open issues, and areas requiring stakeholder discussion. Ad- nature of human monitoring involved. ditionally, introducing prioritization, e.g., using the MoSCoW Mapping to Regulatory Texts: Currently, input artifacts do principle (must, should, could), may serve as a basis for stake- not explicitly link identified threats and mitigations to relevant holder deliberation and provide additional feedback. regulatory provisions, such as the GDPR [21]. This lack of traceability limits the applicability of the generated reports in 5.4 Threats to Validity compliance verification (e.g., GDPR Data Protection Impact Our initial proof-of-concept demonstrates the feasibility and Assessments). To address this limitation, we plan to enrich successful application of our framework and the generated privacy threats and mitigation information with explicit mappings reports. However, while the preliminary validation is based on a to relevant regulatory texts, including, for example, specific limited number of use cases, they represent real-world scenarios articles and recitals of the GDPR. This could be achieved through from our industry collaborators. To address this limitation, we the integration of Retrieval Augmented Generation (RAG) [41] are currently collaborating with industry partners to extend the context pipelines supplying the agents with relevant regulatory set of use cases and apply the framework to more complex and passages when processing STRIDE artifacts, or through the knowledge sources via the Model Context diverse scenarios. In the initial prototype implementation and integration of external 6 preliminary validation, we primarily used Claude Sonnet 4.5 Protocol (MCP) . For example, mitigation tactics, such as data (due to availability) for our agents. As part of our future work, minimisation (cf. Fig. 3), could be linked to Article 5(c) of the we plan to evaluate our framework using diverse LLMs (cf. GDPR [21]. Such additional augmentations would enable the planned extended evaluation, Section 6). Despite this limitation, generation of specialized reports tailored to legal stakeholders Claude Sonnet 4.5 did provide promising results for the generated and support compliance verification, further extending RO2. reports. 6 See https://modelcontextprotocol.io.
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders
7
Agent-Augmented Support: Currently, our workflow relies on manually created design artifacts as input for report generation. In our ongoing work, we intend to extend the framework towards increased automation through LLM-supported artifact generation. This includes using custom agents that generate requirements from use cases, or conduct an automatic STRIDE analysis with tools such as STRIDE GPT7 . Another challenging topic for future work is the automated generation of DFDs from software requirements artifacts, where the LLM must make sound architectural decisions to satisfy all system requirements. While such automation has the potential to reduce modelling effort, it also shifts the burden towards validation. This is especially critical for privacy and security aspects of SE, where incorrect threat assessments or missing mitigations may lead to non-compliance with the law, making rigorous human-in-theloop validation essential. As part of this, we will investigate perceived accuracy and completeness of generated artifacts from the perspective of security professionals, as well as the associated cognitive load, for example, using established methods such as NASA-TLX.
7
Conclusion
Comprehensive evaluation: The preliminary validation results presented in this paper serve as initial evidence that LLMgenerated privacy reports from our framework accurately reflect relevant requirements and privacy analysis artifacts while improving accessibility to non-technical stakeholders. However, a more comprehensive evaluation is required to fully address RO1 and RO2. First, we plan to conduct an interview study with non-technical stakeholders and legal experts to better understand their expectations regarding privacy reports. These insights will then inform the second iteration of our privacy report generation process. Additionally, we plan to evaluate the resulting reports through a comparative evaluation, comparing LLM-generated and manually created reports, including both technical experts and non-technical stakeholders. Based on the feedback received from our expert interview, we will also explore automated validation methods (e.g., round-trip reconstruction of reports and technical artifacts, LLM-as-a-Judge [42]) to identify information loss or inconsistencies in generated reports and reduce manual validation efforts.
elo, “Efficient and scalable runtime monitoring for Cyber-Physical System,” IEEE Systems Journal, vol. 12, no. 2, pp. 1667–1678, 2016. [5] A. Stojadinović, N. Stojanović, and L. Stojanović, “Dynamic monitoring for improving worker safety at the workplace: use case from a manufacturing shop floor,” in Proc. of the 9th ACM Int’l Conf. on Distributed Event-Based Systems, pp. 205–216, 2015. [6] M. Vierhauser, J. Cleland-Huang, S. Bayley, T. Krismayer, R. Rabiser, and P. Grünbacher, “Monitoring CPS at runtime-a case study in the UAV domain,” in Proc. of the 44th Euromicro Conf. on Software Engineering and Advanced Applications, pp. 73–80, IEEE, 2018. [7] A. Buerkle, H. Matharu, A. Al-Yacoub, N. Lohse, T. Bamber, and P. Ferreira, “An adaptive human sensor framework for human– robot collaboration,” The International Journal of Advanced Manufacturing Technology, vol. 119, no. 1, pp. 1233–1248, 2022. [8] H. Liu and L. Wang, “An ar-based worker support system for human-robot collaboration,” Procedia Manufacturing, vol. 11, pp. 22–30, 2017. [9] Alex Greenberg – Siemens, “The future of bridging humans, robots, and humanoids with process simulate software.” https://blogs.sw.siemens.com/tecnomatix/thefuture-of-bridging-humans-robots-and-humanoidswith-process-simulate-software-video, 2025. [10] European Data Protection Board, “Employee monitoring: French sa fined amazon france logistique eur 32 mil-
In this paper, we presented a novel Requirements Engineering framework that integrates privacy threat analysis with automated, non-technical stakeholder-oriented report generation. By combining established security analysis techniques, such as STRIDE and Data Flow Diagrams, with a modular LLM-based transformation pipeline, the approach enables the systematic derivation of structured, non-technical privacy reports from technical Software Engineering and Requirements Engineering artifacts. Our preliminary validation demonstrates both the feasibility of the approach and its potential in future work. At the same time, the results highlight important challenges, including the consistency of input artifacts, the need for improved report structuring, and the importance of human validation in the loop. Overall, the findings suggest that LLM-supported transformations can serve as an effective mechanism to operationalize privacy-by-design principles within requirements engineering processes. As part of our research roadmap and ongoing work, we will extend the framework to incorporate additional privacy analysis methods, support artifact creation with LLMs, and conduct an empirical Persona-specific Reports: Stakeholders involved in the study to further assess the effectiveness and generalizability of decision-making process differ substantially in their technical our framework. expertise and informational needs. For example, worker representatives may primarily be interested in understanding what data References is collected and the privacy implications, whereas legal experts [1] J. Leng, W. Sha, B. Wang, P. Zheng, C. Zhuang, Q. Liu, T. Wuest, require detailed mappings to regulatory provisions to evaluate D. Mourtzis, and L. Wang, “Industry 5.0: Prospect and retrospect,” compliance. Currently, our framework generates a single report Journal of Manufacturing Systems, vol. 65, pp. 279–295, Oct. targeting non-technical stakeholders. Future work will introduce 2022. stakeholder persona-specific report generation to tailor reports [2] K. Michailidis, R. Strazdina, and M. Kirikova, “Continuous requirements engineering for digital transformation.,” in BIR to the specific concerns and expectations of different stakeholder Workshops, CEUR Proc., pp. 26–40, 2021. groups, supporting RO2. This persona-based customisation can be achieved by adapting LLM agent roles and prompt config- [3] T. Kopp, M. Baumgartner, and S. Kinkel, “Success factors for introducing industrial human-robot interaction in practice: An urations. A key research challenge, therefore, is identifying empirically driven framework,” The International Journal of which information is relevant for each persona, ensuring that the Advanced Manufacturing Technology, vol. 112, pp. 685–704, Jan. simplification of the content does not compromise completeness 2021. or correctness of the information (cf. Section 5). [4] X. Zheng, C. Julien, R. Podorozhny, F. Cassez, and T. Rakotoariv-
7 See https://github.com/mrwadams/stride-gpt.
Preprint – Transforming Privacy Artifacts into Accessible Reports for Non-Technical Stakeholders lion.” https://www.edpb.europa.eu/news/nationalnews/2024/employee-monitoring-french-sa-finedamazon-france-logistique-eu32-million_en, 2024. [11] European Data Protection Board, “Hamburg commissioner fines h&m 35.3 million euro for data protection violations in service centre.” https://www.edpb.europa.eu/news/nationalnews/2020/hamburg-commissioner-fines-hm-353million-euro-data-protection-violations_en, 2020. [12] T. Singh and A. Johnston, “How much is too much: employee monitoring, surveillance, and strain,” in Proc. of the 15th Int’l Conf. on Computational Intelligence and Security, 2019. [13] E. Oz, R. Glass, and R. Behling, “Electronic workplace monitoring: what employees think,” Omega, vol. 27, no. 2, pp. 167–177, 1999. [14] Eurofound, Employee monitoring and surveillance – The challenges of digitalisation. Publications Office of the European Union, 2020. [15] R. Connolly and C. McParland, “Dataveillance: Employee monitoring & information privacy concerns in the workplace,” Journal of Information Technology Research, vol. 5, no. 2, pp. 31–45, 2012. [16] O. Kosenkov, E. Zabardast, D. Fucci, D. Mendez, and M. Unterkalmsteiner, “Privacy by design: Aligning gdpr and software engineering specifications with a requirements engineering approach,” Information and Software Technology, p. 107946, 2025. [17] C. Negri-Ribalta, M. Lombard-Platet, and C. Salinesi, “Understanding the gdpr from a requirements engineering perspective—a systematic mapping study on regulatory data protection requirements,” Requirements Engineering, vol. 29, no. 4, pp. 523–549, 2024. [18] S. Abualhaija, M. Ceci, N. Sannier, D. Bianculli, S. Lannier, M. Siclari, O. Voordeckers, and S. Tosza, “Llm-assisted extraction of regulatory requirements: A case study on the gdpr,” in Proc. of the IEEE 33rd Int’l Requirements Engineering Conf., pp. 142–154, IEEE, 2025. [19] N. Manasreh, P. Spoletini, M. Valero, V. Nino, and I. SanchezCardona, “Designing age-friendly apps: Mining functional, usability, and privacy requirements from existing mobile applications,” in Proc. of the IEEE 33rd Int’l Requirements Engineering Conf. Workshops, pp. 605–611, IEEE, 2025. [20] L. Conklin, “Threat Modeling Process | OWASP Foundation.” https://owasp.org/www-community/Threat_Modeling_Process, 2025. [21] “Regulation (EU) 2016/679 of the European Parliament and of the Council on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation),” Apr. 2016. [22] L. Sion, K. Yskout, D. Van Landuyt, and W. Joosen, “Solutionaware data flow diagrams for security threat modeling,” in Proc. of the 33rd Annual ACM Symp. on Applied Computing, (New York, NY, USA), pp. 1425–1432, ACM, Apr. 2018. [23] S. Hernan, S. Lambert, T. Ostwald, and A. Shostack, “Threat modeling-uncover security design flaws using the stride approach,” MSDN Magazine-Louisville, pp. 68–75, 2006. [24] A. Shostack, Threat modeling: Designing for security. John Wiley & Sons, 2014. [25] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. ukasz Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems, vol. 30, Curran Associates, Inc., 2017. [26] T. Wu, M. Terry, and C. J. Cai, “AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts,” in Proc. of the 2022 CHI Conf. on Human Factors in Computing Systems, CHI ’22, (New York, NY, USA), pp. 1–22, ACM, Apr. 2022. [27] M. Krishna, B. Gaur, A. Verma, and P. Jalote, “Using LLMs in
8
Software Requirements Specifications: An Empirical Evaluation,” in Proc. of the 32nd Int’l Requirements Engineering Conf., pp. 475– 483, June 2024. [28] A. Vogelsang, “From Specifications to Prompts: On the Future of Generative Large Language Models in Requirements Engineering,” IEEE Software, vol. 41, pp. 9–13, Sept. 2024. [29] X. Hou, Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, and H. Wang, “Large Language Models for Software Engineering: A Systematic Literature Review,” ACM Trans. Softw. Eng. Methodol., vol. 33, pp. 220:1–220:79, Dec. 2024. [30] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu, “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” ACM Transactions on Information Systems, vol. 43, pp. 1–55, Mar. 2025. [31] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, and et al., “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems, vol. 33, pp. 1877–1901, Curran Associates, Inc., 2020. [32] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” Advances in Neural Information Processing Systems, vol. 35, pp. 24824–24837, Dec. 2022. [33] Y.-E. Tian, Y.-C. Tang, K.-D. Wang, A.-Z. Yen, and W.-C. Peng, “Template-Based Financial Report Generation in Agentic and Decomposed Information Retrieval,” in Proc. of the 48th Int’l ACM SIGIR Conf. on Research and Development in Information Retrieval, (New York, NY, USA), pp. 2706–2710, ACM, July 2025. [34] K. Pohl, Requirements Engineering. Springer Berlin Heidelberg, 2 ed., 2025. [35] J. Henriksen-Bulmer, S. Faily, and S. Jeary, “DPIA in Context: Applying DPIA to Assess Privacy Risks of Cyber Physical Systems,” Future Internet, vol. 12, p. 93, May 2020. [36] n8n, “N8n.io - AI workflow automation tool.” https://n8n.io/, 2026. [37] ANTHROPIC PBC, “Improve your prompts in the developer console.” https://claude.com/blog/prompt-improver, 2024. [38] ANTHROPIC PBC, “Prompting best practices.” https://platform.claude.com/docs/en/build-with-claude/promptengineering/claude-prompting-best-practices, 2026. [39] K. Wuyts, L. Sion, and W. Joosen, “Linddun go: A lightweight approach to privacy threat modeling,” in Proc. of the 2020 IEEE European Symp. on Security and Privacy Workshops, pp. 302–309, IEEE, 2020. [40] S. S. Shapiro, “Privacy Risk Analysis Based on System Control Structures: Adapting System-Theoretic Process Analysis for Privacy Engineering,” in 2016 IEEE Security and Privacy Workshops, pp. 17–24, May 2016. [41] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, “Retrieval-Augmented Generation for KnowledgeIntensive NLP Tasks,” in Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, Curran Associates, Inc., 2020. [42] D. Li, B. Jiang, L. Huang, A. Beigi, C. Zhao, Z. Tan, A. Bhattacharjee, Y. Jiang, C. Chen, T. Wu, et al., “From generation to judgment: Opportunities and challenges of llm-as-a-judge,” in Proc. of the 2025 Conf. on Empirical Methods in NLP, pp. 2757–2791, 2025.