ConceptioArchivearXiv CS
arXiv CSopen access

Operationalizing Software Engineering Theories for Practical Validation

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

Operationalizing Software Engineering Theories for Practical Validation Isaque Alves

Fabio Kon

University of São Paulo São Paulo, Brazil [email protected]

University of São Paulo São Paulo, Brazil [email protected]

arXiv:2605.03257v1 [cs.SE] 5 May 2026

Jessica Diaz

Carla Rocha

Universidad Politécnica de Madrid Madrid, Spain [email protected]

Abstract Context: Software Engineering often adapts theory-building frameworks from the social sciences to address socio-technical complexity. The key phases of the theory-building process are conceptual development, operationalization, testing, and application. Operationalization translates abstract concepts into measurable elements for empirical validation. This phase is essential for delivering the practical utility required by an applied science like Software Engineering. Objective: We propose a systematic procedure for the operationalization phase that bridges the gap between abstract concepts and empirical validation, ensuring the resulting theory is both rigorous and practically useful.Method: We extend the operationalization framework proposed by Sjøberg et al. and formulate non-causal hypotheses following Dubin’s approach. Our procedure defines variables, selects indicators, and systematically derives hypotheses. Results: We present a replicable, evidence-based methodological guideline that preserves a clear chain of evidence and supports practical validation. We illustrate the procedure using the DevOps Team Taxonomies Theory. Conclusion: This guideline provides a transparent chain of evidence from theory to testable elements, empowering researchers to ground theoretical advancements in empirical evidence and deliver actionable insights for practitioners.

CCS Concepts • Software and its engineering → Empirical software validation; Collaboration in software development; Software infrastructure; • General and reference → Empirical studies; • Social and professional topics → Socio-technical systems.

Keywords Software Engineering, Empirical Research, Continuous TheoryBuilding, Theory Operationalization ACM Reference Format: Isaque Alves, Fabio Kon, Jessica Diaz, and Carla Rocha. 2026. Operationalizing Software Engineering Theories for Practical Validation. In 3rd International Workshop on Methodological Issues with Empirical Studies in Software

This work is licensed under a Creative Commons Attribution 4.0 International License. WSESE ’26, Rio de Janeiro, Brazil © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2382-7/2026/04 https://doi.org/10.1145/3786149.3788307

University of Brasília Brasília, Brazil [email protected] Engineering (WSESE ’26), April 12–18, 2026, Rio de Janeiro, Brazil. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3786149.3788307

1

Introduction

Theory, as a “coherent description, explanation, and representation of observed phenomena” [7], is continuously shaped through the process of producing, confirming, and adapting. Software Engineering (SE), as a socio-technical discipline, requires theory-building methodologies that integrate human, organizational, and technical factors. Unlike social fields with established theory-building foundations [2, 5, 18], SE is relatively new in its systematic methodological guidelines to theory-building, and researchers often adapt methods from the social sciences. Sjøberg et al. [29], for example, expanded and tailored Lynham’s [18] theory-building method to address the socio-technical nature of SE. Sjøberg’s framework presents an iterative cycle for theory-building, encompassing: (1) Conceptual Development, which designs theories at a conceptual level, (2) Operationalization, which translates these concepts to the practical domain and enables, (3) Testing, and (4) Application [29]. In applied disciplines such as SE, theories gain greater relevance and impact when they can be effectively implemented and tested by practitioners [18]. To accomplish this, the operationalization phase enables researchers to apply and test abstract ideas in real-world scenarios by translating concepts into constructs and propositions into actionable and observable hypotheses [29]. For instance, consider the concept of usability in user interfaces. Though inherently abstract, it can be operationalized through measurable indicators such as average task completion time, error rate, and user satisfaction scores obtained via standardized instruments like the System Usability Scale (SUS). These metrics enable empirical evaluation of usability in real-world software systems. This study contributes to the ongoing research on the continuous development and refinement of software engineering theories. We extend the definitions provided by Lynham and Sjøberg et al. by integrating Dubin’s logic to present a systematic procedure for operationalizing theories. This approach bridges abstract concepts and empirical validation, ensuring that the resulting theory maintains academic rigor while delivering practical utility. Based on the Sjøberg framework, we outline a structured guideline for operationalizing SE theories, aiming to provide full traceability between the conceptual and practical elements, ensuring a clear and continuous link from abstract concepts to measurable

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

components. We focus on analyzing “what if” scenarios rather than causality or predictive generalization following Dubin’s non-causal modeling method [5] and Pérez et al. [23] in identifying interactions between concepts as the foundation for operationalization. Dubin emphasized that many typically used causal relationships are merely sequential interactions, noting, for example, that an alarm clock ringing before sunrise does not cause the sun to rise [12]. To illustrate and discuss the systematic procedure, we apply the operationalization process to the DevOps Team Taxonomies Theory (T3), developed in a collaboration between teams from the University of São Paulo and Universidad Politécnica de Madrid. We transform this theory from a conceptual model into clear and precise elements that enable researchers to test, apply, and potentially contradict it with evidence [24], represented as a set of 83 testable hypotheses to be confirmed or refuted at the practical level. This paper thus advances the understanding and application of the operationalization process in theory-building within SE by detailing its execution. We discuss how the hypotheses and propositions facilitate traceable empirical theory evaluation and provide a foundation for future studies on team dynamics in socio-technical systems within SE.

2

Continuous Theory-building

Lynham [18] defines theory-building as a structured method that involves both “theorizing to practice” and “practice to theorizing”. The process is “continuous” because findings from empirical testing and application consistently feed back into the conceptual framework, ensuring the theory is constantly refined and adapted to remain relevant over time. In this work, we adopted and refined the procedures proposed by Sjøberg’s methodological framework [29], which defines theory-building as a continuous process of refinement and adaptation, comprising five stages we depict in Figure 1: • 1. Conceptual Development generates theory by identifying and defining concepts and their relationships to describe the phenomenon. It relies on inductive and abductive processes to synthesize data and establish concepts and the relationship among them (propositions) to describe or explain a phenomena [29]. • 2. Operationalization translates conceptual theory into observable and measurable components. It transforms abstract concepts into constructs by assigning empirical values and converts theoretical propositions into testable hypotheses, preparing the theory for empirical evaluation. • 3. Testing involves empirically evaluating the theory by confirming or disconfirming its hypotheses and constructs. Researchers conduct empirical studies to assess the theory’s predictions and refine its elements based on the results [18, 29]. We commonly employ surveys and interviews to collect empirical data, confirming or refuting testable hypotheses. • 4. Application aims to observe the theory in practice. This phase assesses the theory’s relevance and effectiveness in addressing practical problems within specific contexts [18]. Researchers commonly use case studies and action research as primary techniques to investigate the application and practical implications of the theory in these contexts.

Alves et al.

Figure 1: Methodological framework for theory-building.

• 5. Continuous Refinement and Adaptation focus on refining and adapting the theory based on new evidence. Refinement involves clarifying or expanding existing concepts and relationships, while adaptation introduces new elements to ensure the theory evolves alongside advances in knowledge and practice. Theory Conceptual Development in SE is well-established and can follow a range of methodological guidelines such as: Systematic Literature Reviews (SLR), from secondary studies; Grounded Theory (GT) and case studies, which build theories directly from qualitative data; single-source primary studies, which rely on data from a specific case or organization; and even personal experience, where practitioners distill their knowledge into conceptual frameworks [26]. GT and case-study-based approaches have emerged as prominent methods for crafting robust conceptual frameworks grounded in empirical data. GT systematically derives taxonomies and process theories by closely aligning theoretical constructs with real-world observations, thereby minimizing bias [2, 3, 10]. Theories can take shape as taxonomies, classifications, process models, or ontologies, each serving distinct analytical and practical purposes. Despite offering solid conceptual foundations, software engineering theories frequently lack systematic procedures for operationalization and measurement in practice. The problem of lacking operationalization goes far beyond the inability to test and refine the theory. The main consequence is the theory’s incapacity to guide practice and adoption effectively. Without operationalization, theories remain conceptual, failing to provide actionable insights (such as the specific nuances of sociotechnical aspects found in DevOps), which leaves adoption vulnerable to simplistic or arbitrary approaches. To address the limitations, we outline a systematic operationalization procedure that can be applied regardless of the theory-building procedures adopted. We extend existing definitions provided by Lynham[18] and Sjøberg[29]

Operationalizing Software Engineering Theories for Practical Validation

to present a method for operationalizing theories, thereby ensuring their validity, traceability, and utility in practice.

3

Theory Operationalization Methodology

Theory operationalization addresses two main steps ([29], p. 327): (A) operationalizing theoretical concepts into empirical variables and (B) operationalizing theoretical propositions into empirically testable hypotheses. Figure 2 illustrates the detailed operationalization process we propose in this paper, adapted from Dubin [5], with steps highlighted in a dark blue background indicating adaptations or methods developed for this study. To ensure clarity in the process, we distinguish key theoretical components: constructs, concepts, empirical indicators, propositions, and hypotheses. When we operationalize a concept, it becomes a construct, as it acquires specific definitions regarding measurement through established variables, metrics, and indicators [28]. A hypothesis represents a prediction about the values of the units within a theory, where empirical indicators are used to measure the specified units in each proposition (the relationship among concepts) [5].

3.1

Operationalizing theoretical concepts into empirical variables

The first step of theory operationalization aims to identify concepts, empirical values from a conceptual theory through three steps [1]: (i) identifying the concepts that describe a phenomenon, (ii) selecting one or more variables to represent each concept, and (iii) choosing indicators for each variable. It is also necessary to define the range of values that each indicator may take over time to ensure reliability and consistency. The operationalization process for this study is deeply rooted in GT. We perform coding and constant comparison, which produces the conceptual relationships between codes and their properties as they emerge. Consolidating these emerging concepts and establishing the conceptual relationships between them into constructs and propositions. This is a meticulous phase where we perform an analysis of all identified concepts, exploring the connections and interdependencies among them. To identify the concepts, we examine the data, searching for descriptions of relationships. Concepts are typically expressed in relationships or when describing a phenomenon. The next step is to identify attributes that help explain the concepts. Some of these attributes and indicators are presented implicitly in the data and the provided descriptions. Finally, the indicators serve as a refinement that enables the instantiation of a structure or theory. As we describe different types of structures or frameworks, each with distinct characteristics, these characteristics become indicators of their attributes. These indicators enable the attributes to represent distinct and unique outcomes within the given context.

3.2

Operationalizing theoretical propositions into empirically testable hypotheses

The subsequent step consists of translating theoretical propositions into testable hypotheses that integrate the previously defined variables and indicators. According to Dubin [5], “A hypothesis is the

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

predictions about the values of the units of a theory in which empirical indicators are employed for the named units in each proposition.” After defining the variables and indicators, we identify and select the most strategic propositions, focusing on formulating simple and testable hypotheses (see Figure 2). We now detail each aspect of this operationalization, a structured four-step process: 3.2.1 Identifying the propositions. A proposition expresses the relationship between two or more concepts. By translating these concepts into variables and indicators in the previous step, we can derive propositions grounded in both data and theoretical reasoning. This allows us to investigate why particular structures emerge, their consequences, and the trade-offs and challenges they entail. Some relationships are explicitly stated in the data, such as the proposition presented by Leite et al. [13] (e.g., “API-Mediated Departments reduce bottlenecks”). Others were implicit, requiring inference based on descriptions of causes and consequences. Finally, to ensure traceability and reliability, we should provide quotations and explanations to justify the propositions, ensuring they remain well-grounded in the data. 3.2.2 Prioritizing Propositions for Testing. As the number of propositions increases, the number of potential hypotheses grows combinatorially. “The general rule is that a new hypothesis is established each time a different empirical indicator is employed for any of the units designated in a proposition” [5]. To avoid an explosion in the number of hypotheses, Dubin [5] recommends selecting strategic propositions – that establish boundaries/changes in the described system – to avoid an overwhelming number of hypotheses and follow the principle of parsimony (not multiplying entities unnecessarily). This step is crucial for the testing process, as the selected propositions serve as the input guiding the formulation of the test protocol. The testing questions will be derived from these propositions, and in the context of surveys, interviews, or questionnaires, they will directly reference them. The selected propositions describe state changes within the system, excluding those that do not contribute additional information about its current state. In practice, this involves selecting propositions that instantiate specific values of empirical indicators, thereby delineating the boundaries or transitions in the described system. The guiding principle is to ensure that the theoretical framework and team structure explicitly express the relationship between key concepts. 3.2.3 Formulating hypothesis. To formulate a simple and testable hypothesis, we focus on three aspects: (i) determining the nature of the semantic relationship between concepts, (ii) checking and simplifying the hypothesis complexity, and (iii) applying abductive reasoning to refine hypotheses. Our methodology proposes formulating non-causal hypotheses to align with a fundamental principle of scientific inquiry. As articulated by the philosopher David Hume [11], we cannot empirically observe a “necessary connection” that links a cause to its effect. Instead, our experience is limited to observing the “constant conjunction” of events. By framing our hypotheses around non-causal relationships, we provide a simpler description of the phenomena we can actually measure and test. This approach shifts the goal from

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

Alves et al.

Figure 2: Operationalization requires two steps: (A) defining concept variables and indicators, and (B) translating propositions into strategic, testable hypotheses. This results in refined hypotheses and an empirical testing protocol. proving causality to the more practical one of collecting evidence to demonstrate that a conjunction of events is regular and uniform. Following Dubin’s framework, we classify these non-causal relationship semantics into Categoric, Sequential, or Determinant interactions [5]. A Categoric Interaction relies on the presence or absence of concepts without implying change or progression. It corresponds to existence relationships, which establish whether one variable requires the presence or absence of another (e.g., “The presence of responsibility sharing is associated with the presence of collaboration”). A Sequential Interaction describes a progression over time, where one concept typically precedes another without implying direct causation. It aligns with gradient relationships, where the presence of one variable gradually modifies another’s value (e.g., “As responsibility sharing increases, collaboration tends to shift from occasional to daily”). Finally, a Determinant Interaction defines a structured dependency, where one concept’s value systematically changes alongside another’s. This corresponds to correlation relationships, which describe how two variables change together without establishing causation (e.g., “The frequency of collaboration is proportional to the level of responsibility sharing”). First, we classify relationship semantics into these non-causal types [5]. Second, we can reduce the number of hypotheses for parsimony by limiting them to two variables instead of introducing unnecessary complexity. For example, a complex hypothesis like IF (A OR B) THEN C involves three variables (A, B, and C) but can be split into two simpler hypotheses: IF (A) THEN C or IF (B) THEN C. However, formulating a complex hypothesis is necessary if C only occurs when both A and B are true simultaneously. In this work, we prioritize simple hypotheses to facilitate future testing, as they require verifying only two variables. In contrast, refuting a complex hypothesis like IF (A AND B) THEN C demands proving both A and B are true while showing C does not occur. Finally, we propose to apply abductive reasoning to refine hypotheses by generating plausible explanations that align with observed phenomena [31]. This process involves evaluating each hypothesis for its plausibility, internal consistency, and semantic alignment, ultimately enhancing the relevance of the hypotheses in realworld contexts [20]. Hypotheses deemed unlikely, dysfunctional, or logically inconsistent with the expected behavior described by the propositions are problematic and must be discussed or removed.

3.2.4 Selecting Essential Hypotheses for Testing. After establishing the hypotheses, we instantiate the constructs to identify the specific hypotheses that can represent and test the main scenarios described by the theory. Therefore, it is necessary to revisit the list of hypotheses and select those that describe propositions and a specific testing scenario. This step aims to refine the hypotheses to the representations of the theory and ensure that it aligns with the practical observations.

4

Applying the Methodology: DevOps Team Taxonomies Theory (T3)

This section presents our illustrative application, the DevOps Team Taxonomies Theory (T3) introduced by Diaz et al. [4], with which we outline an operationalization process. This paper is part of a broader research effort focused on the continuous building and refinement of the T3 theory through conceptual development, operationalization, testing, and refinement. The DevOps team taxonomy classifies team structures, providing a foundation for understanding current practices and planning DevOps adoption in the industry.

4.1

Introducing T3

T3 [4] harmonizes existing DevOps taxonomies (secondary data), and was produced by combining Scientific Papers [14, 16, 17, 19, 22, 27, 32] and Grey Literature [21, 25, 30] into a unified theory, undergoing a rigorous process of mitigating confirmation bias [26]. Diaz et al. [4] also proposed a systematic process for theory-building that integrates GT, Inter-Coded Agreement (ICA), and Sjøberg’s approach. T3 consists of a comprehensive theory of DevOps Team Taxonomies and follows a methodological guideline ensuring traceability and reproducibility with available data. More details on T3 are available in Section 6, Appendix A. The theory distinguishes four archetypal team-structure instantiations: • Bridge DevOps Team: helps and supports development and operation teams by deploying and hosting applications in the platforms they build, monitor, and support. The engineers of the Bridge DevOps teams are the DevOps practices facilitators; hence, they create, deploy, and manage the infrastructure (environments) and the deployment (CI/CD)

Operationalizing Software Engineering Theories for Practical Validation

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

Figure 3: T3 instantiation of an Enabler (Platform) Team. The team leverages technology to automate workflows and deliver platform services to a Product Team. The communication construct is annotated with captions identifying constructs, variables, and indicators. pipelines. They are usually the bridge interface between developers and IT Operations, driving the DevOps values and practices. • Enabler DevOps Team: Organizations create specific teams to satisfy product team necessities. It includes platform servicing and tools (mainly for infrastructure and deployment pipelines), consulting, training, evangelization, mentoring, and human resources. Thus, they behave as enabler teams by providing these capabilities. The Enabling DevOps Teams are named in different ways, e.g., DevOps Centers of Excellence, chapters, guilds, platform/SRE teams, among others. • Product Team: is entirely responsible for and has complete autonomy over a product or service, including scoping, managing, architecting, building, and operating it. This approach allows the product team to innovate quickly with a strong customer focus by aligning development and operations objectives with business goals. • Development and Operations Teams: have well-defined and differentiated roles with their departments and objectives. The development team, for example, focuses on implementing features and is led by a project manager. The interaction between teams typically occurs as a transfer of work. Developers deliver the finished code to the operations team, which creates and manages the infrastructure while also being responsible for the deployment. According to Gregor’s classification [9], T3 is an Analysis theory, providing descriptions and conceptualizations of “what is”. Furthermore, due to our epistemological positioning, and since the theory is constructed based on qualitative analysis of a set of data, it must necessarily be limited to a substantive (local) theory, as opposed to a formal (all-inclusive) theory [8].

4.2

Operationalizing T3

In this section, we apply the proposed operationalization methodology outlined in Figure 1 to the DevOps taxonomies in T3.

4.2.1 Operationalizing theoretical concepts into constructs. The operationalization process begins by translating a theory’s concepts into measurable constructs. This phase involves a structured analysis of the raw data and a traceability to the elements of operationalized theory. Here, as an example, we present Team construct, a foundational element of the DevOps theory. The Team concept is central to the theory; it organizes and defines relationships between different entities. Specializations, such as the Development Team, Operation Team, and Product Team, further justify Team as an abstract construct representing various organizational structures. Quotations such as “Collaboration and communications among team members can considerably increase by establishing cross-functional teams” help justify the concept of Team as a constructor. To define the Autonomy attribute, we search the data for information that qualifies and describes the constructors. We captured the Autonomy concept through the code “team self-organization & autonomy”. The indicator values (dependent and self-organization) directly link to specific team types found in the data. A “Product Team” is described as being “cohesive, small and multidisciplinary” with “a high level of sharing of the product ownership,” which is associated with “self-organization”. On the other hand, traditional teams have siloed structures where teams are often “dependent” due to “well-defined and differentiated roles” and have a “transfer of work” rather than shared responsibility. During the T3’s conceptual development, we apply the framework for describing SE theories proposed by Sjøberg [29], who already suggests executing Part A of the operationalization (see Figure2) as part of theory building. Therefore, T3 specifies the DevOps Team taxonomies using 10 constructs and 19 variables, depicted in a UML class diagram. Table 1 lists the team and collaboration constructors followed by their variables and indicators. Additionally, we bring 28 propositions (See Table 2) that establish relationships between variables and their indicators. All evidence of this process can be found in Section 6. Finally, Figure 3 illustrates the Enabler (Platform) Team instantiation of T3. The Team construct is one of the elements explaining the phenomenon of DevOps team structures. This construct includes

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

Alves et al.

several associated variables, such as autonomy, blame culture, and role definition. When instantiated, each variable is assigned a specific indicator. For example, in the case of Enabler (Platform) Teams in Figure 3, the Team construct is characterized by self-organization autonomy, no blame culture, and full sharing of skills, knowledge, stack, and tools. These teams also maintain high quality and engage in daily collaboration. Table 1: Example of Constructs, variables, and indicators derived from Team and Collaboration concepts (from [4]). Constructs

Team: It is an artificial (abstract) concept that represents a team structure of an IT department.

ID

Proposition

P1

A team culture based on responsibility/ownership sharing enables collaboration. ... Automated application life-cycle management is a platform servicing Automated infrastructure management is a platform servicing Enabler (platform) teams provide automated application life-cycle management

... P26 P27 P28

Variables

Indicators

autonomy

{dependent, organization}

blame

{true, false}

alignment of dev & ops goals

{local optimization, product thinking}

responsibility/ ownership sharing

{full sharing, medium sharing, minimal or null sharing}

self-

skills/knowledge {full sharing, medium sharing sharing, minimal or null sharing} stack & tools sharing

{full sharing, medium sharing, minimal or null sharing}

crossfunctionality/ skills

{true, false}

role definition/ attributions

{true, false}

Inherited mem- {product teams, horibers zontal teams, bridge teams, enabler teams, dev teams, ops teams} Collaboration: be- frequency tween teams. From the lack or even- quality tual collaboration, to daily collaboration.

Table 2: Examples of propositions provided by T3 (from [4]).

{daily, eventual} {high, low}

4.2.2 Operationalizing theoretical propositions into empirically testable hypotheses. The relationships between concepts will shape the hypotheses’ expressions in categoric, sequential, or determinant formats. As presented in Section 3, a new hypothesis is established each time a different empirical indicator is employed. Given a large number of propositions (28), variables (17), and indicators (33), we expect a high number of hypotheses. The indicators developed during the operationalization of theoretical concepts reflect more possible values than actual situations. To avoid an explosion in the number of hypotheses, Dubin [5] recommends selecting propositions that contain an instantiation

Table 3: Hypotheses Derived from Proposition P1. The hypotheses, highlighted in blue, represent the Enabler (Platform) Team structure.

P1 - categoric

Collaboration

frequency quality

daily eventual high low

Team responsibility/ownership sharing full medium minimal or sharing sharing null sharing h1.1 h1.2 h1.3 h1.4 h1.5 h1.6 h1.7 h1.8 h1.9 h1.10 h1.11 h1.12

of empirical indicator values establishing boundaries/changes in the described system. In our case, of the 28 propositions, only P26 and P27 are considered non-strategic (see Table 2). Automated Infrastructure Management and Automation Application Life Cycle Management are specializations of Platform Service. They are taxonomy elements but do not affect the system’s dynamics or the relationships between its concepts. We applied abductive reasoning to retain only the hypotheses most likely to describe real situations, favoring parsimony. For instance, the absence of responsibility/ownership sharing in a team implies that collaboration may be neither daily nor high quality. This process reduced the initial set of hypotheses from 115 to 83, and further refinement for specific contexts—such as the Enabler (Platform) Team—resulted in 30 hypotheses. P1. A team culture based on responsibility/ownership sharing enables collaboration establishes a categoric relationship between team and collaboration. The first concept, team, has nine associated variables (see Table 1), and we will focus in: responsibility/ownership sharing. This variable has three associated values: full sharing, medium sharing, and minimal or null sharing. The second concept, collaboration, has two associated variables, each with two values. Therefore, the number of possible hypotheses is 12 (Table 3). Nevertheless, the semantics of the relationship indicate that some form of responsibility/ownership sharing must exist (we discard h1.3, h1.6, h1.9, and h1.12) and that its existence implies an increase in collaboration (gradient relationship). In this case, the concept of collaboration has two variables, each with two possible values. Therefore, we formulate the simple testable hypotheses such as:

Operationalizing Software Engineering Theories for Practical Validation

Figure 4: The path from concepts to testable hypotheses (via constructs, variables, indicators, and propositions). P1 is used to generate the research questions, and the hypothesis table is color-coded by team structure.

H1.1 (h1.1 and h1.4): A team culture based on the full sharing of responsibilities makes it possible to move from eventual collaboration between team members to daily collaboration. Figure 4 summarizes the results of T3’s operationalization and illustrates how propositions and hypotheses are applied in the testing phase, making the link between operationalization and testing explicit. In this example, we test the theory with semi-structured interviews: propositions guide the interview protocol, expert responses are analyzed to confirm or refute the hypotheses, and full traceability identifies which theory elements require revision when a hypothesis is refuted.

5

Discussions and Implications

In this paper, we operationalize T3, enabling its continuous testing and refinement. The supplementary material provides the instantiation of four DevOps team structures, detailing their key concepts, variables, and indicators. In this section, we discuss the results, the role of operationalization, and outline the key implications of this work.

5.1

Practical Utility

Many organizations aspire to have high-performance DevOps teams; however, it is often difficult to accurately assess their current state or determine how to progress toward higher performance. The DevOps Research and Assessment (DORA) metrics provide a standardized framework to evaluate the effectiveness of DevOps practices,

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

focusing on four key indicators: Lead Time for Changes, Deployment Frequency, Change Failure Rate, and Time to Restore Service (MTTR) [6]. While these metrics effectively capture observed performance and maturity levels, they fall short of offering actionable guidance. Consider an organization with an Enabler (Platform) Team that supports product teams yet struggles to increase deployment frequency – a DORA throughput metric typically associated with shorter lead time for changes. DORA metrics show progress, but not how to achieve them or which practices to adopt. Through operationalization, this issue maps to the Collaboration and Automation constructs. One scenario analysis could reveal limited responsibility and ownership sharing across teams, hindering automation. Consequently, the organization can implement targeted Enabler-Team interventions – such as culture dissemination, consulting, and seamless tool integration – to strengthen Shared Responsibility. Greater shared ownership fosters more frequent, higher-quality collaboration, enabling deeper automation, higher deployment frequency, and ultimately, reduced lead time for changes. In this context, operationalization enables practitioners to (1) assess organizational maturity through defined attributes and variables, effectively conducting an organizational diagnosis; (2) outline a clear path for adoption or evolution by identifying which practices require change (e.g., increasing deployment frequency, reducing silos); and (3) establish indicators and metrics to monitor the effectiveness of these changes. Implication #1 Operationalization defines clear, measurable constructs that make a theory’s assessment criteria explicit and translate them into concrete, step-by-step guidelines for improving performance and advancing organizational maturity.

5.2

Traceability and Continuous Adaptation

Operationalization formally connects the conceptual to the practical and testable levels, defining constructs, variables, and hypotheses (see Figure 4), ensuring empirical traceability. This traceability underpins academic rigor by creating a clear chain of evidence and supporting robust, replicable results. It also enables researchers to adjust theoretical constructs based on empirical findings accurately. For example, when a hypothesis requires revision, traceability clarifies where and how these adjustments should be made. Subsequent movements – such as DevSecOps, MLOps, and AIOps – extend the DevOps paradigm by introducing additional quality dimensions and team dynamics into the delivery lifecycle, including security, data governance, and intelligent automation. From a theory-building perspective, these movements should not be regarded as isolated frameworks, but rather as conceptual evolutions within the DevOps taxonomy, inheriting its foundational constructs of team autonomy, continuous delivery, and cross-functional collaboration. Through operationalization, it becomes possible to derive new theoretical instantiations (e.g., MLOps or DevSecOps) as contextual refinements or “forks” of the core DevOps theory. This approach preserves theoretical continuity while enabling the systematic adaptation of constructs, variables, and indicators to domain-specific concerns, thereby ensuring cumulative knowledge development and empirical testability.

WSESE ’26, April 12–18, 2026, Rio de Janeiro, Brazil

Implication #2: Viewing DevSecOps, MLOps, and AIOps as evolutions of the DevOps taxonomy – rather than separate frameworks – supports cumulative theory-building in socio-technical systems. This approach allows researchers to extend existing constructs and indicators through systematic operationalization, maintaining theoretical continuity while adapting to new domains.

5.3

Operationalization for the Testing Protocol

The value of an abstract taxonomy versus an operational one becomes clear when designing a testing protocol. For example, Leite et al.[15] characterize siloed departments (i.e., Development and Operations Teams) through properties such as “Developers and operators have well-defined and differentiated roles.” While this descriptive taxonomy offers strong conceptual grounding, its abstract nature complicates the design of measurable tests. In a subsequent study, Leite et al. [13] employed open-ended interviews to validate their model, which increases the analytical rigor required to interpret unstructured data and relate it to abstract theoretical constructs. On the other hand, operationalization guides the testing protocol, enabling questions based on the propositions and validations through a non-causal hypothesis. For example, we can validate and discuss aspects of teams and collaboration (see Figure 4) without requiring a contextualization and explanation of the entire theoretical framework, thereby avoiding biases in data collection through surveys or interviews. The non-causal hypotheses enhance methodological rigor by moving the goal from proving absolute causation (e.g., Daily team collaboration reduces organizational silos) to the description of events (e.g., Teams with daily collaboration are associated with fewer organizational silos). However, daily collaboration is not the only factor responsible for reducing silos; other aspects, such as strong leadership support, standardized communication protocols, or shared performance metrics, may also contribute significantly to reducing silos. Implication #3: Operationalization serves as a methodological guide for the testing phase. Defining measurable constructs and non-causal hypotheses, it provides structure and direction to testing procedures, ensuring a transparent linkage between empirical findings and theoretical propositions. This systematic connection enables researchers to identify which aspects of a theory are empirically supported and which require revision, thereby strengthening the iterative process of continuous theory-building.

6

Data Availability

All data related to this research providing a detailed chain of evidence, is available at https://bit.ly/Operationalization.

Acknowledgment This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) - Finance Code 001, CNPq proc. 308327/2021-7, and FAPESP proc. 2023/00811-0.

References [1] Pritha Bhandari. 2022. Operationalization. A Guide with Examples, Pros & Cons. https://www.scribbr.com/methodology/operationalization/ [2] Kathy Charmaz. 2014. Constructing Grounded Theory. Sage 2nd Ed., London.

Alves et al.

[3] Juliet Corbin and Anselm Strauss. 1990. Basics of Qualitative Research: Grounded Theory Procedures and Techniques. SAGE Publication, London. [4] Jessica Díaz, Jorge Pérez, Isaque Alves, Fabio Kon, Leonardo Leite, Paulo Meirelles, and Carla Rocha. 2024. Harmonizing DevOps taxonomies—A grounded theory study. Journal of Systems and Software 208 (2024), 111908. [5] Robert Dubin. 1978. Theory building: a practical guide to the construction and testing of theoretical models. Free Press, New York. [6] Nicole Forsgren, Jez Humble, and Gene Kim. 2018. Accelerate: The science of lean software and devops: Building and scaling high performing technology organizations. IT Revolution, Portland. [7] Dennis A Gioia and Evelyn Pitre. 1990. Multiparadigm perspectives on theory building. Academy of management review 15, 4 (1990), 584–602. [8] Barney Glaser and Anselm Strauss. 1967. The Discovery of Grounded Theory: Strategies for Qualitative Research. Aldine de Gryter, New York. [9] Shirley Gregor. 2006. The nature of theory in information systems. MIS quarterly 30, 3 (2006), 611–642. [10] Rashina Hoda. 2022. Socio-Technical Grounded Theory for Software Engineering. IEEE Trans. Softw. Eng. 48, 10 (oct 2022), 3808–3832. doi:10.1109/TSE.2021.3106280 [11] David Hume. 2016. An enquiry concerning human understanding. Routledge, New York. 183–276 pages. [12] James Jaccard and Jacob Jacoby. 2019. Theory construction and model-building skills: A practical guide for social scientists. Guilford publications, Nova York. [13] Leonardo Leite, Nelson Lago, Claudia Melo, Fabio Kon, and Paulo Meirelles. 2022. A theory of organizational structures for development and infrastructure professionals. IEEE Transactions on Software Engineering 49, 4 (2022), 1898–1911. [14] Leonardo Leite, Gustavo Pinto, Fabio Kon, and Paulo Meirelles. 2021. The organization of software teams in the quest for continuous delivery: A grounded theory approach. Information and Software Technology 139 (2021), 106672. [15] Leonardo Leite, Gustavo Pinto, Fabio Kon, and Paulo Meirelles. 2021. The organization of software teams in the quest for continuous delivery: A grounded theory approach. Information and Software Technology 139 (2021), 106672. [16] Daniel López-Fernández, Jessica Díaz, Javier García, Jorge Pérez, and Ángel González-Prieto. 2021. DevOps team structures: Characterization and implications. IEEE transactions on software engineering 48, 10 (2021), 3716–3736. [17] Welder Pinheiro Luz, Gustavo Pinto, and Rodrigo Bonifácio. 2019. Adopting DevOps in the real world: A theory, a model, and a case study. Journal of Systems and Software 157 (2019), 110384. [18] Susan A. Lynham. 2002. The General Method of Theory-Building Research in Applied Disciplines. Advances in Developing Human Resources 4, 3 (2002), 221–241. [19] Ruth W. Macarthy and Julian M. Bass. 2020. An Empirical Taxonomy of DevOps in Practice. In 2020 46th Euromicro Conference on Software Engineering and Advanced Applications (SEAA). IEEE, New Jersey, 221–228. [20] Lorenzo Magnani. 2011. Abduction, reason and science: Processes of discovery and explanation. Springer Science & Business Media, New York. [21] Niall Richard Murphy, Liz Fong-Jones, Betsy Beyer, Todd Underwood, Laura Nolan, and Dave Rensin. 2018. Site Reliability Engineering book, Chapter 1 How SRE Relates to DevOps. https://sre.google/workbook/how-sre-relates/ [22] Kristian Nybom, Jens Smeds, and Ivan Porres. 2016. On the impact of mixing responsibilities between devs and ops. In International Conference on Agile Software Development. Springer, Springer International, Publishing, Cham, 131–143. [23] Jorge Pérez, Jessica Díaz, Ángel González-Prieto, and Sergio Gil-Borrás. 2024. Theory building for empirical software engineering in qualitative research: Operationalization. arXiv preprint arXiv:2412.02384 1-22 (2024), 1898–1911. [24] Karl Popper. 1959. The logic of scientific discovery. Routledge, New York. [25] Puppet and CircleCI. 2020. 2020 State of DevOps Report. https://www2.circleci. com/2020-state-of-devops-report.html [26] Paul Ralph. 2018. Toward methodological guidelines for process theories and taxonomies in software engineering. IEEE Transactions on Software Engineering 45, 7 (2018), 712–735. [27] Mojtaba Shahin, Mansooreh Zahedi, Muhammad Ali Babar, and Liming Zhu. 2017. Adopting continuous delivery and deployment: Impacts on team structures, collaboration and responsibilities. In Proceedings of the 21st International Conference on evaluation and assessment in software engineering. Association for Computing Machinery, New York, 384–393. [28] Dag IK Sjøberg and Gunnar Rye Bergersen. 2022. Construct validity in software engineering. IEEE Transactions on Software Engineering 49, 3 (2022), 1374–1396. [29] Dag I.K. Sjøberg, Tore Dybå, Bente Anda, and Jo Erskine Hannay. 2008. Building Theories in Software Engineering. Springer London, London, 312–336. [30] Matthew Skelton and Manuel Pais. 2019. Team topologies: organizing business and technology teams for fast flow. It Revolution, Portland. [31] Douglas Walton. 2014. Abductive reasoning. University of Alabama, Alabama. [32] Xin Zhou, Huang Huang, He Zhang, Xin Huang, Dong Shao, and Chenxin Zhong. 2022. A Cross-Company Ethnographic Study on Software Teams for DevOps and Microservices: Organization, Benefits, and Issues. In 2022 IEEE/ACM 44th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, Association for Computing Machinery, New York, 1–10.

Record · ID 155371 · SHA-256 4a6c380980e47bf1
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.