Article
Towards an Ontology for the Foundations of Software Languages Draft as of May 19, 2026. Submitted for review/publication. Ralf Lämmel 1 1
University of Koblenz, Faculty of Comuter Science, SoftLang Team; [email protected]
arXiv:2605.17374v1 [cs.SE] 17 May 2026
Abstract The notion of software languages subsumes programming languages, modeling languages, and yet many other types of languages used in software engineering. The emerging ontology ‘Foundations of Software Languages’ (FSL) organizes the foundations underlying software languages. We are concerned with language categories, language concepts, associated tools and methodological approaches, the formal systems or other formal entities underlying software languages, and the embedding of software languages into into software engineering activities. The primary objective of FSL is to serve as a knowledge resource in Computer Science education by connecting several subject areas in a principled manner. The first release of FSL (V1), as discussed in this paper, was built through a relatively standard methodology involving common steps for expectations, reuse, conceptualization, formalization, and validation. We leveraged GenAI to support ontology engineering (discovery, classification, linkage, completion, and transformation). Keywords: Software language; Ontology; Programming language; Software engineering; Megamodeling; Technological space; Generative AI.
1. Introduction This work presents a significant effort towards building an ontology for the foundations of software languages — FSL for short. This emerging ontology is a knowledge resource that is meant to support Computer Science education across fields such as programming language theory, software engineering, and software language engineering. FSL and accompanying artifacts regarding validation etc. are available online.1 The present paper covers the first release of FSL (V1) and fully presents its development. In the remainder of the introduction, we discuss the broader context of FSL, i.e., software language engineering (SLE), we motivate the need for an ontology for the foundations of software languages (FSL), we demonstrate the ontology at work (‘in a nutshell’), we summarize contributions and methodology (based on ontology engineering), and we conclude with the roadmap of the paper. 1.1. Background — Software Language Engineering The notion of ‘software language’ [1,2] integrates more specialized views on different types of ‘computer languages’, as they have developed in different fields of computer science, for example, i) programming languages and domain-specific languages in the context of programming language theory and language implementation including compiler construction; ii) modeling languages in the context model-driven engineering; iii) query languages in the context of databases or semantic data. Basically, every artificial language used in the software lifecycle and in software engineering is a software language. The 1
https://github.com/softlang/fsl
2 of 40
encompassing notion of ‘software language engineering’ assumes that software languages themselves require engineering, including lifecycle support, and that one needs to consider them often in the context of different technological spaces [3]. The SLE community, as perhaps best identifiable through its conference series,2 over the last 20 years, has made countless contributions on top of the aforementioned ‘assumption’. Some important developments concern language workbenches, safer metaprogramming approaches, bridges between technological spaces, support for advanced (e.g., bidirectional) transformation scenarios.
Figure 1. Selected software language categories in FSL
1.2. In Need of an Ontology for the Foundations of Software Languages As a field matures and grows, it is natural to see the need for robust knowledge resources. SLE is no different in this respect, and such resources are perhaps even more important for SLE, as it is an interdisciplinary and knowledge-integrating field. In some 2
https://dblp.org/db/conf/sle/
3 of 40
of the sub-areas of SLE, taxonomy- or ontology-like contributions have been developed. For example, programming language classification [4,5] is an established effort. However, there have been no successful efforts to provide a relatively comprehensive ontological foundation for SLE. We will discuss all available ‘pre-ontological’ material in some detail, when we discuss the aspects of ontology reuse in Sec. 2.2. A few highlights are given here. A noteworthy (and somewhat ongoing) effort has been SLEBOK (The SLE Body of Knowledge) which should feature “artifacts, definitions, methods, techniques, best practices, open challenges, case studies, teaching material, and other components that will afterwards help students, researchers, teachers, and practitioners to learn from, to better leverage, to better contribute to, and to better disseminate the intellectual contributions and practical tools and techniques coming from the SLE field” [6]. Some candidate pieces of such an SLEBOK have been published [7– 10]. By contrast, FSL focuses on formal knowledge representation as opposed to BOK-style recipes or best practices or experience reports. A quasi-ontological approach with relevance for SLE is megamodeling [11–13] where software systems are modeled in terms of types of artifacts, technologies, and languages involved, subject to rich relationships including conformance, set-membership, part-whole, and structured correspondence. Such work has not so far led to a comprehensive ontology for the foundations of software languages. The present paper and its FSL go well beyond specialized taxonomies and megamodels. 1.3. The FSL Ontology in a Nutshell Let us introduce FSL by a series of illustrative query scenarios and associated result graphs, thereby essentially exercising competency questions, as is standard practice in ontology engineering [14].
Figure 2. Selected software language concepts in FSL
In our first scenario, we expect to be able to query different classes (types) of software languages with the leaves of the shown classification hierarchy corresponding to actual lan-
4 of 40
guages. Fig. 1 shows the corresponding query result while including only some categories and individuals for the sake of presentability. In our second scenario, we expect to be able to query languages in a manner that we can retrieve the language concepts for the languages. A simplified expectation would be here to query languages by ‘paradigm’. Fig. 2 shows a more refined approach in FSL — while paradigmatic concepts appear (see imperative and OO paradigm as well as declarative and logic paradigm), there are many more refined language concepts. For example, there are concepts related to abstraction, control, and typing. In the figure, we focus on Java as a programming language versus OCL as a non-programming language. We can observe what concepts are specific to either language versus what concepts are shared. Only relatively abstract concepts are shared; see, for example, ‘paradigm concept’. Fig. 3 is less about a competency question or an expectation. This figure simply summarizes the top-level class hierarchy of FSL. The language taxonomy is intentionally left out, as we saw already part of it in Fig. 1. The tool taxonomy is left out for now; we will encounter some types of tools below. The most general subclasses of ‘Conceptual entity’ and ‘Formal entity’ are included. Let us explain a few nodes. ‘Formal entity’ is root to (Hilbert-style [15,16]) formal systems (such as lambda and process calculi, formal grammars, etc.), methodological approaches (such as metaprogramming, abstract syntax definition, and denotational semantics), and yet other formally defined entities. ‘Conceptual entity’ is root to different types of semantic annotations including language concepts (exercised in Fig. 2) and technological spaces explored separately below. In our next, the third scenario, we expect to be able to observe how formal entities (such as formal systems and methodological approaches) engage in relationships among themselves and with languages and tools; see Fig. 4. For instance, we can observe that denotational semantics uses lambda calculus and abstract syntax definition; also, model transformation is supported by the languages ATL and QVT and by the tool category of model transformation engines. In our next, the fourth scenario, we look into software engineering activities, while focusing here on implementation-related activities (as opposed to, for example, requirements, design, and deployment) and we expect to observe what kinds of languages, tools, and artifacts are relevant here; see Fig. 5. For example, code generation may be supported by a model transformation language and a model transformation engine. In our last, the fifth scenario, we expect to understand the typical ingredients of a technological space. That is, we query for all (types of) languages, tools, artifacts, formal, and conceptual entities as they are associated with the space; we also expect to observe the relationships between all these associated entities. Fig. 6 shows the resulting model for ‘Modelware’ [17]. 1.4. Contributions and Methodology We provide the first relatively comprehensive effort to organize the field of software languages including their engineering in an emerging ontology derived by a relatively formal and standard ontology engineering process. All previous efforts (see Sec. 2.2.2 specifically) are either focused on rather specific aspects (e.g., programming language classification) or they do not leverage ontology engineering (like in our own work on 101companies [18]). Our methodology is relatively standardized in terms of executing ontology engineering with the following specifics. Obviously, the methodology is specialized towards software languages and their engineering. In this context, we invoke crucial notions such as technological spaces, language concepts, and software engineering activities. Much in align-
5 of 40
Figure 3. FSL entity types — top-level view
ment with current trends in ontology engineering, we leverage GenAI to semi-automatically perform some steps of ontology engineering. We provide explained traces of GenAI usage. 1.5. Roadmap of the paper Sec. 2 develops the ‘Materials and Methods’ of our research. At its heart, the section outlines nine phases of bringing FSL into existence. There is also a part on the use of GenAI. Sec. 3 presents quantitative and qualitative results based on the execution of the phases of our methodology. Sec. 4 summarizes our findings and presents future work. Sec. 5 concludes the paper.
6 of 40
Figure 4. Selected formal entities in FSL and their associations among themselves and with languages and tools
2. Materials and Methods Sec. 2.1 reviews the method of ontology engineering and explains how it is applied in the methodology for FSL. Sec. 2.2 focuses on an important aspect of ontology engineering, i.e., reuse, to clarify what the knowledge resources underlying FSL are and to what extent existing foundational ontologies have been leveraged for FSL. Importantly, Sec. 2.3 describes
7 of 40
Figure 5. Selected SE activities in FSL(all related to implemented) and their associations with languages, tools, and artifacts
the phases of developing FSL in terms of the (non-) goals for each stage. Finally, Sec. 2.4 explains the use of GenAI in this research and publication effort. 2.1. Ontology Engineering FSL — the key result of the paper — has been created by means of an ontology engineering methodology. Accordingly, we rehash the field of ontology engineering below. We begin by recalling the basics in terms of scholarly work that paved the way. Afterwards, we summarize the general process for ontology engineering. As our ontology engineering effort employs GenAI in an essential manner, we also discuss related work on this aspect. Finally, as we are particularly interested in validating FSL for its first release, we also discuss related work on this aspect. 2.1.1. Ontology Engineering — Basics Ontology engineering is a mature field, perhaps starting in the early 1990ies with the notion of portable, reusable ontology specifications [19], with the clarification of foundational principles, methods, and applications [20], with moving methodology from ‘ontological art’ to engineering [21], also leading eventually (towards the end of the 1990ies) to a state-of-the-art in ontology engineering [22,23] and to the availability of a clear guideline for building ontologies [24]. It is instructive to observe the characteristics of more recent ontology engineering work. NeOn with its scenario-based, reuse-heavy methodology is quite established [25]. Challenges regarding ontology languages (e.g., the path from SKOS to OWL) and effective reuse (for sizable ontologies) are well pronounced [26]. The (necessary) alignment of the industrial lifecycle aligned with Semantic Web practice (LOT) [27] has received much interest; the need for an agile and collaborative methodology for organizations as well [28]. Also, the combination of ontologies and knowledge graphs has been a recent concern [29]. LOT [27] and NeOn [25] can serve as the backbone of any current methodology (like ours).
Figure 6. The technological space of Modelware in FSL 8 of 40
9 of 40
2.1.2. Ontology Engineering — the Recommended Process A sensible current process, distilled from [21,23–28], is: Expectation Specify purpose, scope, users, and competency questions. Reuse Survey and reuse existing ontologies, patterns, vocabularies, and data models first. Conceptualization Conceptualize with domain experts, iteratively and collaboratively. Formalization Formalize (‘encode’) in OWL et al. and document decisions. Validation Validate continuously: logical consistency, competency-question coverage, data-shape validation, and quality checks. Release Publish, version, and maintain the ontology as part of a broader KG/data lifecycle, also attracting and incorporating feedback from the relevant community. We follow this process in a refined manner; we also acknowledge that FSL is a very early stage of development. The expectation has been set in the introduction (specifically Sec. 1.3) with a number of scenarios (competency questions). Reuse is considered below (Sec. 2.2). Conceptualization, formalization, and validation are covered by our phases presented below (Sec. 2.3). Domain experts, up to this point, were the author and GenAI. FSL has been released on GitHub; a more advanced release process for new versions is beyond the scope of this paper, but see the future work discussion on CI/CD (Sec. 4.2.2). 2.1.3. Ontology Engineering — Use of AI (NLP, GenAI) Such use is, of course, established by now. (We focus on GenAI usage in what follows.) For example, ontology learning with GenAI is considered in a various places [30–35]; more specifically, a survey of LLM tasks in ontology engineering has been compiled [32] and the problem LLM-assisted ontology engineering from datasets has been researched [33]; see also the systematic literature review on LLMs for OE [34]. Our work is a response to the programmatic call to accelerate work on knowledge graphs and ontology engineering with LLMs [31]. Our methodology is ‘hybrid’ in how human–GenAI collaborative ontology design with expert validation is performed [35]; in our approach, the initiative remains fully with the human author; GenAI helps with productivity. 2.1.4. Ontology Engineering — Validation An overall assumption in validation is that structural, functional, and usabilityprofiling measures need to be ‘tested’ [36]. There is also the notion of ontology quality assessment [37]. There is actually tooling explicitly for ontology evaluation or validation [38] such that the many aspects are modularly covered. Validation may also rely on systematic use of constraints [39] (here: a combination of SHACL + SPARQL + Wikidata constraints) and there is actually a tricky interplay between validation and inference [40] (in the context of temporal ontology constraints). Of course, LLMs are also used for validation, for example, by means of translating textual requirements into executable validation queries [41]. Validation in terms of effective use of Wikipedia (by means of ‘browsing’ also) [42] is also heavily exercised in our methodology. 2.2. Reuse-related aspects Reuse of knowledge resources (preferably, ontologies) is a key requirement in ontology engineering, which we want to address here upfront. We look at two directions for possible reuse: a) upper ontologies that may provide a foundation for the more domain-specific ontology FSL we are aiming at; b) any knowledge structures in the immediate field of software languages that could suggest entities and relationships for FSL.
10 of 40
2.2.1. Upper ontologies with reuse potential FSL reuses some standard ontologies, for example, FOAF3 and SKOS4 , but the interesting question is whether FSL could reuse, in a meaningful manner, an upper ontology. That question is a common one in ontology development [43]. Let us discuss common options. The Basic Formal Ontology (BFO)5 is an upper ontology often reused across sciences; it focuses on space and time. Currently, space is not relevant for FSL. Time would be relevant for FSL, if we were to cover, for example, processes properly, for example, in software development. The General Formal Ontology (GFO)6 is somewhat similar to BFO, but it attempts to include several aspects of modern philosophy. GFO includes functions (in the teleological sense, i.e., relating to purpose); they could be useful in coordinating purpose-related dimensions in FSL. The upper ontology of Business Objects Reference Ontology (BORO) [44] — not available in OWL — can be used, for example, to model legacy software systems in the context of software re-engineering. FSL does not go into any depth of software architecture. The CIDOC Conceptual Reference Model7 is an ontology for the cultural heritage domain with a subset, CRM Core, which is an upper ontology for space-time, material and immaterial things, whereas FSL is mostly about immaterial things. The COmmon Semantic MOdel (COSMO) ontology8,9 is an upper ontology for semantic interoperability (i.e., an ontology to enable software systems to exchange data with unambiguous, shared meaning). Because of its focus on contemporary English (as opposed to the highly specific nature of software languages and their foundations) and also because of its huge size, it is not straightforward to extract a fragment with relevance for FSL from it. The Descriptive Ontology for Linguistic and Cognitive Engineering (DOLCE)10 (also its OWL version DOLCE-Ultralite11 ) is an upper ontology with classes for ‘objects’ (any physical, social, or mental object to participate in processes and to be spatiotemporally located), ‘events’, ‘processes’, ‘qualities’ (such as shape), ‘abstracts’ (i.e., non-spatiotemporal entities) with relationships for participation (of objects in processes), inherence (for qualities to inhere in entities), parthood (i.e., mereological relations), and dependence (i.e., qualities depending on their bearers). Except for some aspects, such as mereology, we cannot easily judge whether the overall class structure of DOLCE provides a helpful foundation for FSL. However, DOLCE-Ultralite has paved the way for Ontology Design Patterns12,13.14 [45], which are of use for FSL, for example, Classification, ConceptTerms, Constituency, Description, Information realization, Partition, PartOf, Tagging, Topic. We use such patterns without any formal import. As a descendant of the BFO and DOLCE line of work, there is now also gUFO: a gentle foundational ontology for semantic web knowledge graphs [46]. We may align FSL with gUFO at some point. 2.2.2. Pre-ontological knowledge in the software languages field Textbooks on programming languages and paradigms (such as [47–61]) are rich knowledge resources along multiple dimensions such as syntax, static semantics (well3 4 5 6 7 8 9 10 11 12 13 14
https://xmlns.com/foaf/spec/ https://www.w3.org/2004/02/skos/ https://github.com/bfo-ontology/BFO/wiki https://www.onto-med.de/ontologies/gfo https://cidoc-crm.org/ https://micra.com/COSMO/ https://github.com/mauna-ai/COSMO.owl https://www.loa.istc.cnr.it/index.php/dolce/ http://www.ontologydesignpatterns.org/ont/dul/DUL.owl http://ontologydesignpatterns.org/index.php/Ontology_Design_Patterns_._org_(ODP) https://dblp.org/db/conf/wop/index.html https://github.com/odpa
11 of 40
formedness and typing), dynamic semantics (e.g., based on operational semantics), pragmatics, paradigms, language concepts and mechanisms. We make sure that FSL covers such dimensions. We do not aim at the most detailed coverage of these programming language-related dimensions, but we instead aim to generalize across other types of software languages. Few of these textbooks are in a conceptualized format. One could call out Mosses’ action semantics [48], as it features a highly structured system of (tiny) executable language definition modulets. Research on domain-specific languages (DSLs) (or more specifically, domain-specific modeling languages (DSMLs)) and the corresponding research areas of generative programming, and software language engineering [2,62–71] expose the lifecycle of software languages and many different ecosystems for supporting the lifecycle, thereby also suggesting some aspects of FSL. Some of this work is so highly structured that it directly suggests entities and relationships of FSL. For instance, Lämmel, Schauss et al.’s work [70,72] on a chrestomathy of DSL implementations systematically organizes features of language implementation based on metaprogramming systems and also adds semantic labels for different classes, for example, technology, programming concepts, and involved languages. Classification hierarchies (taxonomies) or slightly richer quasi-ontological structures have been proposed for a number of types of software languages: programming languages [4,5], visual languages [73], model transformation languages (and approaches and tools) [74–77], business rules modeling languages [78], software architecture description languages [79], and possibly others. Domain-specific languages — as a larger container of software languages — were also categorized by means of a mapping study [80]. Even the entirety of software languages (or computer languages per Wikipedia terminology) were categorized through a mining and cleaning-based approach aiming at alignment with Wikipedia categories [81]. Rather than aiming at the classification of languages, there are additional categorization efforts covering some aspects of programming (programming languages) or software engineering, for example, a classification of typing approaches in programming languages [82], a taxonomy for software re-/reverse engineering [83], and a chrestomathy of programming techniques and technologies relying heavily on feature modeling and other means of semantic data enrichment [18]. A quasi-ontological approach in software engineering is megamodeling [11,13,84–88] where software systems are modeled in terms of types of artifacts, technologies, and languages involved, subject to rich relationships including conformance, set-membership, part-whole, and structured correspondence. These megamodels provide ontological insights into particular patterns of technology usage or system development, for example, parsing [89], model-driven engineering [90], or software transformations (coupled transformations, specifically) [91]. Megamodels also relate to the notion of technological spaces [3], which can be seen as an ontological pattern relating primary and secondary artifact types, some notion of schema or metamodel that supports conformance, and some forms of languages for querying, transformation, and others. Typical technological spaces are grammarware [92], model-driven engineering [93], SQLware, XMLware, JSONware, and Ontoware. We incorporate some central aspects of megamodeling and technological spaces into (the first release of) FSL. Software language engineering is software engineering for languages. Thus, the ‘body of knowledge’ on software engineering applies [94–96]. There are also other bodies of knowledge in neighboring territories, most notably in model-based engineering [97]. We can draw from these BOKs for software engineering-related entity types in FSL.
12 of 40
2.3. Phases of FSL Development The overall objective was to arrive at a ‘MVP’-like (minimum viable product) first release of FSL, to be able to publish this first release, to enable its use and review, to publish the underlying ontology engineering (OE) effort in a scholarly manner, and thereby to lay the foundation of continuing the work at a larger scale with more domain experts and contributors involved. To this end, we went through the phases as summarized in Table 1. Phase 1 — Taxonomy Initialization Phase 2 — Category Discovery Phase 3 — Taxonomy Enrichment Phase 4 — Property Discovery Phase 5 — Technological Space Coverage Phase 6 — Modularization Phase 7 — SL Concept Coverage Phase 8 — SE Activity Coverage Phase 9 — Review Table 1. The phases of ontology engineering for V1 of FSL.
Here is a summary per phase: Phase 1 — Taxonomy Initialization (Sec. 2.3.1) This is the trivial starting point. Using a manually designed small seed set of SLE formalisms, we begin with a simple taxonomy. We start exercising an overarching OE principle, i.e., to connect all FSL resources, as much as possible, to knowledge resources. We assume that the seed set will effectively serve as a focus group for validation in any of the subsequent phases. Phase 2 — Category Discovery (Sec. 2.3.2) We deliberately do not distinguish categories versus individuals in Phase 1. Instead, we establish this distinction separately in this phase so that we can perform a small test regarding the use of GenAI for individual/category distinction (i.e., classification). We have certain expectations regarding the categories to be discovered. For example, we have multiple closely related calculi in the seed set. Phase 3 — Taxonomy Enrichment (Sec. 2.3.3) With some categories in place, we want to make the classification more meaningful by aiming at a very limited completion effort focusing on categories and individuals in the immediate ‘neighborhood’ of the existing entities. As a result, the new entities are in the scope of our expertise, which simplifies the task of validating the use of GenAI for completion. Phase 4 — Property Discovery (Sec. 2.3.4) We begin to introduce properties for FSL entities such as ‘uses’ and ‘isSpecifiedBy’. We leverage the fact that FSL features different types of entities, in particular: more calculus-like entities (e.g., lambda calculus) versus more approach-like entities (e.g., denotational semantics). Property discovery is continued through phases 5, 7, 8. Phase 5 — Technological Space Coverage (Sec. 2.3.5) Based on our understanding of technological spaces as an important quasi-ontological concept in the SLE context, we integrate such spaces into the emerging ontology. That is, we agree on a number of established technological spaces as well as relationships between them and languages, tools, artifacts, and formalisms. Tools and artifacts are introduced at this point, subject to new root classes of the ontology. Phase 6 — Modularization (Sec. 2.3.6) The emerging ontology eventually requires modularization to remain manageable and understandable. Modularization can centralize
13 of 40
the tbox and group entities based on the top-level entity type. That is, on top of the tbox, there are aboxes for languages, tools, formalisms (and such), and more conceptual entities. Naming and details of modular decomposition can evolve over the remaining phases. Phase 7 — SL Concept Coverage (Sec. 2.3.7) An important partnering subject area of SLE, also known for rich knowledge resources, is ‘programming language theory’ and ‘programming paradigms’ or more broadly ‘programming language concepts’. At this stage, we aim at capturing programming language concepts in FSL, associating programming language entities with them and generalizing the hierarchy of programming language concepts to become meaningful for all software languages. Phase 8 — SE Activity Coverage (Sec. 2.3.8) Likewise, SLE is software engineering for languages; also, languages serve software engineering. Thus, we shall also determine a more SE-centered sub-taxonomy to serve as a foundation for linking the rest of FSL to software engineering. We suggest that SE shall be approached in terms of SE activities (such as requirements, design, implementation, etc.), with refinement into more specific activities (such as code generation under implementation). In this manner, those activities can also be associated with languages, tools, artifacts, etc. or, in fact, types thereof. Phase 9 — Review (Sec. 2.3.9) Up to this point, we primarily pushed for ontology completion with validation and normalization as subordinate goals. During this phase, we review the emerging FSL ontology in depth. For example: What (SHACL) shapes should be devised? How could classes be defined definitionally or what disjointness should be claimed? Also, we devise a set of (SPARQL) queries to summarize the emerging knowledge graph at the meta-level, thereby helping us to arrive at a regular and explainable structure. We now describe the phases in detail. For each phase, we present goals and nongoals. The overall phase-like process, as summarized above, was designed up front as a methodology, but some details (a few of the more subordinate goals or non-goals) in the following description only emerged during execution of the process. 2.3.1. Phase 1 — Taxonomy Initialization This phase relates to the ‘expectation’ step and (the beginning of) the ‘conceptualization’ step in the process for ontology engineering; see Sec. 2.1.2. That is, we specify our expectation by identifying a seed set of stereotypical entities of interest to form a taxonomy. We also specify prioritized subject areas of interest to be associated with the entities in the seed set. We submit to the following (non-) goals: Goals Subject Area Agreement (SAA ) We agree on subject areas of interest (such as software language engineering or programming language theory) and prioritize them. Seed-Set Agreement (SSA ) We agree on the seed set of entities while respecting the aforementioned priorities, that is, the higher the priority of an area, the more associated entities in the seed set. FSL, in this phase and any phase to come, must cover the seed set. Knowledge Resource Linkage (KRL ) We want all elements from the seed set to be linked to external (knowledge) resources. This could be, for example, DBpedia or Wikipedia. In this phase, we look up those resources manually to be certain.
14 of 40
Non-goals Individual/Category Distinction (ICD ) At this stage, we do not aim yet at distinguishing individuals versus categories. We withhold some expectations for such a distinction until later. This is obviously a very low bar, but still a good way to get off the ground. 2.3.2. Phase 2 — Category Discovery We have not yet distinguished individuals versus categories (see the non-goal ICD above) — in part to test whether automated (GenAI-based) classification (category discovery) can be used in our domain context. Once we have established categories, we can expect to use them for few-shot classification subsequently, as more entities arise. We submit to the following (non-) goals: Goals Individual/Category Distinction (ICD ) This was a non-goal before. Some of the entities in the seed set may be categories; alternatively, we may also need to introduce some categories to group entities. For instance, we expect to discover some category related to the calculi in the seed set. In the end, all seed-set entities must be categorized or serve as a category. Non-goals Subject Area Preservation (SAP ) During process execution, we realized that it would be challenging, if not distracting, to maintain Subject Area Agreement (SAA ) along with enriching the taxonomy — also in the view of the inital priorities serving as a constraint. Thus, we allowed ourselves to continue, for now, without subject areas. 2.3.3. Phase 3 — Taxonomy Enrichment As we have started with a fairly small seed set, there is much room for completion. For now, we make a very conservative step. That is, essentially, we look for siblings of entities and classes already at hand. This again serves as small test in GenAI-based completion. We submit to the following (non-) goals: Goals Double Instantiation Enforcement (DIE ) There must not be a class that is instantiated by not at least two individuals because this would be a sign of spurious or overfitted or insufficiently exercised classification. Double Subclassing Enforcement (DSE ) There must not be a class that is subclassed by only one class for essentially the same reason as under DIE. OWL 2 Punning Enforcement (OPE ) The categorical and the individualized levels cannot always be fully separated in ontology engineering. Consider the following example in the FSL domain: ‘process calculus’ may be an entity of interest in certain assertions, but it can also serve as a classifier for more specific process calculi. Therefore, we employ OWL 2 punning (metamodeling) in FSL. Non-goals Individual/Category Completeness (ICC ) In this phase, we limit ourselves to discover only entities needed to address goals DIE and DSE. In general, the first release of FSL does not aim at a reasonable complete taxonomy, but only at satisfying certain coverage criteria.
15 of 40
2.3.4. Phase 4 — Property Discovery We have focused on classification so far. In this phase, we aim at discovering relationships other than subclassing and classification on ontological entities so that we move beyond a mere taxonomy. We submit to the following (non-) goals: Goals Object Property Introduction (OPI ) Pre-ontological related work (Sec. 2.2.2) suggests relationships related to, for example, ‘usage’ (uses) and ‘parthood’ (hasPart); also, the formalistic nature of some ontology entities shall be exercised by a relationship ‘isSpecifiedBy’ so that one entity (e.g., a formal system) is specified by (e.g., in the sense of its semantics) another entity. (As an aside, we also need annotation and datatype properties, in addition to object properties.) Here are illustrative examples: :DenotationalSemantics :uses :AbstractSyntax . :AttributeGrammar :hasPart :ContextFreeGrammar . :CommunicatingSequentialProcesses :isSpecifiedBy :OperationalSemantics .
Entity/Assertion Discovery (EAD ) In exploring different candidate object properties and their nuances in their semantics, new individuals and categories naturally arise or they are actively discovered for populating the new properties with assertions. Subject Area Recovery (SAR ) The association of FSL entities with areas is now to be modelled by an ‘hasArea’ property, without any constraints regarding prioritization, but with cardinality expectations such that entities should associate with at least 1 area and with 3 areas at maximum. Those associations are mostly discoved automatically (GenAI-based) from here on. Subject area associations can be automatically validated for the original seed set (Sec. 2.3.1). Non-goals Entity/Assertion Completeness (EAC ) This non-goal generalizes the earlier nongoal Individual/Category Completeness (ICC ) so that we include assertions. That is, neither do we aim at completeness for the entities (individuals and categories) in FSL, nor do we aim at (relative) completeness for the assertions on top of the entities that are in FSL (first release). We limit ourselves to satisfying certain coverage criteria such as the one of the next phase. 2.3.5. Phase 5 — Technological Space Coverage As we have argued in Sec. 2.2.2, technological spaces provide an important ontological structure for software languages and their foundations. By adding coverage for them in FSL, we expect to also discover important types of entities and relationships. We submit to the following (non-) goals: Goals Technological Space Coverage (TSC ) We have to introduce a new root category of technology spaces and populate it with actual spaces. Technological spaces may be associated with languages, tools, and (types of) artifacts via a ‘hasSpace’ property. Megamodeling Coverage (MMC ) The general structure of a technological space can be understood as an abstract megamodel with each individual space corresponding to a more concrete megamodel. In this manner, we expect to recover typical relationships such as ‘conformsTo’, ‘elementOf’, ‘definedBy’, ‘transforms’, and
16 of 40
‘processedBy’; we also expect to discover the essential notions of ‘artifact’ and ‘tool’. We also expect to discover certain types of languages, tools, and artifacts. Entity/Assertion Discovery (EAD ) — continued As we decide to include certain technological spaces into FSL, we will undoubtly need to include more entities supporting these spaces and more assertions to make the new entities participate in existing properties. For example, we need to need to include additional types of languages, for example, query and transformation languages and metalanguages. Knowledge Resource Linkage (KRL ) — continued Linking to DB- or Wikipedia is not sufficient for such highly specialized community notions such as megamodeling and technological spaces as well as the involved key relations such as conformance. In cases like this, we shall use scholarly references to link FSL entities to supporting knowledge resources. Non-goals Entity/Assertion Completeness (EAC ) — continued Completeness with regard to technological spaces specifically is also not a goal, even though one could argue that there are only finitely many. However, we want to avoid getting lost in obscure space options or nuances of abstraction (e.g., Tableware versus SQLware), as this would otherwise massively drive up the size of the ontology at an too early stage. Likewise, we do not aim at full specification of all spaces in terms of the assertions conceivable. 2.3.6. Phase 6 — Modularization At this stage, FSL is expected to be of a size and complexity that it helps to introduce an explicit modular structure for better scoping and usability of the emerging ontology. We submit to the following (non-) goals: Goals Modules per Entity Types (MET ) We aim at a simple modular structure as follows. There shall be a core module — tbox — with all the various top-level types and relationships between them. There is one module per top-level entity type, thereby grouping all assertions by subject. There is one more module that combines all these modules. All modules but tbox are essentially aboxes. Top Level Refactoring (TLR ) During process execution, we realized that top level of the class hierachy was not clearly enough organized to enable the modular splitting. We had to refactor the class hierarchy to have a manageable number of demarcated roots. Non-goals None 2.3.7. Phase 7 — SL Concept Coverage As we have argued in Sec. 2.2.2, (programming) language classification provides an important ontological dimension for software languages and their foundations. We begin by discovering programming language concepts including those related to programming language paradigms. Subsequently, we aim at generalizing the hierarchy of language concepts to cover types of software languages other than programming languages. We submit to the following (non-) goals: Goals
17 of 40
Programming Paradigm Coverage (PPC ) The hierarchy of language concepts must include the most established programming paradigms, for example, imperative, functional, logic and object-oriented programming. PL Concept Coverage (PCC ) The hierarchy of language concepts must include branches for concepts other than paradigms, for example, concepts related to typing, abstraction, and evaluation strategy. We do not expect that the distinction of programming paradigms can dominate the concept hierarchy; we rather assume that multiple dimensions of classification go side by side. Programming Language Coverage (PLC ) We need to perform entity/assertion discovery for actual programming languages that meet the following two criteria: a) they are popular so that the concept-based elaboration for them is meaningful; b) we are familiar enough with these languages so that we can provide or validate assertions for them such that actual languages are associated with actual concepts. SL Concept Coverage (SCC ) The hierarchy of language concepts must be scaled up further to become meaningful to language categories other than programming languages. As discussed in Sec. 2.2, there exist classification schemes for various language categories; they shall be modeled — to some extent. For instance, we expect to include concepts immediately related to querying, transformation, and metamodeling. We also need to discover assertions for them so that actual software languages associate with actual concepts. Non-goals Entity/Assertion Completeness (EAC ) — continued There are competing classification schemes for programming languages and other types of languages; there are highly detailed (if not idiosyncratic) classification schemes for languages other than programming languages. In the current cycle, we only aim to include and exercise major concepts. We defer a proper taxonomy integration effort starting from a multitude of classification schemes to another time. 2.3.8. Phase 8 — SE Activity Coverage Software languages (and SLE as a discipline) are highly intertwined with software engineering (SE). We can contribute to the foundations of software languages by starting a ‘chase’ from the software lifecycle and its underlying SE activities to associate them to related languages, tools, and artifacts or, in fact, types thereof. We submit to the following (non-) goals: Goals Software Lifecycle Coverage (SLC ) The phases of the software lifecyle (e.g., requirements, design, implementation, deployment, maintenance) provide a good toplevel layer for the hierarchy of SE activities to be included into FSL. For example, under ‘implementation’ we may expect SE activities such as ‘programming’ and ‘code generation’. Technological Space Coverage (TSC ) — continued More or less specific SE activities can be naturally associated with (types of) languages, tools, and artifacts. In this manner, we also contribute to TSC because the entities associated with SE activities can also be associated with spaces (e.g., via ‘hasSpace’). Non-goals
18 of 40
Entity/Assertion Completeness (EAC ) — continued We do not aim at a comprehensive model of SE activities and the corresponding discovery of assertions towards entity types. In particular, we do not aim at covering the SE Body of Knowledge [94–96] in terms of SE activities and their relationships to (types of) languages, tools, and artifacts. 2.3.9. Phase 9 — Review Review is a continuous activity along ontology engineering, but we concentrate some goals in the last phase of the process for the first release of the emerging FSL ontology: Goals Reasoning-Based Consistency (RBC ) The ontology must pass (w.l.o.g.) the HermiT reasoner. Inconsistencies could arise from the use of constraint forms such as class disjointness or equivalence. Shape-Based Validation (SBV ) The ontology must pass (w.l.o.g.) SHACL shapes for basic validation constraints on all entity types. In particular, we may insist on certain CWA-style assertions. Query-Based Reporting (QBR ) The ontology must be explained through (w.l.o.g) SPARQL queries demonstrating the various patterns in which entities associate with each other through assertions. Text-Based Transformation (TBT ) All bulk changes due to identified shortcomings during review must be properly documented as textual specifications suitable for GenAI-based execution of ontology transformations. Non-goals Release-Blocking Issues (RBI ) In Sec. 4.2.1, we discuss relatively obvious limitations regarding the first release of FSL. We do not accept any of these shortcomings as release-blocking issues because it is important to get more SL(E) experts involved and therefore to expose FSL as early as possible. 2.4. Use of Generative AI We now summarize the usage of GenAI in this research. We explain why GenAI was used in this research and how this usage improved the methodology and its execution (productivity). We also clarify why GenAI was not used in certain scopes. GPT 5.2 or 5.3 were used for all cases of GenAI usage mentioned subsequently. Overall, GenAI was used to improve productivity during execution of the ontology-engineering methodology, while the initiative always remained with the researcher. See Table 2 for an overview; details follow: 1.
2.
Methodology GenAI played no role here because the choice of methodology was relatively straightforward within the given context. We wanted a methodology that is well in line with ontology engineering best practices, i.e., with common steps for expectations, reuse, conceptualization, formalization, and validation. The methodology is original (i.e., domain-specific) insofar as technological spaces, programming concepts, and software engineering activities are important, but these types of entities are closely related to the expertise of the author and the domain of FSL. Publication This paper is almost exclusively written by the author, apart from AIbased language assistance; two relatively minor cases of GenAI usage are worth pointing out: a)
Related work The author is fully aware of the literature in the FSL domain, which is also evident from the citations at hand and from the fact that much
19 of 40
Aspect
AI?
Comment
Methodology
−
Follow OE best practices and SLE needs
Publication ▷ Paper structure ▷ Related work ▷ Illustrations ▷ Other sections and parts
•
Ontology ▷ Entities ▷ ▷ Discovery ▷ ▷ Classification ▷ ▷ Linking to Wikipedia
▷ Properties ▷ ▷ Discovery ▷ ▷ Naming ▷ ▷ Assertions ▷ Axioms ▷ ▷ Domain/range axioms ▷ ▷ Class axioms ▷ ▷ Property axioms ▷ Transformations
−
• •
−
• • • • • • −
• • • − −
Defined by journal and methodology Used for search on some specific topics Some illustrations are generated with the LLM Not counting AI-based improvement of text
Seeded entity discovery Few-shot classification Entity linking (LLM+Google) Ontology-specific properties require expert initiative Exploration of naming options informed by prior art Few-shot completion
•
Part of manual discovery process for properties Ontology-specific axioms require expert initiative Some obvious options were auto-generated
•
Textual descriptions of OWL transformations
Table 2. GenAI usage in this paper and in the underlying research (none: −; some: •, •, •)
b)
3.
of the author’s contributing work dates back quite a while (culminating in the Software Languages Book [2]). However, we leveraged GenAI to better cover topics in ontology engineering (see Sec. 2.1) and more specifically CI/CD in that space (see future work discussion in Sec. 4.2). We used GenAI for search and discovery in tandem with search on Google Scholar. Illustrations Some of the figures of Sec. 1.3 were created via controlled GenAI usage to improve productivity. Validation of these artifacts was straightforward. We also used chain-of-thought prompting to separate raw data extraction from visualization, thereby simplifying validation. By contrast, all tables (like those needed for quantification of results in Sec. 3) are based on custom Python/SPARQL code (included in the FSL GitHub project) to avoid any sort of hallucination, to be precise regarding quantification, and to enable reproducibility.
Ontology With the FSL ontology being the primary result of this research and with the underlying ontology engineering process being the primary concern of the research methodology, we must discuss GenAI usage from an ontology-focused perspective. a)
b)
Entities Across the phases, we discover individuals and categories — often with the help of GenAI. The common pattern is that we provide initial, ‘authoritative examples’ (as in a seed set) request additional candidates from GenAI (seeded entity discovery) and request classification (few-shot classification), followed by review and selection by us. This is often an iterative process. For example, rather than trying to enumerate all conceivable software engineering activities ourselves or mapping one specific resource to a formal taxonomy ourselves, we leverage GenAI to collect types of activities at varying degrees of detail; see Sec. 2.3.8. Review often relies on supporting Wikipedia pages for a given subject. Properties Properties were identified (‘discovered’) solely by the author, who took responsibility for shaping the emerging ontology in this respect. That is, the initial introduction of a property — formally by means of an OWL property declaration — is always a human-led initiative. In theory, we could ask GenAI for additional proposals for properties; this may be a sensible thing to do, as
20 of 40
c)
d)
we assume that the underlying model knows of ontologies and can compare the emerging FSL ontology with existing ontologies and relate to best practices. In practice, so far, we find it demanding enough to work through our own expectations regarding properties. Naming for properties can be difficult, also sometimes in combination with deciding on the primary direction of an invertible property. For example, we were often struggling with names such as ‘has...’, ‘uses’, ‘serves’, ‘facilitates’. We used GenAI to discuss naming questions on the grounds of specifying the intended semantics of the relationship. We could have searched the web or looked through ontologies for inspiration instead. Using GenAI simply increased productivity. As the ontology increased in size, it helped to use GenAI for assertion discovery (few-shot completion) based on examples that we provided and subject to review of suggested assertions. For example, a number of programming languages (with which we are familiar) were associated with programming concepts by GenAI; see Sec. 2.3.7. Axioms domain/range axioms are part of the conception of property declarations, which is an author-conducted responsibility. Authoring class axioms (other than subclassing, see ‘classification’), i.e., ‘DisjointClasses’, etc., is also an author-conducted responsibility, as such axioms require very specific domain knowledge. Some of the more obvious property axioms (e.g., ‘inverseOf’) could be suggested by GenAI, when encountering common scenarios (e.g., ‘IrreflexiveProperty’ in the context of mereology). Transformations Throughout the phases, the emerging ontology was rarely edited ‘manually’ and, if so, only locally (such as for naming and comments). All other steps of introducing new entities, properties, assertions were verbally described and realized as OWL-level transformations executed by means of GenAI. The ability to leverage GenAI here was massively helpful in improving productivity. The initiative regarding these transformations remained completely with the author. Validation was straightforward — diff-based and subject to exploration in Protégé.
3. Results We will present quantitative and qualitative results phase by phase. The intermediate stages of FSL can be observed in its GitHub repo; there is a designated overview for the first release covered by the present paper.15 Where GenAI was used to operationalize steps of some phases, transcripts of the conversations are available from the overview. 3.1. Phase 1 — Taxonomy Initialization Based on our understanding of the notion of foundations of software languages, we assume certain ‘subject areas of interest’; see the goal Subject Area Agreement (SAA ) (Sec. 2.3.1); we assign priorities to them to rank their relevance, but also to align with our expertise; see Table 3 for the result. Accordingly, we define a seed set with entities for FSL; see the goal Seed-Set Agreement (SSA ) (Sec. 2.3.1); see Table 4 for the result. The selection is biased by our background and expertise, but the sampling is relatively broad because of the need to cover diverse subject areas of interest; the higher the priority of a subject area, the more corresponding entities there are in the seed set. The seed set will effectively serve as a focus group for validation in all subsequent phases. 15
https://github.com/softlang/fsl/tree/main/misc/1st_release
21 of 40
Subject area to be sampled
Priority
Software Language Engineering (SLE) Compiler Construction (CC) Programming Language Theory (PLT) Theoretical Computer Science (TCS) Software Engineering (SE) Formal Methods (FM) Knowledge Representation & Reasoning (KR) Computational Linguistics (CL) Natural Language Processing (NLP)
High High High Medium Low Low Low Low Low
Table 3. Subject areas of interest Foundational entity
Reason for inclusion
Area(s)
Context-free grammar (CFG) Parsing expression grammar (PEG) Extended Backus-Naur Form (EBNF) Regular grammar Attribute grammar Term-rewriting system Lambda calculus Untyped lambda calculus Simply typed lambda calculus System F Lambda cube Denotational semantics Operational semantics Axiomatic semantics Process calculus Communicating sequential processes (CSP) Calculus of communicating systems (CCS) UML state machine Hoare logic Description logic Dependency grammar
A fundamental formalism in parsing A more modern approach to parsing A metasyntax (i.e., a syntax for syntaxes) A fundamental formalism in parsing, in fact, scanning An important language-implementation technique A formal and rule-based transformation approach All lambda calculi used across PLT, TCS, etc. The basic, untyped lambda calculus The simply typed lambda calculus A lambda calculus with polymorphic types All typed lambda calculi The denotational style of semantics specification The operational style of semantics specification The axiomatic style of semantics specification All calculi for modelling concurrent systems A concrete calculus on concurrency A concrete calculus on concurrency A modeling language for finite automata A prominent approach for program verification A prominent family of logics for ontologies Grammatical theories based on dependency
CC, SLE CC, SLE CC, SLE CC, SLE CC, SLE CC, SLE PLT PLT PLT PLT PLT PLT PLT PLT PLT, TCS PLT, TCS PLT, TCS FM, SE FM KR CL, NLP
Table 4. Seed set for FSL
All entities were linked to Wikipedia (according to goal KRL as of Sec. 2.3.1), that is, we favor Wikipedia for systematic linkage of external knowledge resources for now. A common discussion topic in ontology engineering with regard to linking to Wikipedia, Wikidata, or DBpedia is which of these sources to use and what type of link to use (‘sameAs’ or ‘seeAlso’ or ‘primaryTopic’ or ‘page’). At the risk of sounding unorthodox, we decided in favor of Wikipedia because of how important human validation and therefore understanding the narrative of Wikipedia pages is. Also, we were looking into the common ‘sameAs’ and ‘seeAlso’ challenge and decided eventually to prefer FOAF-style ‘isPrimaryTopicOf’ and ‘page’, which are also prepared for the identity issues on Wikipedia. We also allow links with anchors for subpages. 3.2. Phase 2 — Category Discovery The main goal of this phase was Individual/Category Distinction (ICD ) (Sec. 2.3.2); see Fig. 7 for the resulting hierarchy. As one can observe, some of the entities from the seed set are set up to function as categories, whereas most entities from the seed set are considered individuals. In order to properly categorize all of the individuals, we had to include a number of categories. In this paper, we did not (plan to) discover additional individuals. It is also worth noting that we have a flat classification hierarchy, that is, there is no subclassing yet between the categories. 3.3. Phase 3 — Taxonomy Enrichment The main goals of this phase were Double Instantiation Enforcement (DIE ) and Double Subclassing Enforcement (DSE ) (Sec. 2.3.3); see Fig. 8 for the resulting hierarchy. As one can observe, we mainly added instances in response to DIE. As our classificatiobn hierarchy is flat, there was nothing to be done with regard to DSE. Clearly, this property should be monitored throughout subsequent expansion. However, we added a new category
22 of 40
Legend: rectangles for individuals; ellipses for categories. Nodes added in this phase are highlighted to visualize the delta between this phase and the seed set. Figure 7. Result of Phase 2 — Category Discovery
Legend: rectangles for individuals; ellipses for categories. Added categories and individuals are highlighted to visualize the delta between this phase and the previous phase. Figure 8. Result of Phase 3 — Taxonomy Enrichment
23 of 40
for ‘Classical Lambda calculus’ to better deal with the overloaded meaning of ‘Lambda calculus’. 3.4. Phase 4 — Property Discovery This phase aimed at the introduction of an initial set of properties. Property discovery was continued in phases 5, 7, 8, as new types of entities were included. In Table 5, all properties (not just those initially introduced in this phase) are listed with an explanatory comment to get a first impression of the kind of properties discovered. In Table 6, additional metadata about the properties is provided. From the number of assertions, we can infer the modest size of FSL in its first release, calling for future growth (‘completion’). 3.5. Phase 5 — Technological Space Coverage All goals of this phase (Sec. 2.3.5 are directly aimed at the addition of the notion of technological spaces to FSL. Table 7 summarizes the result — essentially by showing how (types of) languages, tools, and artifacts associate with technological spaces. We refer back to Fig. 6 for illustration where we see indeed all such different types of entities which engage in ‘conformsTo’, ‘processes’, ‘serves’ et al. relationships. We also experimented with a variation on Knowledge Resource Linkage (KRL ) in this phase in so far that scholarly references had to be provided for technological spaces and the related notion of megamodeling as well as key concepts involved in megamodeling, notably conformance and membership; see Table 8 for a first attempt. We are not satisified with this status. In particular, there is no place yet in FSL, where the notion of megamodeling would be directly reified; we use ‘MegamodelArtifact’ as a proxy for now. As an aside, the phase-by-phase discussion of results focuses on the main goals per phase, while in reality extra modifications were performed on the emerging ontology, as it could be observed from the intermediate versions and, in many cases, also from the transcripts of the LLM conversations. For example, in this phase, we actually i) established the ultimate name of FSL; ii) we introduced ‘extends’ for some prior uses of ‘hasPart’; iii) we performed another renaming on what (later) would be renamed again to the final category ‘FormalEntity’. 3.6. Phase 6 — Modularization The main goal of this phase was the modularization of FSL. To this end, we would first address the goal Top Level Refactoring (TLR ) (Sec. 2.3.6) to review and improve the top-level of the entity-type hierarchy. The top-level plus some selected branches were previously shown in Fig. 3. Quantities for the top-level entity types are shown in Table 9 — as of the end of preparing the first release. Because of punning/metamodeling — OWL 2 Punning Enforcement (OPE ) (Sec. 2.3.3) — we count subclasses also as ‘entities’. In addressing the goal Modules per Entity Types (MET ), we obtained the modules listed in Table 10 with the dependencies show in Fig. 9 — as of the end of preparing the first release; we should note that the module ‘pe’ was only introduced in Phase 7 — SL Concept Coverage (Sec. 2.3.8). The import dependencies are in line with the idea that the tbox is completely centralized in one module; all other modules are aboxes — with subclass- and possibly class-related constraints though. The namespace dependencies are a consequence of how subjects of a given namespace associate with objects of other namespaces. As an aside, we realized only in this phase that we had forgotten an important subject area of interest: MDE (model-driven engineering). This turned out not to be a problem because relevant entities were discovered during Phase 5 — Technological Space Coverage (Sec. 2.3.5), as ‘modelware’ is an important technological space which closely aligns with MDE.
Relates a reified assertion about language concepts with the concept Relates a reified assertion about language concepts with the language Relates an artifact to another artifact in the generalized sense of metamodels classifying models Documents the ontology-wide expectation regarding commenting Relates an artifact to another artifact in the generalized sense of models conforming to metamodels Relates an artifact to a language in the sense of a definition (specification) Relates an artifact to a language in the set-theoretic sense Relates a software language to another software language that it extends or generalizes Documents the ontology-wide expectation regarding formatting Relates an activity or a tool to an artifact kind that it typically affects (thereby complementing in- and output) Relates an entity to an important subject area of interest Relates a technological space to a kind of artifact that it is concerned with Relates an entity to the BibTeX key of a bibliographic entry that documents, discusses, or supports it Relates a language concept to a dominated form Relates an activity or a tool to an artifact kind that it typically requires or consumes as input Relates an activity or a tool to an artifact kind that it typically produces as output Relates a composite language concept to a (primary) component Relates a reified assertion about language concepts with the primary mechanism for realization Relates a composite language concept to a (secondary) component Relates an entity to a technological space to which it characteristically belongs Relates a reified assertion about language concepts with the type of support Relates a software engineering activity to a language kind that is characteristically used to perform or express the activity Relates a software engineering activity to a tool kind that characteristically supports carrying out the activity Relates a technological space to a kind of tool that it is concerned with Relates an entity to a methodological approach used to formally specify or define its behavior or meaning Documents ontology-wide expectations for linking classes and individuals to descriptive web pages Documents ontology-wide expectations for explicit punning-based typing of classes and individuals Relates a language to a concept that it supports natively Relates a tool entity to an artifact entity that it characteristically processes Relates an entity to a methodological approach that it serves, realizes, supports, or instantiates in an important way Relates a language to a concept that it supports standardly Relates a language to a concept that it supports Relates a tool entity to a language entity that it targets Relates an ontology entity to another ontology entity that it uses conceptually, methodologically, or technically
assertedForConcept assertedForLanguage classifies commentingPolicy conformsTo defines elementOf extends formattingPolicy hasAffectedArtifact hasArea hasArtifact hasBibTeX hasDominantForm hasInputArtifact hasOutputArtifact hasPrimaryComponent hasPrimaryMechanism hasSecondaryComponent hasSpace hasSupportKind hasSupportingLanguage hasSupportingTool hasTool isSpecifiedBy linkingPolicy metamodelingPolicy nativelySupports processes serves standardlySupports supports targets uses
Table 5. FSL properties — Overview I/II
Comment
Property
24 of 40
422 379 186 126 39 38 29 27 26 23 23 22 21 19 16 15 14 13 13 13 13 12 12 11 9 9 9 7 6 5 3 3 2 1
hasArea nativelySupports hasSpace supports targets hasTool processes uses elementOf hasSupportingLanguage hasPrimaryComponent hasSecondaryComponent hasArtifact hasInputArtifact hasOutputArtifact hasSupportingTool metamodelingPolicy assertedForConcept assertedForLanguage hasDominantForm hasSupportKind serves hasPrimaryMechanism hasBibTeX commentingPolicy formattingPolicy linkingPolicy hasAffectedArtifact conformsTo classifies extends standardlySupports isSpecifiedBy defines
Table 6. FSL properties — Overview II/II
# Assertions
Property
LanguageConcept LanguageEntity LanguageConcept LanguageSupportKind MethodologicalApproach ImplementationMechanismConcept string
ArtifactEntity ArtifactEntity ArtifactEntity SoftwareLanguage LanguageConcept MethodologicalApproach LanguageEntity
LanguageSupportAssertion LanguageSupportAssertion LanguageSupportAssertion LanguageSupportAssertion Entity LanguageSupportAssertion Entity
ArtifactEntity ArtifactEntity SoftwareLanguage LanguageEntity FormalEntity ArtifactEntity
EngineeringActivity
SubjectArea LanguageConcept TechnologicalSpace LanguageConcept LanguageEntity ToolEntity ArtifactEntity Entity LanguageEntity LanguageEntity LanguageConcept LanguageConcept ArtifactEntity ArtifactEntity ArtifactEntity ToolEntity
Range
Entity LanguageEntity Entity LanguageEntity ToolEntity TechnologicalSpace ToolEntity Entity ArtifactEntity EngineeringActivity CompositeConcept CompositeConcept TechnologicalSpace
Domain
specifies isDefinedBy
classifies conformsTo isExtendedBy
isServedBy
isArtifactOf
isTargetedBy isToolOf isProcessedBy isUsedBy hasElement
isSpaceOf
isAreaOf
Inverse
Legend: Type O for object properties, A for annotation properties, D for datatype properties.
O O O O O O O O O O O O O O O O A O O O O O O D A A A O O O O O O O
Type
25 of 40
26 of 40
Metric
Count
instances subclasses spaces_with_languages languages_with_spaces spaces_with_tools tools_with_spaces spaces_with_artifacts artifacts_with_spaces
13 0 13 101 13 26 13 21
Legend: ‘instances’ for the number of technological spaces; ‘subclasses‘ being 0 to mean that no subclassing on spaces is attempted; ‘spaces_with_. . .’ for the number of spaces with assertions to a given entity type; likewise for ‘. . ._with_spaces’. Table 7. Quantities for Phase 5 — Technological Space Coverage
Entity
Reference
MegamodelArtifact MegamodelArtifact Grammarware Modelware TechnologicalSpace conformsTo conformsTo conformsTo conformsTo conformsTo elementOf
[85] [84] [92] [93] [3] [98] [99] [100] [101] [102] [102]
Table 8. hasBibTeX assertions for Phase 5 — Technological Space Coverage
Entity type
# Entities
# Instances
# Subclasses
Language entity Formal entity Conceptual entity Tool entity Artifact entity Property entity
103 70 280 55 28 51
69 44 70 23 0 51
34 26 129 32 28 0
Table 9. Quantities for the top-level entity types of FSL
Ontology
Purpose
fsl tbox fe pe le te ae ce ie
FSL as the union of imports Shared TBox of FSL ABox for formal entities ABox for programming languages and their concepts ABox for (other) software languages and their concepts ABox for software tools ABox for software artifacts ABox for conceptual entities ABox for pending issues
Table 10. FSL’s modules
27 of 40
Imports
Namespaces
Figure 9. Dependencies for FSL’s modules
3.7. Phase 7 — SL Concept Coverage In accordance with the goal Programming Paradigm Coverage (PPC ) (Sec. 2.3.7), there are corresponding concepts included in FSL; refer back to Fig. 2 for an illustration. Quantitive results regarding the other goals (PL Concept Coverage (PCC ), Programming Language Coverage (PLC ), and SL Concept Coverage (SCC )) (Sec. 2.3.7) are summarized in Table 11. For example, we see that a significant number of concepts are ‘in use’ and both programming languages and software languages of other types are annotated with concepts. Metric
Count
concepts instances subclasses used_instances used_subclasses properties software_languages programming_languages
156 125 31 118 31 8 47 27
Legend: ‘concepts’ for the number of concepts with the breakdown ‘instances’ and ‘subclasses’ (all being individuals due to punning); ‘used_instances’ and ‘used_subclasses’ for clarifying what concepts are actually exercised in ontological assertions (thereby also providing evidence of punning); ‘properties’ for the number of ontological properties exercised in assertions with concept involvement; ‘software_languages’ for the number of actual software languages with concept-based assertions with ‘programming_languages’ for the fraction thereof being specifically programming languages. Table 11. Quantities for Phase 7 — SL Concept Coverage
3.8. Phase 8 — SE Activity Coverage In Table 12, quantitative results regarding the goals of this phase — Software Lifecycle Coverage (SLC ) and Technological Space Coverage (TSC ) (Sec. 2.3.8) — are summarized. The top-level of the resulting hierarchy of SE activities exactly corresponds to the software lifecycle. All subclasses are just more nuanced. The activity types are not yet exercised much. Technological space coverage is best claimed by identifying types of languages, tools, and artifacts being associated with SE activity types; coverage is slim here at this stage.
28 of 40
Metric
Count
instances immediate_subclasses nonimmediate_subclasses used_nonimmediate_subclasses properties language_instances language_subclasses tool_instances tool_subclasses artifacts
0 10 80 17 6 0 10 0 9 6
Legend: ‘instances’ (being 0) as a testament to the fact that FSL does not capture actual activities in actual projects (but applications of FSL very well may do so); ‘immediate_subclasses’ (being 10) illustrates the coverage of major activities in the software lifecycle with ‘nonimmediate_subclasses’ being much larger serving as an indication of the nuanced types of SE activities covered of which (see ‘used_nonimmediate_subclasses’) however only relatively few are already exercised; ‘properties’ for the number of ontological properties exercised in assertions with SE activity involvement; ‘language_. . .’ and ‘tools_ldots’ for the numbers of languages and tools exercised in assertions with SE activity involvement where the zeroes at the ‘. . ._instances’ level indicate that punning is at play such that SE activity types associate with types of languages and tools — rather than specific languages and tools; ‘artifacts’ for the number of artifact types because FSL does not capture actual artifacts in actual projects or repositories and general activity types are also just concerned with types of artifacts. Table 12. Quantities for Phase 8 — SE Activity Coverage
3.9. Phase 9 — Review Reviewing was obviously a permanent task during the development process for the first release of FSL, but we concentrated a more systematic effort in the last phase. We will walk through the corresponding goals as of Sec. 2.3.9 now. Property
# Uses
disjointWith inverseOf minQualifiedCardinality onClass onProperty someValuesFrom unionOf
19 28 4 4 6 2 19
Table 13. OWL constraints (Phase 9 — Review)
The goal Reasoning-Based Consistency (RBC ) (Sec. 2.3.9) does not require any documentation other than that the HermiT Reasoner completes fine on the first release of FSL. However, let us also summarize the constraints for the first release of FSL; see Table 13. One may draw the conclusion that FSL (as of the first release) is relatively underspecified; see the discussion of future work (Sec. 4.2.1). The goal Shape-Based Validation (SBV ) (Sec. 2.3.9) is addressed in Table 14 where we summarize the shapes that readily constrain the first release of FSL; see GitHub for details. There are definitely some areas for improvement. For example, FSL uses punning in a manner that would better be documented and enforced more strongly. The goal Query-Based Reporting (QBR ) (Sec. 2.3.9) has been essentially demonstrated throughout Sec. 3 because all tables and figures are based on systematically authored queries; see Table 15 for a summary.
29 of 40
ClassDeclarationsMustHaveCommentShape Class-declarations must be ‘documented’ by means of a comment. PropertyDeclarationsMustHaveCommentShape Ditto for property declarations. ClassDeclarationsMustHaveFoafLinkShape Class-declarations must be linked with the ‘real world’ by a FOAF link. PropertyDeclarationsMustHaveFoafLinkShape Ditto for property declarations. MetamodelingShapes The different types of entities (especially in terms of their subclasses) must conform to a schema of punning/metamodeling. TechnologySpacesMustHaveConformsToShape Technological spaces must engage in at least one conformance relationship involving two types of space-related artifacts which in turn relate to each other in the sense of conformance. TechnologySpacesMustBeSpecifiedShape Technological spaces must be ‘specified’ sufficiently in terms of related languages, tools, and artifacts. EngineeringActivitiesMustBeSpecifiedShape Software engineering activities must be ‘specified’ in terms of artifacts for input, output, or otherwise consumed or affected, also in terms of supporting languages and tools. MethodologicalApproachesMustBeSpecifiedShape Methodological approaches must be ‘specified’ in the sense that they are connected to other types of entities, most notably in the sense that some formal entity or language serves an approach.
Table 14. SHACL shapes (Phase 9 — Review)
Query theme
# Queries
Purpose
phase2 phase3 entities properties spaces language_concepts se_activities modules constraints references
1 1 1 1 8 8 10 2 1 1
Initial FSL taxonomy tree of phase 2 FSL taxonomy tree for delta between phases 2 and 3 FSL entity type-usage metrics FSL property-usage metrics Specification aspects for technological spaces in FSL Specification aspects for language concepts in FSL Specification aspects for SE activities in FSL Graphs for import and namespace dependencies of FSL Constraint usage metrics for FSL Use of scholarly references in FSL
Table 15. SPARQL queries (Phase 9 — Review)
The goal Text-Based Transformation (TBT ) (Sec. 2.3.9) is more like a guideline for all phases and certainly for the final review phase so that transformations should be described and recorded, thereby resulting in better transparency/reproducibility. In this manner, we are committing to a basic principle that is established all across software engineering and computer science; in the more schema-based context at hand, we also have contributed to such a transformation-based development of specifications — especially in the context of grammars [103–106] (even though typically based on specifications in an appropriate transformation language rather than text-based descriptions, but the latter is now practical with GenAI support). Throughout the development of the first release of FSL, we described several transformations in the LLM-based conversations, as evident from the corresponding transcripts, but in the review phase, we deliberately switched to explicit capture of the transformation in an ontological format; there were three batches of issues; see Fig. 10 for the last one which happens to be relatively small, thus fit for presentation. The idea is that ‘target’ refers to the primary resource that requires transformation, ‘critique’ formulates the problem with the present ontology; ‘suggestion’ describes the actual transformation; ‘resolveAfter’ helps with ordering of issue resolution. The set of issues at hand is obviously concerned with cleaning up some related property naming and entity placement (in namespaces) and module-import dependencies. In Table 16, all issues — ontologically coded as illustrated above — are listed; see GitHub for details. The issues from Fig. 10 are also included in the list. See also the discussion of future work for some outstanding issues (Sec. 4.2.1) and the need to use CI/CD for a more systematic process of issue tracking and resolution (Sec. 4.2.2).
30 of 40
@prefix ... -- omitted <http://www.softlang.org/ontologies/ie> a owl:Ontology ; .., rdfs:comment "Sub-ontology / ABox for issue entities. ..."@en ; owl:imports <http://www.softlang.org/ontologies/tbox> . :SupportsNameClashIssue a tbox:IssueEntity ; tbox:target tbox:supports ; tbox:target ce:supports ; tbox:critique "There are two different types of ’supports’ in use. tbox:supports relates tools to languages, whereas ce:supports relates languages to concepts. They cannot be usefully combined under a shared abstraction, unless we are willing to have longer property names that, perhaps, include the range into the property name for disambiguation."@en ; tbox:suggestion "Rename tbox:supports to tbox:targets, since a tool may be said to target a language."@en . :LanguageConceptsNamespaceIssue a tbox:IssueEntity ; tbox:target ce:LanguageConcept ; tbox:target ce:CompositeConcept ; tbox:target ce:LanguageSupportKind ; tbox:target ce:LanguageSupportAssertion ; tbox:target ce:LanguageConceptApplicability ; tbox:target ce:ImplementationMechanismConcept ; tbox:target ce:supports ; tbox:target ce:nativelySupports ; tbox:target ce:standardlySupports ; tbox:target ce:assertedForConcept ; tbox:target ce:assertedForLanguage ; tbox:target ce:hasDominantForm ; tbox:target ce:hasPrimaryMechanism ; tbox:target ce:hasApplicability ; tbox:target ce:hasComponent ; tbox:target ce:hasPrimaryComponent ; tbox:target ce:hasSecondaryComponent ; tbox:target ce:hasSupportKind ; tbox:critique "Several schema-level language-concept classes and properties are currently declared in the ce module, although they belong to the shared tbox vocabulary. This was previously considered acceptable, but the module boundary should now be cleaned up."@en ; tbox:suggestion "Move all targeted class and property declarations from ce to tbox, change their namespace from ce to tbox, and update all references accordingly."@en ; tbox:resolveAfter :SupportsNameClashIssue . :ModuleImportIssue a tbox:IssueEntity ; tbox:target <http://www.softlang.org/ontologies/ae> ; tbox:target <http://www.softlang.org/ontologies/le> ; tbox:target <http://www.softlang.org/ontologies/pe> ; tbox:target <http://www.softlang.org/ontologies/te> ; tbox:critique "Ideally, all modules other than tbox and fsl should only import tbox. This is violated for no good reason."@en ; tbox:suggestion "Remove the violating imports from the targeted modules."@en ; tbox:resolveAfter :LanguageConceptsNamespaceIssue .
Figure 10. Illustrative sets of issues fixed in Phase 9 — Review
4. Discussion We will now summarize our findings and suggest directions for future research. 4.1. Summary of findings We draw the following conclusions from the conducted research: An ontological definition of the software languages field Up to now, software languages and their engineering have been conceptualized only informally and pragmatically. FSL provides a formalized conceptualization in terms of taxonomic and ontological building blocks. FSL classifies software languages and connects them with language
31 of 40
FormalSystemsDisjointnessIssue conformsToIssue FoundationalToFormalIssue describesIssue LanguagesDisjointnessIssue hasArtifactIssue LinkingPolicyIssue hasReferenceGeneralizationIssue MakePropertiesEntitiesIssue hasSpaceGeneralizationIssue MigrateToBibTeXIssue hasToolIssue MissingPoliciesIssue isAreaOfNormalizationIssue PropertyGraphIssue1 isProcessedByToolIssue PropertyGraphIssue2 servesDomainIssue ToolsDisjointnessIssue supportsNormalizationIssue tboxDependsOnCEIssue tboxCommentingIssue tboxDependsOnFEIssue tboxFormattingIssue LanguageConceptsNamespaceIssue ArtifactsDisjointnessIssue ModuleImportIssue ConceptsDisjointnessIssue SupportsNameClashIssue ConformanceLiteratureIssue FormalEntitiesDisjointnessIssue Table 16. Issues covered along the first release of FSL
tools, language concepts, the underlying formalisms, software engineering activities and technological spaces (as to how they involve languages). A software language-focused instantiation of ontology engineering Every domain requires specific efforts for ontology engineering. We identified the specific aspects required to cover software languages in the FSL context. That is, we included phases of covering language concepts, the underlying formalisms, software engineering activities and technological spaces. All these types of entities are now integrated in FSL and released for use, review, and enhancement on GitHub. Proven usefulness of GenAI in the software languages field We received significant help from GenAI regarding entity discovery and classification, ontology completion, and linkage of knowledge resources. GenAI was also specifically useful for ontology manipulation based on textual specifications of intended transformations that were found to be reliably executable via GenAI. 4.2. Future research directions There are three major directions: i) evolution of FSL in a relatively standard manner; ii) leverage of CI/CD for the continuous improvement of FSL; iii) pairing of the FSL ontology with suitable chrestomathic efforts. We will discuss these three directions in turn. 4.2.1. Evolution of FSL There are many relatively obvious issues regarding the necessary and unsurprising evolution of FSL. We have filed corresponding issues on GitHub.16 We group these issues here as follows: Validation Some issues are about SHACL-based hardening of validation. Completion Other issues are about further ontology completion, as some incompleteness scenarios are readily known, for example, not all technological spaces are sufficiently specified. Exception elimination In fact, several issues deal with validation and completion at once — in the sense that the existing SHACL shapes make exceptions due to observed incompleteness so that validation passes for now. 16
https://github.com/softlang/fsl/issues?q=is%3Aissue%20label%3A%22past%20V1%22
32 of 40
Consistency Other issues are concerned with ontology consistency regarding structure, constraints, and documentation. For example, the systematic and relatively exhaustive use of disjointness constraints and investigation of the use of definitional axioms (‘EquivalentClass’) are part of this theme. Also, the class hierarchy needs some work. For example, ‘tbox:ConceptualEntity‘ combines very different types of entities, which creates some confusion for domain/range properties. Also, the use of policies (for documentation, metamodeling and linking) is somewhat ad hoc at this point. Also, we use domain/range+cardinality constraints in some areas, where these are rather validation-related expectations, as the OWA does not serve our expectations here; we were using these constraints to pass intentions to the LLM. Bibliography We also aim at systematic bibliographic grounding. 4.2.2. CI/CD for FSL CI/CD (Continuous Integration and Continuous Delivery/Deployment) for ontologies would certainly include aspects such as issue templates, SHACL and SPARQL tests, competency-question regression tests, reasoner checks, documentation generation, and defined and checked release criteria. CI/CD is a developing field in ontology engineering and knowledge management. Before we explain how we specifically want to apply CI/CD to FSL and how we expect to make a contribution, we summarize the state of the art. ROBOT [107] is a tool for automating ontology workflows; while it is not framed explicitly in the CI/CD context, the approach covers some of the aspects listed above. ODK (Ontology Development Kit) [108] is a toolkit for building, maintaining, and standardizing biomedical ontologies; it standardizes automatically executable workflows for quality control, dependency management, and releases. (Software-like deployment — the D in CI/CD — is often seen as release engineering in the ontology context.) OntoFlow [109] provides a general workflow-based view on ontology engineering. OnToology [110] automates documentation, evaluation, releasing, and versioning for Git-based ontology development; this is a good example of relatively early CI-style support automation, before the CI/CD label was used more explicitly. By contrast, Ontolo-CI [111] is an explicitly CI-labeled approach which leverages GitHub Actions for RDF and ontology-adjacent validation using ShEx. ACIMOV [112] is a methodology which thus goes beyond tool support and proposes a development method based on modularity, Git workflows, automated syntactic/semantic checks, documentation generation, and publication. SAREF [113] is a strong example of current CI/CD state of the art for ontology engineering; it covers compliance checks, dependency handling, SHACL-based verification, and automated portal generation/publication. Even with knowledge resources that are not strict ontologies, DevOps-/CI/CD-like methods have been explored. For instance, the newer DBpedia Release Cycle [114] concerns a major knowledge graph rather than ontology engineering per se, but the approach is strongly CI/CD-shaped: automated release workflow, testing methodology, regular releases, productivity/agility, and maintainability. The idea of using DevOps principles for semantic data quality dates back much longer [115]. In our own work on online software chrestomathies with semantic labeling [18,70,72,116,117], we also explored simple forms of quality control and dependency management regarding, for example, consistency of Wiki content with underlying source code as well as correct use of a RDFS-like vocabulary applied for feature modeling or architectural modeling. Much of this CI/CD work for ontologies should be relatively straightforward to apply to FSL, for example, GitHub Actions, workflows for SHACL and reasoning, template-based issue generation, competency-question regression tests, release management, documentation generation, etc. — we are currently working on such aspects. Given the early state of FSL’s development and the particularities of FSL’s domain, a few particular concerns are
33 of 40
worth mentioning, which in some cases indeed require research: a) FSL and Wikipedia are linked in a challenging manner, which requires constant monitoring regarding correctness, drift, and discovery in combination with semi-automatic workflows for corrective measures; b) FSL is highly incomplete in terms of available candidate ABox knowledge such that prioritized discovery would be needed to arrive at a manageable process for improvement with the human in the loop; c) FSL is inspired by much pre-ontological knowledge, as we have reported, but the alignment with those resources should be controlled more formally, subject to appropriate semi-formalization of the underlying resources and semi-automated integration thereof. 4.2.3. Pairing of FSL with chrestomathies As discussed in Sec. 2.2.2, we consider software chrestomathies [118] and megamodels in the software (language) engineering context [119] important pre-ontological knowledge resources informing ontology engineering for FSL. Ultimately, FSL as an ontology should converge with a suitable collection of megamodels and a suitable chrestomathy. Let us sketch such a future pairing: YAS YAS17 (‘Yet another SLR (Software Language Repository)’) is a rich collection (chrestomathy) of language-related artifacts serving the demonstration of various language-related aspects (such as specification, implementation, analysis, and transformation). YAS is the repository providing the foundation of the author’s software languages book [2]. Interestingly, YAS internally uses two types of megamodeling languages for multi-level relationship maintenance [119]. It would be relatively obvious to incorporate a referencing layer into FSL to associate with YAS for demonstrative purposes. We actually plan such a layer in the ongoing effort of preparing the next edition of the book, as the book is also meant to benefit from FSL and to cover ontology engineering as a subject (in the sense of ontologies being considered software languages, too). 101companies The 101companies chrestomathy on software languages, technologies, and concepts [18,72,116] and a related metaprogramming-focused spinoff [70,72] suggest another referencing layer for FSL in principle, but 101companies is no longer actively maintained and does not cover current software (language) engineering practice. However, the underlying idea of implementing a small software system idiomatically and in alignment with an underlying feature model across many languages and technologies remains a viable path for a knowledge resource. Coverage of software engineering activities was limited for 101companies; broader coverage would be natural for FSL. Formal entities FSL suggests another chrestomathic direction: formal entities in the context of software languages. Such a chrestomathy would cover formal systems and other formal entities relevant in the context of software languages. In the most direct manner, all formalisms covered in FSL should be properly demonstrated. In a less obvious manner, the existing category of programming (or software) language concepts could be illustrated by appropriately labeled language definitions and uses of formalisms as well as sample programs.
5. Conclusions We have presented the initial development of FSL — an emerging ontology for the foundations of software languages. FSL covers software languages, their categories, tools 17
https://github.com/softlang/yas/
34 of 40
related to them, concepts underlying those languages, software engineering activities involving languages, and technological spaces contextualizing language usage. FSL is already useful for teaching and research in the author’s context in several ways. First, FSL now provides a knowledge resource to inject structure into courses on software language engineering, programming language theory, and ontology engineering. Second, FSL will be leveraged to cover the software language category of ontologies in the next edition of the Software Languages Book (i.e., the 2nd edition following [2]). Some of these efforts will advance FSL; community contributions will be very helpful.
References 1. 2. 3.
4. 5. 6.
7.
8.
9.
10.
11.
12. 13. 14.
Favre, J.M.; Gasevic, D.; Lämmel, R.; Winter, A. Guest Editors’ Introduction to the Special Section on Software Language Engineering. IEEE Trans. Softw. Eng. 2009, 35, 737–741. Lämmel, R. Software Languages: Syntax, Semantics, and Metaprogramming; Springer, 2018. https: //doi.org/10.1007/978-3-319-90800-7. Kurtev, I.; Bézivin, J.; Aksit, M. Technological Spaces: An Initial Appraisal. In Proceedings of the Proceedings of the Federated Conferences CoopIS, DOA, and ODBASE 2002, Industrial Track, 2002. Introduces the concept of technological spaces as ecosystems of related languages, artifacts, tools, and transformations., https://www.researchgate.net/profile/Ivan-Kurtev/ publication/228580557_Technological_Spaces_An_Initial_Appraisal/links/0fcfd507dd6e6 aab8a000000/Technological-Spaces-An-Initial-Appraisal.pdf. Babenko, L.P.; Rogach, V.D.; Yushchenko, E.L. Comparison and classification of programming languages. Cybern. Syst. Anal. 1975, 11, 271–278. Doyle, J.R.; Stretch, D.D. The Classification of Programming Languages by Usage. Int. J. Man–Machine Stud. 1987, 26, 343–360. Combemale, B.; Lämmel, R.; Van Wyk, E. SLEBOK: The Software Language Engineering Body of Knowledge (Dagstuhl Seminar 17342). Dagstuhl Reports 2018, 7, 45–54. https://doi.org/10.4 230/DagRep.7.8.45. Steimann, F.; Freitag, M. The Semantics of Plurals. In Proceedings of the Proceedings of the 15th ACM SIGPLAN International Conference on Software Language Engineering, New York, NY, USA, 2022; SLE 2022, p. 36–54. https://doi.org/10.1145/3567512.3567516. Lee, Y.; Gopinathan, K.; Yang, Z.; Flatt, M.; Sergey, I. DSLs in Racket: You Want It How, Now? In Proceedings of the Proceedings of the 17th ACM SIGPLAN International Conference on Software Language Engineering, New York, NY, USA, 2024; SLE ’24, p. 84–103. https: //doi.org/10.1145/3687997.3695645. Steimann, F.; Stunic, R. The Linguistic Theory behind Blockly Languages. In Proceedings of the Proceedings of the 17th ACM SIGPLAN International Conference on Software Language Engineering, New York, NY, USA, 2024; SLE ’24, p. 113–129. https://doi.org/10.1145/3687997. 3695636. Jansen, N.; Lüpges, A.; Rumpe, B. Lessons Learned from Developing the MontiCore Language Workbench: Challenges of Modular Language Design. In Proceedings of the Proceedings of the 18th ACM SIGPLAN International Conference on Software Language Engineering, New York, NY, USA, 2025; SLE ’25, p. 112–127. https://doi.org/10.1145/3732771.3742717. Bézivin, J.; Jouault, F.; Rosenthal, P.; Valduriez, P. Modeling in the Large and Modeling in the Small. In Proceedings of the European MDA Workshops MDAFA 2003 and MDAFA 2004, Revised Selected Papers. Springer, 2005, Vol. 3599, LNCS, pp. 33–46. Favre, J.; Lämmel, R.; Varanovich, A. Modeling the Linguistic Architecture of Software Products. In Proceedings of the Proc. MODELS. Springer, 2012, Vol. 7590, LNCS, pp. 151–167. Bagge, A.H.; Zaytsev, V. Languages, Models and Megamodels. In Proceedings of the Post-proc. of SATToSE 2014. CEUR-WS.org, 2015, Vol. 1354, CEUR Workshop Proceedings, pp. 132–143. Keet, C.M.; Khan, Z.C. Characterising Competency Questions for Ontologies. In Proceedings of the Proceedings of the Joint Ontology Workshops (JOWO) - Episode XI: The Sicilian Summer under the Etna, co-located with the 15th International Conference on Formal Ontology in Information Systems (FOIS 2025), Catania, Italy, September 8-12, 2025; Beverley, J.; Keet, C.M.;
35 of 40
15. 16. 17. 18.
19. 20. 21.
22.
23. 24.
25.
26. 27.
28.
29.
30.
31.
32.
33.
Lamba, N.; Lambrix, P.; Tiwari, S.; Giorgis, S.D.; Righetti, G.; Sacco, G.; Moreira, J.L.R.; Terkaj, W.; et al., Eds. CEUR-WS.org, 2025, CEUR Workshop Proceedings. https://ceur-ws.org/Vol-41 76/caos-9.pdf. Hilbert, D.; Bernays, P. Grundlagen der Mathematik; Vol. 1, Springer: Berlin, 1934. Hilbert, D.; Bernays, P. Grundlagen der Mathematik; Vol. 2, Springer: Berlin, 1939. Bézivin, J. Model Driven Engineering: An Emerging Technical Space. In Proceedings of the GTTSE 2005, Revised Papers. Springer, 2006, Vol. 4143, LNCS, pp. 36–64. Favre, J.; Lämmel, R.; Schmorleiz, T.; Varanovich, A. 101companies: A Community Project on Software Technologies and Software Languages. In Proceedings of the Proc. TOOLS. Springer, 2012, Vol. 7304, LNCS, pp. 58–74. Gruber, T.R. A Translation Approach to Portable Ontology Specifications. Knowledge Acquisition 1993, 5, 199–220. https://doi.org/10.1006/knac.1993.1008. Uschold, M.; Grüninger, M. Ontologies: Principles, Methods and Applications. The Knowledge Engineering Review 1996, 11, 93–136. https://doi.org/10.1017/S0269888900007797. Fernández-López, M.; Gómez-Pérez, A.; Juristo, N. METHONTOLOGY: From Ontological Art Towards Ontological Engineering. In Proceedings of the Proceedings of the AAAI Spring Symposium on Ontological Engineering, 1997, pp. 33–40. https://aaai.org/papers/0005-ss9706-005-methontology-from-ontological-art-towards-ontological-engineering/. Gómez-Pérez, A. Ontological Engineering: A State of the Art. Expert Update: Knowledge Based Systems and Applied Artificial Intelligence 1999, 2, 33–43. https://oa.upm.es/6493/1/Ontological_ Engineering_A_st.pdf. Staab, S.; Studer, R.; Schnurr, H.P.; Sure, Y. Knowledge Processes and Ontologies. IEEE Intelligent Systems 2001, 16, 26–34. https://doi.org/10.1109/5254.912382. Noy, N.F.; McGuinness, D.L. Ontology Development 101: A Guide to Creating Your First Ontology, 2001. Stanford Knowledge Systems Laboratory Technical Report KSL-01-05, https: //protege.stanford.edu/publications/ontology_development/ontology101.pdf. Suárez-Figueroa, M.C.; Gómez-Pérez, A.; Fernández-López, M. The NeOn Methodology Framework: A Scenario-Based Methodology for Ontology Development. Applied Ontology 2015, 10, 107–145. https://doi.org/10.3233/AO-150145. Tudorache, T. Ontology Engineering: Current State, Challenges, and Future Directions. Semantic Web 2020, 11, 125–138. https://doi.org/10.3233/SW-190382. Poveda-Villalón, M.; Fernández-Izquierdo, A.; Fernández-López, M.; García-Castro, R. LOT: An Industrial Oriented Ontology Engineering Framework. Engineering Applications of Artificial Intelligence 2022, 111, 104755. https://doi.org/10.1016/j.engappai.2022.104755. Spoladore, D.; Pessot, E.; Trombetta, A.; Comai, S.; Matteucci, M. A Novel Agile Ontology Engineering Methodology for Supporting Organizations in Collaborative Ontology Development. Computers in Industry 2023, 151, 103979. https://doi.org/10.1016/j.compind.2023.103979. Pernisch, R.; Poveda-Villalón, M.; Conde-Herreros, D.; Chaves-Fraga, D.; Stork, L. When Ontologies Met Knowledge Graphs: Tale of a Methodology. In The Semantic Web: ESWC 2024 Satellite Events; Springer, 2025; Vol. 15344, pp. 286–290. https://doi.org/10.1007/978-3-031-78 952-6_43. Babaei Giglou, H.; D’Souza, J.; Auer, S. LLMs4OL: Large Language Models for Ontology Learning. In The Semantic Web – ISWC 2023; Springer, 2023; pp. 408–427. https://doi.org/10.1 007/978-3-031-47240-4_22. Shimizu, C.; Hitzler, P. Accelerating Knowledge Graph and Ontology Engineering with Large Language Models. Journal of Web Semantics 2025, 85, 100862. https://doi.org/10.1016/j.websem. 2025.100862. Garijo, D.; Poveda-Villalón, M.; Amador-Domínguez, E.; Wang, Z.; García-Castro, R.; Corcho, O. LLMs for Ontology Engineering: A Landscape of Tasks and Benchmarking Challenges. In Proceedings of the ISWC 2024 Special Session on Harmonising Generative AI and Semantic Web Technologies, 2025, Vol. 3953, CEUR Workshop Proceedings. https://ceur-ws.org/Vol-3953 /364.pdf. Val-Calvo, M.; Egaña Aranguren, M.; Mulero-Hernández, J.; Almagro-Hernández, G.; Deshmukh, P.; Bernabé-Díaz, J.A.; Espinoza-Arias, P.; Sánchez-Fernández, J.L.; Mueller, J.; FernándezBreis, J.T. OntoGenix: Leveraging Large Language Models for Enhanced Ontology En-
36 of 40
34.
35.
36.
37. 38.
39.
40.
41.
42.
43.
44. 45. 46.
47. 48. 49. 50. 51. 52. 53. 54. 55. 56. 57.
gineering from Datasets. Information Processing & Management 2025, 62, 104042. https: //doi.org/10.1016/j.ipm.2024.104042. Li, J.; Poveda-Villalón, M.; Garijo, D. Large Language Models for Ontology Engineering: A Systematic Literature Review. Semantic Web 2025. Early access / accepted manuscript, https://www.semantic-web-journal.net/system/files/swj4001.pdf. Kampars, J.; Mosans, G.; Jogi, T.; Roters, F.; Vajragupta, N. LLM-Supported Collaborative Ontology Design for Data and Knowledge Management Platforms. Frontiers in Big Data 2025, 8, 1676477. https://doi.org/10.3389/fdata.2025.1676477. Gangemi, A.; Catenacci, C.; Ciaramita, M.; Lehmann, J. Modelling Ontology Evaluation and Validation. In The Semantic Web: Research and Applications; Springer, 2006; Vol. 4011, pp. 140–154. https://doi.org/10.1007/11762256_13. Wilson, R.S.I.; Goonetillake, J.S.; Indika, W.A.; Ginige, A. A Conceptual Model for Ontology Quality Assessment. Semantic Web 2023, 14, 1051–1097. https://doi.org/10.3233/SW-233393. Hammouda, N.; Mahfoudh, M.; Boukadi, K. MoOnEv: Modular Ontology Evaluation and Validation Tool. Procedia Computer Science 2024, 246, 3532–3541. https://doi.org/10.1016/j. procs.2024.09.203. Ferranti, N.; De Souza, J.F.; Ahmetaj, S.; Polleres, A. Formalizing and Validating Wikidata’s Property Constraints Using SHACL and SPARQL. Semantic Web 2024, 15, 2333–2380. https: //doi.org/10.3233/SW-243611. Robaldo, L.; Batsakis, S. On the Interplay Between Validation and Inference in SHACL: An Investigation on the Time Ontology. Semantic Web 2024, 15, 567–599. https://doi.org/10.3233/ SW-240030. Tüfek Özkaya, N.; Thuluva, A.S.; Bandyopadhyay, T.; Just, V.P.; Sabou, M.; Ekaputra, F.J.; Hanbury, A. Validating Semantic Artifacts With Large Language Models. In The Semantic Web: ESWC 2024 Satellite Events; Springer, 2025; Vol. 15344, pp. 92–101. https://doi.org/10.1007/97 8-3-031-78952-6_9. Yu, J.; Thom, J.A.; Tam, A. Ontology Evaluation Using Wikipedia Categories for Browsing. In Proceedings of the Proceedings of the Sixteenth ACM Conference on Information and Knowledge Management, 2007, pp. 223–232. https://doi.org/10.1145/1321440.1321474. Partridge, C.; Mitchell, A.; West, M. A survey of Top-Level Ontologies: To inform the ontological choices for a Foundation Data Model. Technical report, constructioninnovationhub.org.uk, 2020. https://doi.org/10.17863/CAM.58311. Partridge, C. Business Objects: Re-Engineering for Reuse; Butterworth-Heinemann, 1996. Gangemi, A.; Presutti, V., Ontology Design Patterns. In Handbook on Ontologies; Springer, 2009; pp. 221–243. https://doi.org/10.1007/978-3-540-92673-3_10. Almeida, J.P.A.; Guizzardi, G.; Sales, T.P.; Fonseca, C.M. gUFO: A Gentle Foundational Ontology for Semantic Web Knowledge Graphs, 2026, [arXiv:cs.AI/2603.20948]. https://arxiv.org/abs/ 2603.20948. Stoy, J.E. Denotational Semantics: The Scott-Strachey Approach to Programming Language Semantics; MIT Press, 1977. Mosses, P.D. Action Semantics; Cambridge University Press, 1992. Gunter, C. Semantics of Programming Languages: Structures and Techniques; MIT Press, 1992. Tennent, R.D. Denotational semantics. In Handbook of logic in computer science; Oxford University Press, 1994; Vol. 3, pp. 169—-322. Slonneger, K.; Kurtz, B. Formal Syntax and Semantics of Programming Languages; Addison Wesley, 1995. Sethi, R. Programming Languages: Concepts and Constructs; Addison Wesley, 1996. 2nd edition. Strachey, C. Fundamental Concepts in Programming Languages. Higher Order Symbol. Comput. 2000, 13, 11–49. Pierce, B. Types and Programming Languages; MIT Press, 2002. Pierce, B. Advanced Topics in Types and Programming Languages; MIT Press, 2004. Nielson, F.; Nielson, H.R.; Hankin, C. Principles of Program Analysis, corrected 2nd printing ed.; Springer, 2004. Krishnamurthi, S. Programming Languages: Application and Interpretation; Brown University, 2007. https://cs.brown.edu/~sk/Publications/Books/ProgLangs/.
37 of 40
58. 59. 60. 61. 62. 63. 64. 65. 66. 67. 68.
69.
70. 71. 72.
73. 74. 75. 76. 77.
78. 79.
80. 81.
82. 83. 84.
Friedman, D.; Wand, M. Essentials of Programming Languages; MIT Press, 2008. 3rd edition. Scott, M. Programming Language Pragmatics; Morgan Kaufmann, 1996. 3rd edition. Sestoft, P. Programming Language Concepts; Springer, 2012. Sebesta, R.W. Concepts of Programming Languages; Addison-Wesley, 2012. 10th edition. Czarnecki, K.; Eisenecker, U. Generative Programming: Methods, Tools, and Applications; AddisonWesley Professional, 2000. Wile, D.S. Lessons Learned from Real DSL Experiments. In Proceedings of the Proc. HICSS-36. IEEE, 2003, p. 325. Wile, D.S. Lessons learned from real DSL experiments. Sci. Comput. Program. 2004, 51, 265–290. Mernik, M.; Heering, J.; Sloane, A.M. When and how to develop domain-specific languages. ACM Comput. Surv. 2005, 37, 316–344. Gray, J.; Fisher, K.; Consel, C.; Karsai, G.; Mernik, M.; Tolvanen, J. DSLs: the good, the bad, and the ugly. In Proceedings of the Companion OOPSLA 2008. ACM, 2008, pp. 791–794. Ceh, I.; Crepinsek, M.; Kosar, T.; Mernik, M. Ontology driven development of domain-specific languages. Comput. Sci. Inf. Syst. 2011, 8, 317–342. Voelter, M.; Benz, S.; Dietrich, C.; Engelmann, B.; Helander, M.; Kats, L.C.L.; Visser, E.; Wachsmuth, G. DSL Engineering – Designing, Implementing and Using Domain-Specific Languages; dslbook.org, 2013. Bryant, B.R.; Jézéquel, J.; Lämmel, R.; Mernik, M.; Schindler, M.; Steinmann, F.; Tolvanen, J.; Vallecillo, A.; Völter, M. Globalized Domain Specific Language Engineering. In Proceedings of the Globalizing Domain-Specific Languages – International Dagstuhl Seminar, Dagstuhl Castle, Germany, October 5–10, 2014 Revised Papers. Springer, 2015, Vol. 9400, LNCS, pp. 43–69. Schauss, S.; Lämmel, R.; Härtel, J.; Heinz, M.; Klein, K.; Härtel, L.; Berger, T. A Chrestomathy of DSL Implementations. In Proceedings of the Proc. SLE. ACM, 2017. 12 pages. Wasowski, ˛ A.; Berger, T. Domain-Specific Languages: Effective Modeling, Automation, and Reuse; Springer, 2023. https://doi.org/10.1007/978-3-031-23669-3. Lämmel, R.; Leinberger, M.; Schmorleiz, T.; Varanovich, A. Comparison of feature implementations across languages, technologies, and styles. In Proceedings of the Proc. CSMR-WCRE. IEEE, 2014, pp. 333–337. Marriott, K.; Meyer, B. On the Classification of Visual Languages by Grammar Hierarchies. J. Vis. Lang. Comput. 1997, 8, 375–402. http://dx.doi.org/10.1006/jvlc.1997.0053. Mens, T.; Gorp, P.V. A Taxonomy of Model Transformation. ENTCS 2006, 152, 125–142. Czarnecki, K.; Helsen, S. Feature-based survey of model transformation approaches. IBM Syst. J. 2006, 45, 621–646. Tamura, G.; Cleve, A. A Comparison of Taxonomies for Model Transformation Languages. Paradigma 2010, 4, 1–14. Gomes, C.; Barroca, B.; Amaral, V. Classification of Model Transformation Tools: Pattern Matching Techniques. In Proceedings of the Proc. MODELS. Springer, 2014, Vol. 8767, LNCS, pp. 619–635. Skalna, I.; Gawel, B. Model Driven Architecture and classification of business rules modelling languages. In Proceedings of the Proc. FedCSIS, 2012, pp. 949–952. Medvidovic, N.; Taylor, R.N. A Classification and Comparison Framework for Software Architecture Description Languages. IEEE Trans. Softw. Eng. 2000, 26, 70–93. http://doi. ieeecomputersociety.org/10.1109/32.825767. Kosar, T.; Bohra, S.; Mernik, M. Domain-Specific Languages: A Systematic Mapping Study. Inf. Softw. Technol. 2016, 71, 77–91. Lämmel, R.; Mosen, D.; Varanovich, A. Method and Tool Support for Classifying Software Languages with Wikipedia. In Proceedings of the Proc. SLE. Springer, 2013, Vol. 8225, LNCS, pp. 249–259. Cardelli, L.; Wegner, P. On Understanding Types, Data Abstraction, and Polymorphism. ACM Comput. Surv. 1985, 17, 471–522. Chikofsky, E.J.; II, J.H.C. Reverse Engineering and Design Recovery: A Taxonomy. IEEE Softw. 1990, 7, 13–17. Favre, J.M.; Nguyen, T. Towards a Megamodel to Model Software Evolution Through Transformations. In Proceedings of the Proceedings of the International Workshop on Software Evolution
38 of 40
85.
86. 87.
88.
89. 90.
91. 92.
93.
94. 95.
96.
97.
98.
99.
through Transformations (SETra), 2004. Introduces megamodeling, a framework for modeling languages, artifacts, and transformations; provides the structural basis for representing technological spaces and their interconnections., https://www.researchgate.net/publication/220369 017_Towards_a_Megamodel_to_Model_Software_Evolution_Through_Transformations. Favre, J.M.; Lämmel, R.; Varanovich, A. Modeling the Linguistic Architecture of Software Products. In Proceedings of the MODELS 2012. Springer, 2012, Lecture Notes in Computer Science, pp. 151–167. Models the linguistic architecture of software products, strengthening the megamodeling perspective for relating languages, artifacts, and technological spaces., https://doi.org/10.1007/978-3-642-33666-9_11. Lämmel, R.; Varanovich, A. Interpretation of Linguistic Architecture. In Proceedings of the Proc. ECMFA. Springer, 2014, Vol. 8569, LNCS, pp. 67–82. Härtel, J.; Härtel, L.; Heinz, M.; Lämmel, R.; Varanovich, A. Interconnected Linguistic Architecture. The Art, Science, and Engineering of Programming Journal 2017, 1. 27 pages. Available at http://programming-journal.org/2017/1/3/. Lämmel, R. Megamodels on the Catwalk. In Proceedings of the Proceedings of the 9th International Conference on Model-Driven Engineering and Software Development, MODELSWARD 2021, Online Streaming, February 8-10, 2021; Hammoudi, S.; Pires, L.F.; Seidewitz, E.; Soley, R., Eds. SCITEPRESS, 2021, p. 7. Zaytsev, V.; Bagge, A.H. Parsing in a Broad Sense. In Proceedings of the Proc. MODELS. Springer, 2014, Vol. 8767, LNCS, pp. 50–67. Rocco, J.D.; Ruscio, D.D.; Härtel, J.; Iovino, L.; Lämmel, R.; Pierantonio, A. Understanding MDE projects: megamodels to the rescue for architecture recovery. Softw. Syst. Model. 2020, 19, 401–423. https://doi.org/10.1007/S10270-019-00748-7. Lämmel, R. Coupled Software Transformations Revisited. In Proceedings of the Proc. SLE. ACM, 2016, pp. 239–252. Klint, P.; Lämmel, R.; Verhoef, C. Toward an Engineering Discipline for Grammarware. ACM Transactions on Software Engineering and Methodology 2005, 14, 331–380. Establishes grammarware as a coherent engineering ecosystem, providing a concrete example of a technological space in practice., https://doi.org/10.1145/1072997.1073000. Bézivin, J. Model Driven Engineering: An Emerging Technical Space. In Generative and Transformational Techniques in Software Engineering; Springer, 2006; Vol. 4143, Lecture Notes in Computer Science, pp. 36–64. Interprets model-driven engineering as a technological space, illustrating how a space organizes languages, metamodels, and transformations., https://doi.org/10.1007/11877028_2. Bourque, P.; Dupuis, R.; Abran, A.; Moore, J.W.; Tripp, L.L. The Guide to the Software Engineering Body of Knowledge. IEEE Softw. 1999, 16, 35–44. https://doi.org/10.1109/52.805471. Dupuis, R.; Bourque, P. Guide to the Software Engineering Body of Knowledge Diffusion and Experimentation Strategy. In Proceedings of the Thirteenth Conference on Software Engineering Education and Training, 6-8 March, 2000, Austin, Texas, USA. IEEE Computer Society, 2000, pp. 49–50. https://doi.org/10.1109/CSEE.2000.827017. Bourque, P. The SWEBOK Guide - More Than 20 Years down the Road. In Proceedings of the 32nd IEEE Conference on Software Engineering Education and Training, CSEE&T 2020, Virtual Conference, Germany, November 9-12, 2020; Daun, M.; Hochmüller, E.; Krusche, S.; Brügge, B.; Tenbergen, B., Eds. IEEE, 2020, pp. 1–2. https://doi.org/10.1109/CSEET49119.2020.9206209. Burgueño, L.; Ciccozzi, F.; Famelis, M.; Kappel, G.; Lambers, L.; Mosser, S.; Paige, R.F.; Pierantonio, A.; Rensink, A.; Salay, R.; et al. Contents for a Model-Based Software Engineering Body of Knowledge. Softw. Syst. Model. 2019, 18, 3193–3205. https://doi.org/10.1007/S10270-019-007 46-9. Atkinson, C.; Kühne, T. Model-Driven Development: A Metamodeling Foundation. IEEE Software 2003, 20, 36–41. Clarifies how metamodeling underpins classification and typing in model-driven development, which is relevant to understanding conformance., https://doi.org/ 10.1109/MS.2003.1231149. Atkinson, C.; Gerbig, R. Demystifying Ontological Classification in Language Engineering. In Proceedings of the European Conference on Modelling Foundations and Applications, 2016.
39 of 40
Explains ontological classification in language engineering and therefore sheds light on how classification and typing relate to conformance., https://doi.org/10.1007/978-3-319-42061-5_6. 100. Degueule, T.; Combemale, B.; Blouin, A.; Barais, O.; Jézéquel, J.M. Safe Model Polymorphism for Flexible Modeling. Computer Languages, Systems & Structures 2017, 49, 176–195. Discusses limitations of standard conformance and proposes safer forms of polymorphism, directly informing the modeling of conformance., https://doi.org/10.1016/j.cl.2016.09.001. 101. Rossini, A.; de Lara, J.; Guerra, E.; et al. A formalisation of deep metamodelling. Formal Aspects of Computing 2014, 26, 1115–1152. Provides a formal account of deep metamodelling, helping distinguish multi-level classification and typing notions related to conformance., https: //doi.org/10.1007/s00165-014-0307-x. 102. Favre, J.M. Foundations of Meta-Pyramids: Languages vs. Metamodels – Episode II: Story of Thotus the Baboon. In Proceedings of the Language Engineering for Model-Driven Software Development. Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 2005, Vol. 4101, Dagstuhl Seminar Proceedings (DagSemProc), pp. 1–28. Contrasts languages and metamodels in a way that sharpens distinctions among conformance, classification, and typing., https://doi.org/10.4230/ DagSemProc.04101.7. 103. Lämmel, R.; Verhoef, C. Semi-automatic grammar recovery. Softw. Pract. Exp. 2001, 31, 1395– 1438. https://doi.org/10.1002/SPE.423. 104. Lämmel, R.; Verhoef, C. Cracking the 500-Language Problem. IEEE Softw. 2001, 18, 78–88. https://doi.org/10.1109/52.965809. 105. Lämmel, R. Grammar Adaptation. In Proceedings of the Proc. FM. Springer, 2001, Vol. 2021, LNCS, pp. 550–570. 106. Klint, P.; Lämmel, R.; Verhoef, C. Toward an engineering discipline for grammarware. ACM Trans. Softw. Eng. Methodol. 2005, 14, 331–380. 107. Jackson, R.C.; Balhoff, J.P.; Douglass, E.; Harris, N.L.; Mungall, C.J.; Overton, J.A. ROBOT: A Tool for Automating Ontology Workflows. BMC Bioinformatics 2019, 20, 407. https://doi.org/ 10.1186/s12859-019-3002-3. 108. Matentzoglu, N.; Goutte-Gattat, D.; Tan, S.Z.K.; Balhoff, J.P.; Carbon, S.; Caron, A.R.; Duncan, W.D.; Flack, J.E.; Haendel, M.; Harris, N.L.; et al. Ontology Development Kit: a toolkit for building, maintaining and standardizing biomedical ontologies. Database 2022, 2022, baac087. https://doi.org/10.1093/database/baac087. 109. Dziwis, G.; Wenige, L.; Meyer, L.P.; Martin, M. OntoFlow: A user-friendly Ontology Development Workflow. In Proceedings of the Proceedings of the International Workshop on Semantic Industrial Information Modelling (SemIIM) @ ESWC 2022, 2022, Vol. 3355, CEUR Workshop Proceedings. https://ceur-ws.org/Vol-3355/ontoflow.pdf. 110. Alobaid, A.; Garijo, D.; Poveda-Villalón, M.; Santana-Perez, I.; Fernández-Izquierdo, A.; Corcho, O. Automating ontology engineering support activities with OnToology. Journal of Web Semantics 2019, 57, 100472. https://doi.org/10.1016/j.websem.2018.09.003. 111. Publio, G.C.; Labra Gayo, J.E.; Colunga, G.F.; Menéndez, P. Ontolo-CI: Continuous Data Validation With ShEx. In Proceedings of the Proceedings of the Poster and Demo Track of the 18th International Conference on Semantic Systems (SEMANTiCS 2022), 2022, Vol. 3235, CEUR Workshop Proceedings. https://ceur-ws.org/Vol-3235/paper6.pdf. 112. Hannou, F.Z.; Charpenay, V.; Lefrançois, M.; Roussey, C.; Zimmermann, A.; Gandon, F. The ACIMOV Methodology: Agile and Continuous Integration for Modular Ontologies and Vocabularies. In Proceedings of the Joint Proceedings of Workshops at JOWO 2023, 2023. https://www.emse.fr/~zimmermann/Papers/mk2023.pdf. 113. Lefrançois, M.; Gnabasik, D. The SAREF Pipeline and Portal—An Ontology Verification Framework. In Proceedings of the The Semantic Web – ISWC 2023. Springer, 2023, pp. 134–151. https://doi.org/10.1007/978-3-031-47243-5_8. 114. Hofer, M.; Hellmann, S.; Dojchinovski, M.; Frey, J. The New DBpedia Release Cycle: Increasing Agility and Efficiency in Knowledge Extraction Workflows. In Proceedings of the Semantic Systems. In the Era of Knowledge Graphs. Springer, 2020, pp. 1–18. https://doi.org/10.1007/ 978-3-030-59833-4_1.
40 of 40
115. Meissner, R.; Junghanns, K. Using DevOps Principles to Continuously Monitor RDF Data Quality. In Proceedings of the Proceedings of the 12th International Conference on Semantic Systems, 2016. https://doi.org/10.1145/2993318.2993351. 116. Favre, J.; Lämmel, R.; Leinberger, M.; Schmorleiz, T.; Varanovich, A. Linking Documentation and Source Code in a Software Chrestomathy. In Proceedings of the Proc. WCRE. IEEE, 2012, pp. 335–344. 117. Lämmel, R.; Schmorleiz, T.; Varanovich, A. The 101haskell Chrestomathy: A Whole Bunch of Learnable Lambdas. In Proceedings of the Proc. IFL. ACM, 2013, p. 25. 118. Lämmel, R. Software chrestomathies. Sci. Comput. Program. 2015, 97, 98–104. 119. Lämmel, R. Relationship Maintenance in Software Language Repositories. The Art, Science, and Engineering of Programming Journal 2017, 1. 27 pages. Available at http://programming-journal. org/2017/1/4/.