Solution of the Hempel’s statistical ambiguity problem and Causal AI Evgenii Vityaev1* 1*
Sobolev Institute of Mathematics of the SB RAS, Koptuga 4, Novosibirsk, 630090, Russia.
Corresponding author(s). E-mail(s): [email protected];
arXiv:2607.12826v1 [cs.AI] 14 Jul 2026
Abstract This paper addresses Carl Hempel’s longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are derived from statistical laws. To avoid such predictions, Carl Hempel proposed the Requirement of Maximal Specificity (RMS) for the statistical laws used in the inference. An analysis of the RMS refinements made by Wesley Salmon, Alberto Coffa, and James Fetzer led to the following definition of maximally specific statistical laws: ”the lawlike premises of an adequate explanation must specify all and only those properties whose presence or absence made a difference to the occurrence of its explanandum-phenomenon.” However, there was no proof of a solution to the statistical ambiguity problem based on this definition. We use Nancy Cartwright’s definition of causes that raise probabilities across background contexts, and then introduce the concept of Causal Rules. Then we define a special semantic probabilistic inference procedure that incrementally refines these causal rules by incorporating all statistically relevant information. This procedure yields Maximally Specific Causal Relationships (MSCRs), for which we prove (Theorem 1) that predictions derived from them are consistent. This resolves the statistical ambiguity problem. The semantic probabilistic inference procedure provides a probabilistic causal learning system, which may be used in such new areas as Causal AI and Causal Machine Learning. They fundamentally explore causal inference as a tool for understanding cause-and-effect relationships within complex systems. Properties similar to RMS remain under discussion. Several notions related to RMS are considered: invariant feature learning, invariant causal prediction, and spurious association. Keywords: explanation, inductive-statistical inference, statistical ambiguity, causal inference, consistency
1
1 Introduction In recent years, within the fields of Causal AI and Causal Machine Learning, causal inference has become a critical tool for understanding cause-and-effect relationships. Integrating these relationships into machine learning models enables the creation of causal models for a deeper understanding of real-world systems. This paradigm provides more robust, interpretable, and actionable insights in areas such as medicine, finance, and autonomous systems. To develop the probabilistic causal learning system that infers effects without contradictions we need to solve Hempel’s statistical ambiguity problem. It concerns the problem that contradictory predictions derived from inductive-statistical inference Hempel (1965, 1968). To avoid such contradictions, Carl Hempel introduced the Requirement of Maximal Specificity (RMS) for statistical laws, which means that maximally specific statistical laws must incorporate all information relevant to prediction. Subsequent analysis of RMS revealed numerous problems and disputes (see next section on Historical background). As a result, a solution to the ”statistical ambiguity” problem was not reached, and the consistency of predictions for maximally specific statistical laws was not proven, and the problem remained unsolved. This article resumes the consideration of the ”statistical ambiguity problem” and RMS. We define RMS based on Hempel’s definition and define the notion of maximally specific cause-effect relationships (MSCR). Then we prove a theorem (Theorem 1) that predictions based on MSCRs are consistent. This solves Hempel’s ”statistical ambiguity problem” for maximally specific cause-effect relationships. For the works in Causal AI and Causal Machine Learning it is rather important. Based on Cartwright’s definition of causes that raise probabilities of their effects in various background contexts Cartwright (1979); Stanford (2008) we define a stronger notion of Causal Rules (CR). Cartwright: C causes E if and only if P (E | C &B ) > P (E | ¬C &B ) for every background context B. The background in the Causal Rules consists of all other conditions of the rule. Causal Rule (see Definition 6): The rule A1 & . . . &Ak ⇒ A0 , is a causal rule iff every literal A1 , . . . , Ak is a cause of A0 P (A0 | A1 & . . . &Ak ) > P (A0 | ¬Ai &B ), i = 1, . . . , k , relative to background B = &{{A1 , . . . , Ak } \ Ai }, containing all literals except Ai . We introduce a refinement procedure for causal rules that incrementally strengthens the conditional probabilities of causal rules by adding all relevant information until the causal rules can no longer be refined, thus producing Maximally Specific Causal Relationships. It provides the learning method (Definition 8) for MSCR discovery. MSCRs may be considered as causal models, as they identify the key variables that underpin reliable cause-and-effect relationships. This method thus produces a causal learning system. Let MSR is the set of all MSCRs. It is proved in Theorem 1 that I-S inference of predictions based on any subset of MSR is consistent. The set MSR thus provides 2
a consistent set of probabilistic knowledge. The set MSR may be considered as a consistent probabilistic theory in the same sense as a logical theory that is consistent and infers consistent conclusions. The problem of statistical ambiguity for statistical laws of the form A1 & . . . &Ak ⇒ A0 , where A1 , . . . , Ak , A0 are literals was considered in Vityaev (2006); Vityaev and Martinovich (2015), and for the form ϕ ⇒ ψ , where ϕ and ψ are propositional formulas was considered in Vityaev and Odintsov (2019). Our paper is organized as follows. Section 2 (RMS history) is devoted to the historical analysis of RMS and reveals discussions around it. Section 3 considers relation of RMS to Causal AI and Causal Machine Learning. Section 4 is devoted to definition of probability on propositional formulas. Section 5 considers rules, their probabilities and the definition of MSCR. In Section 6, we define I-S inference as a prediction operator for MSR and prove that applying it to a consistent set of MSCRs produces a consistent set. Finally, Section 7 is conclusion.
2 RMS history The problem of statistical ambiguity in inductive-statistical (I-S) explanation was first identified by Hempel. It generated a substantial literature attempting to articulate and refine the conditions under which probabilistic explanations can avoid contradictory conclusions. This section traces the evolution of the Requirement of Maximal Specificity (RMS) from its origins in Hempel’s work through its subsequent critiques and reformulations, culminating in the identification of causal factors as the essential element for resolving the ambiguity. The Principle of Total Evidence and Its Limitations. Before addressing the Requirement of Maximal Specificity directly, it is essential to distinguish it from another fundamental principle of probabilistic inference: the requirement of total evidence. As articulated by Carnap and others, the requirement of total evidence states that in any rational application of probabilistic inference, the probability assigned to a hypothesis must be determined by reference to all available evidence. In Hempel’s formulation, this principle directs that when assessing the credibility of a statement, one must consider the entire body of knowledge K: “What degree of belief, or what probability, is it rational to assign to the statement ’Gi’ in a given knowledge situation?” Hempel (1968). Initially, Hempel had conflated the requirement of maximal specificity with the requirement of total evidence, suggesting that RMS served as a ”rough substitute” for the latter. However, in his 1968 reappraisal, he explicitly retracted this view, recognizing that the two principles address fundamentally different questions. The requirement of total evidence concerns the rational credibility of a statement based on all available information. The Requirement of Maximal Specificity, by contrast, addresses the explanatory status of an argument: it specifies conditions under which two sentences in K — a statistical law p(G, F ) = r and a singular sentence Fi — can serve to explain, relative to K , why i is G. As Hempel emphasizes: The point of an explanation is not to provide evidence for the occurrence of the explanandum phenomenon, but to exhibit it as nomically expectable. And the probability attached to an I − S explanation is
3
the probability of the conclusion relative to the explanatory premises, not relative to the total class K . Thus, the requirement of total evidence simply does not apply to the determination of the probability associated with an I − S explanation, and the requirement of maximal specificity is not ”a rough substitute for the requirement of total evidence.” Hempel (1968) This clarification established RMS as an independent principle with its own distinct rationale, specifically designed to address the problem of explanatory ambiguity in statistical contexts. The Problem of Statistical Ambiguity and the Initial Formulation of RMS. The problem that RMS was designed to solve emerges from the nature of inductive-statistical explanation itself. As Hempel (1965) demonstrated, an I − S argument has the form:
p(G; F ) = r F (a) G(a) where line indicates an inductive relationship, and r represents the inductive probability of the explanandum given the explanans. Statistical ambiguity arises when two arguments with true premises yield contradictory conclusions. It can be illustrates with a classic example (cited Salmon). Suppose that we have the following statements.
• L1 Almost all cases of streptococcus infection clear up quickly after the administration of penicillin. • L2 Almost no cases of penicillin resistant streptococcus infection clear up quickly after the administration of penicillin. • C1 Jane Jones had streptococcus infection. • C2 Jane Jones received treatment with penicillin. • C3 Jane Jones had a penicillin resistant streptococcus infection. Following the above pattern of I-S explanation it is possible to construct two contradictory arguments based on these statements. On the base of L1 and C 1 ∧ C 2 one can explain why Jane Jones recovered quickly (E). The second argument with premises L2 and C 2 ∧ C 3 explains why Jane Jones did not (¬E ). The premises of both arguments are consistent with each other, they could all be true. The probability of both argument may be close to 1. To block such conflicting explanations, Hempel introduced the Requirement of Maximal Specificity Hempel (1965). RMS: For an I − S argument to be acceptable relative to a knowledge state K, for any predicate H such that K contains both ∀x(H(x) ⇒ F (x)) and H(a), there must exist a statistical law p(G; H) = r in K with the same probability r.
The basic idea is that if H provides more specific information about the object than F , then the law based on H should be preferred — and if that law has a different probability, the original argument is invalid. Hempel hoped to solve this problem by forcing all statistical laws in an argument to be maximally specific. That is, they should contain all relevant information with 4
respect to the domain in question. In our example, then, the premise C 3 invalidates the first argument, since it is not maximally specific with respect to all information about Jane Jones. So, we can only explain ¬E, but not E. The Problem of Epistemic Relativity. Coffa criticizes Hempel’s insistence on relativizing I − S explanation to a knowledge state K Coffa (1974). Coffa begins by distinguishing between epistemic and non-epistemic concepts. A concept is epistemic if its meaning cannot be given without reference to knowledge; non-epistemic concepts—such as “table,” “chair,” or “truth” — can be characterized independently of any knower. Hempel’s thesis, Coffa explains, is that inductive explanation is not merely epistemic in the sense that it depends on what we know about the evidence (as in the case of “well-confirmed D-N explanation”), but rather that it is a “non-conformational epistemic concept”. This means, as Hempel himself acknowledges, that “there is no concept that stands to his epistemically relativized notion of inductive explanation as the concept of true D-N explanations stands to that of well-confirmed D-N explanation” Coffa (1974). In other words: According to the thesis of epistemic relativity there is no meaningful notion of true inductive explanation. Hence, we could not possibly have reasons to believe that anything is a “true inductive explanation”. Coffa (1974). Coffa’s fundamental objection is that this conclusion is not merely surprising but fatal to the project of inductive explanation. If there is no notion of true inductive explanation, then I − S explanations relativized to K cannot be understood as arguments that we have reason to believe are explanations in the same sense that D − N explanations are. More importantly, Coffa argues that Hempel has misidentified the nature of the problem. The real difficulty is not epistemic but ontological — it is the old problem of the reference class in a new guise. We would like to suggest that when Hempel turned his attention to the theory of inductive explanation what he stumbled upon was the fact that the problem of defining a model of inductive explanation for single events was the other side of the coin of the single case problem. He stumbled, that is, upon the reference-class problem Coffa (1974). The reference-class problem arises because a single event can belong to multiple reference classes with different probabilities for the outcome of interest. Jones belongs to the class of persons with streptococcus infection, the class of persons treated with penicillin, and the class of persons with penicillin-resistant infections. Each yields a different probability for recovery. The frequentist tradition, Coffa notes, held that it is ”strictly meaningless to assign a probability to a single event” Coffa (1974) because all reference classes are, in principle, equally legitimate. Hempel’s epistemic relativization attempts to solve this problem by appeal to knowledge: the appropriate reference class is the most specific class to which the individual is known to belong. But Coffa observes that this strategy has the ironic consequence that “it is ignorance, rather than knowledge, that makes the maximal specificity principle look like a workable demand” Coffa (1974). As knowledge increases, the principle becomes increasingly difficult to satisfy; an omniscient being would find no inductive explanations at all.
5
Coffa’s proposed alternative points toward an ontic rather than epistemic formulation: the requirement should refer to “all relevant aspects of the explanandum” Coffa (1974), where relevance is understood not as statistical correlation but as nomic connection — “a predicate being nomically relevant to another when a law of nature determines that changes in the first one generate changes in the second one” Coffa (1974). This suggestion anticipates later developments in causal approaches to explanation. Salmon’s Statistical Relevance Approach and shift to Ontic Homogeneity. A different line of critique emerged from Wesley Salmon, who questioned not merely the formulation of RMS but its fundamental orientation. Salmon argued that Hempel’s requirement of high probability is misguided; what matters for explanation is not the magnitude of probability but the presence of statistical relevance relations. The following example illustrates this point Salmon (1970):
John Jones was almost certain to recover from his cold within a week, because he took vitamin C, and almost all colds clear up within a week after administration of vitamin C. The difficulty with this example is that colds tend to clear up within a week regardless of the medication administered, and controlled tests indicate that the percentage of recoveries is unaffected by the use of vitamin C. Thus, even though the argument satisfies Hempel’s requirements — a high probability, a statistical law, and true premises — it fails to be explanatory because the putative explanatory factor (taking vitamin C) is irrelevant to the outcome. What is needed, Salmon argues, is not high probability but relevance — the fact that taking vitamin C makes a difference to the probability of recovery. This led Salmon to develop the Statistical-Relevance (S-R) model, in which an explanation consists of partitioning a reference class into cells that are homogeneous with respect to the explanandum property, together with information about which cell contains the individual in question. Crucially, Salmon’s homogeneity requirement is objective rather than epistemic: a class is homogeneous if no further partition can be made that is relevant to the occurrence of the explanandum property, regardless of whether such partitions are known. This represents a fundamental shift from Hempel’s epistemic relativization to an ontic conception of explanation. As Salmon later put it, “the identification of the appropriate explanans is fully objective” once the explanandum has been unambiguously specified (Salmon, 1989). Fetzer’s Requirement of Strict Maximal Specificity and the Turn to Causation. James Fetzer (Fetzer (1981, 1993)) carries the ontic turn further, arguing that neither Hempel’s epistemic RMS nor Salmon’s statistical relevance conditions are sufficient. What is required is not merely statistical homogeneity but causal relevance. Fetzer introduces the Requirement of Strict Maximal Specificity (RSMS), which demands “that the lawlike premises of an adequate explanation must specify all and only those properties whose presence or absence made a difference to the occurrence of its explanandum-phenomenon” Fetzer (1993). Fetzer’s analysis reveals that the problem of statistical ambiguity cannot be resolved at the level of statistical relations alone. The fundamental question is not
6
which reference class yields the highest probability or even which partition yields homogeneous cells, but rather: which factors are genuinely responsible for the outcome? This is an ontological question about causal structure, not an epistemological question about the choice of reference class. Conclusion. The evolution of the requirements for maximum specificity reveals the depth of the Hempel statistical ambiguity problem. Salmon and Fetzer came close enough to solving the problem, but did not solve it. Their research is taken into account in our definition of ”causal rule”. Salmon’s statistical significance is taken into account in the definition of ”causal rules” as Cartwright’s definition of cause. James Fetzer’s requirement of strict maximal specificity is taken into account by determining the statistical significance of each condition in the premise of the rule. Fetzer’s requirement that ”adequate explanations apply to all and only those properties whose presence influenced the explanation” is taken into account in the procedure for clarifying causal rules, which gradually strengthens the conditional probabilities of causal rules by adding all relevant information until the causal rules can no longer be clarified. This procedure allows us to obtain the Maximally Specific Causal Relationships that solve the problem of statistical ambiguity (Theorem 1).
3 Causal AI and Causal Machine Learning In recent years, causal inference has emerged as a critical tool for understanding causeand-effect relationships within complex systems. By incorporating causal reasoning into machine learning, models can move to deeper understanding of the real-world systems. The main theoretical achievement of MSCRs is the formal proof of consistency for predictions. This approach shares the ambition of contemporary causal approaches to identify genuinely operative causes. Properties similar to RMS remain under discussion Kaddour et al. (2022). There are some notions related to RMS in the Causal Machine Learning Kaddour et al. (2022):
• Invariant Feature Learning. The task of identifying features of our data X that are predictive of Y across a range of environments E. This definition is very similar to the N. Cartwright definition of causes that raise probabilities of their effects in various background contexts. It is included in our definition of causal rules. • Invariant Causal Prediction – an algorithm to find the causal feature set, the minimal set of features which are causal predictors of a target variable Kaddour et al. (2022). In our definition of causal rule this property is also fulfilled. • Spurious association. In the training dataset, pictures of cows typically exhibit alpine pasture backgrounds, a spurious association caused by the cow’s natural habitat. Definitions of cause by N. Cartwright and a causal rule exclude these spurious associations. Thus, the introduced notion of “causal rule” remains an orientation for other Causal AI and Causal Machine Learning methods to extract precise knowledge from data predicting without contradictions.
7
4 Logic and Probability background Let F (At) be a set of well-formed formulas constructed from a set of atoms At using connectives &, ∨, →, ¬. Equivalence ↔ is an abbreviation, φ ↔ ψ = (φ → ψ )&(ψ → φ). We define ⊤ as φ ∨ ¬φ, where φ is some fixed formula. V The conjunction of a finite set of formulas T is denoted by T . V (φ) denotes the set of atoms occurring in formula φ. The algebra of formulas F (At) is a set F (At) with naturally interpreted connectives. Definition 1 The models are mappings from At to truth values {0, 1}. Such mappings we call (At-)valuations. Every valuation val : At → {0, 1} extends in a standard way to the set F (At) as val : F (At) → {0, 1}.
Let G be a set of valations. A formula φ is said to satisfiable in G if val(φ) = 1 for some val ∈ G. We say that φ holds on G (and write G |= φ) if val(φ) = 1 for all val ∈ G. Finally, a set T ⊆ F (At) is said to be G-consistent if there is val ∈ G such that val(φ) = 1 for all φ ∈ T . The set of all valuations we denote A. By a logical tautology we mean a formula that holds on A. For a set G of valuations, the relation φ ≡G ψ is defined by φ ↔ ψ holding on G. Then ≡G is a congruence on F (At), the respective quotient is denoted as B G (At). The coset of φ w.r.t. ≡G is denoted as [φ]G and the universe of B G (At) equals {[φ]G | φ ∈ F (At)}. The operations of B G (At) are denoted as &G , ∨G , →G , ¬G and the lattice order as ⊑G . Recall that [φ]G ⊑G [ψ ]G iff [φ]G = [φ&ψ ]G = [φ]G &G [ψ ]G . Definition 2 Finitely additive measure µ is defined on G as µ : 2G → [0, 1]:
1. µ(G) = 1, µ(∅) = 0; 2. µ(A1 ∪. . .∪An ) = µ(A1 )+ . . . + µ(An ) for pairwise disjoint subsets A1 , . . . , An ⊆ G. 3. µ(A) = 0 iff A = ∅ Elements of G may be interpreted as outcomes of experiments. Further, we assume that only essential experiments are included in G, which explains why µ({val}) ̸= 0 for all val ∈ G. For every φ ∈ F (At), we define φG := {val ∈ G | val(φ) = 1} and function ν : F (At) → {0, 1} by the rule ν (φ) = µ(φG ). Then the function ν satisfies the following properties. Proposition 1 1. ν (φ) = 1 iff φ holds on G.
2. ν (φ) = 0 iff φ is not satisfiable on G. 3. ν (φ ∨ ψ ) = ν (φ) + ν (ψ ) iff φ&ψ is not satisfiable in G. Thus we defined the probability on the set of propositional formulas in the sense of Fagin et al. (1990).
8
5 Method. Causal Rules and Semantic Probabilistic Inference By a rule we mean a syntactic object of the form
r = φ ⇒ ψ, where φ, ψ ∈ F (At). Formula φ is called a body of the rule, whereas ψ is a head of the rule: φ = B (r) and ψ = H (r). A rule r cannot be identified with the implication φ → ψ because the probability of r will be defined in a different way. Namely, for a rule r = φ ⇒ ψ , whose body is satisfiable on G, i.e., ν (φ) ̸= 0, we put
ν (r) := ν (ψ|φ) =
ν (ψ &φ) . ν (φ)
In case B (r) is not satisfiable on G, the value ν (r) remains undefined. Notice that the value ν (r) was defined so that it is smaller than the probability of implication. Proposition 2 For every rule r with B(r)G ̸= G, we have ν(r) ≤ ν(B(r) → H(r)). Moreover, the equality ν(r) = ν(B(r) → H(r)) is equivalent to B(r)G ⊆ H(r)G , i.e., to the fact that the implication B(r) → H(r) holds on G.
This statement can be proved analogously to Theorem 2 from Vityaev and Odintsov (2019) Definition 3 Let r1 and r2 be two rules with the same head, H(r1 ) = H(r2 ). We call r1 a specification of r2 , symbolically r1 ≼ r2 , if B(r1 )G ⊆ B(r2 )G ; rule r1 is a proper specification of r2 , r1 ≺ r2 , if B(r1 )G ⫋ B(r2 )G . We say in this case that r2 is a (proper) generalization of r1 .
In other words, one of the two rules with the same head is a proper generalization of the other if its body is weaker from the logical point of view. Definition 4 Let rule r1 be a specification of r2 . We say that r1 is a refinement of r2 , symbolically r1 > r2 , if ν(r1 ) > ν(r2 ).
Evidently, the relation r1 > r2 implies that r1 is a proper specification of r2 . Definition 5 A set R of rules is said to be rich if for every r ∈ R and an arbitrary s such that s > r, there is r′ ∈ R with r′ ≼ s and ν(r′ ) > ν(r).
9
Definition 6 Let R be a rich set of rules. Rule r is a causal rule relative to R, if r ∈ R and r is a refinement of all its proper generalizations from R. We assume that if r ∈ R and (⊤ ⇒ H(r)) ≻ r then ν(r) > ν(⊤ ⇒ H(r)). Proposition 3 Let R be a rich set of rules over a finite set At of atoms. For every rule r ∈ R, there exists its generalization r′ such that r′ is a causal rule relative to R and ν(r′ ) ≥ ν(r).
Proof Let r = φ ⇒ ψ ∈ R. Consider the set ∆ = {[α]G | ν(α ⇒ ψ) ≥ ν(r), [α]G ⊑G [φ]G }. Recall that the condition [α]G ⊑G [φ]G means exactly that α ⇒ ψ ≼ r. Since ∆ is finite as a subset of BG (At), we can choose an element [β]G ∈ ∆ minimal w.r.t. ⊑G . Since R is reach and β ⇒ ψ ≼ r, there is r′ ∈ R with r′ ≼ β ⇒ ψ. Assume that r′ is not a causal rule, then there is a proper generalization γ ⇒ ψ of r′ such that ν(γ ⇒ ψ) ≥ ν(r′ ). Naturally, in this case [α]G ∈ ∆ and since a proper generalization γ ⇒ ψ is a proper generalization of r′ , we have [γ]G ⊏G [β]G , which contradicts the minimality of [β]G . Thus, r′ is the required causal rule. □ Definition 7 A causal rule r is strong relative to R, if there is no causal rule r′ such that r′ is a refinement of r. Definition 8 Semantic Probabilistic Inference (SPI) of ψ is a sequence r1 , . . . , rk of causal rules ri ∈ R with the head ψ such that:
ν (r1 ) > ν (⊤ ⇒ ψ ); ri+1 > ri , i=1,. . . ,k-1; rk – strong relative to R. Finally, among all strong causal rules with a given head we distinguish rules with maximal probability. Definition 9 Let ψ ∈ F (At). A strong causal rule r with head ψ is called a maximal specific causal rule for ψ relative to R, if its probability ν(r) is greatest among all strong causal rules with head ψ.
We will say that r is a maximal specific causal rule relative to R if r is a maximal specific causal rule relative to R for H (r). The set of all maximal specific causal rules we denote as MSCR. Rules from MSCR we consider as satisfying the Requirement of Maximal Specificity, because a specification of such rules does not lead to an increase of probability, which means that their bodies contains all statistically relevant information for the prediction of its head. Proposition 4 Let At be finite. For every rule r such that H(r) = ψ and the value ν(r) is defined, there exists a maximal specific causal rule r′ for ψ such that ν(r′ ) ≥ ν(r).
10
Proof According to Proposition 3 the set ∆ = {[α]G | α ⇒ ψ is a probabilistic causal rule} is non-empty. It follows immediately from Definition 7 that α ⇒ ψ is a strong causal rule iff [α]G is a minimal element of ∆ w.r.t. ⊑G . Since ∆ is finite, the set of its ⊑G -minimal elements is a finite non-empty set. So we can choose in this set an element [β]G with the greatest value ν(β ⇒ ψ). This is a required maximal specific causal rule r′ for ψ. That ν(β ⇒ ψ) ≥ ν(r) follows again from Proposition 3. □
6 Result Let At, a set G of models, and a measure µ on G be fixed. The set of all maximal specific causal rules relative to R in that case we denote as MSCR(R, G, µ). For the set of rules Π ⊆ MSCR(R, G, µ) and T ⊆ F (At) we define an operator of direct predictions:
P rΠ (T ) = T ∪ {H (r) | r ∈ Π, ∃φ1 , ..., φn ∈ T (G |= (φ1 & . . . &φn ) ↔ B (r))} Further, we put: n+1 0 n P rΠ (T ) = T, P rΠ (T ) = P rΠ (P rΠ (T )),
P RΠ (T ) =
[
n P rΠ (T ).
n∈ω
We call P RΠ a prediction operator for Π. Theorem 1 Let At be a finite set of atoms, Π ⊆ MSCR(R, G, µ), and T ⊆ F (At). If T is G-consistent, then P RΠ (T ) is G-consistent too. Proof Obviously, it will be enough to check that the operator of direct predictions produces a G-consistent set of formulas. First of all we show that the set T of formulas and the set Π of rules can be replaced by finite sets. Consider the family of cosets {[φ]G | φ ∈ T }, which is finite since At is finite. For every [φ]G from this family choose a single representative and put it into T ′ . Now we consider the set of pairs: Θ = {([φ]G , [ψ]G ) | φ ⇒ ψ ∈ Π.} ′ A finite set of rules Π is defined as follows. For every pair of cosets ([φ]G , [ψ]G ) ∈ Θ, we choose a single pair of representatives (φ, ψ) and put the rule φ ⇒ ψ into Π′ . It is clear that for every ψ ∈ P rΠ (T ) there is a ψ ′ in P rΠ′ (T ′ ) such that ψ ≡G ψ ′ . In this way, if P rΠ′ (T ′ ) is G-consistent, then P rΠ (T ) is G-consistent too. Further, let Π′′ be the set of such rules r from Π′ that the equivalence (φ1 & . . . &φn ) → B(r) holds on G for some φ1 ,. . . , φn ∈ T ′ . Let Π′′ = {r1 , . . . , rn }. We put T0 = T ′ , Ti+1 = Ti ∪ {H(ri )}. Clearly, Tn = P rΠ′ (T ′ ). Using induction on i we show that every Ti is G-consistent. Assume that Ti is G-consistent, but Ti+1 is not. Let ri = φ ⇒ ψ. By definition of Π′′ there are χ1 , . . . , χn ∈ Ti such V that (χ1 & . . . &χn ) ↔ φ holdsVon G. Let N = Ti \ {χ1 , . . . χn }. Assume that {φ, ¬( N )} is G-consistent, i.e., ν(φ&¬( N )) ̸= 0. In this case V for s = φ&¬( N ) ⇒ ψ we have: V V ν(φ&¬( N )&ψ) ν(φ&ψ) − ν(φ& N &ψ) V V ν(s) = = . ν(φ&¬( N )) ν(φ) − ν(φ& N )
11
V V V V We have G |= Ti+1 ↔ (φ& VN &ψ) and G |= V Ti ↔ (φ& N ) by choice of χ1 , . . .V , χn . Since by assumption ν( Ti+1 ) = 0 and ν( Ti ) ̸= 0, we conclude that V ν(φ& N &ψ) = 0 and ν(φ& N ) ̸= 0. In this way, we have ν(φ&ψ) ν(φ&ψ) V ν(s) = > = ν(ri ). ν(φ) − ν(φ& N ) ν(φ) Since s ≺ ri and ν(s) > ν(ri ), then there is r′ ∈ R such that r′ ⪯ s and ν(r′ ) > ′ ν(ri ). On the other hand, from ri ∈ MSCR(R, G, µ) and ri ≻ r′ we obtain V ν(ri ) ≥ ν(r ). This contradiction proves that V the body of s is not G-consistent: ν(φ&¬( N )) = 0. As a consequence we obtain ν(φ&¬( N )&ψ) = 0. Now we have: ^ ^ ν(φ&ψ) = ν(φ&ψ) − ν(φ&¬( N )&ψ) = ν(φ& N &ψ) = 0. Thus, ν(ri ) = 0. At the same time ri is a causal rule, which implies 0 = ν(ri ) > ν(⊤ ⇒ ψ) ≥ 0. The obtained contradiction concludes the proof. □
7 Conclusion Analysis of the history surrounding the RMS discussions has made it possible to precisely formalize and resolve Carl Hempel’s statistical ambiguity problem. This result does not only sum up the historical debate about RMS requirements, but also may open new directions in the philosophy of science. It follows from the result that IS inference can discover logically consistent natural-scientific theories. It also follows that cyclic causal relations, which reflect the integrity of objects of the perceived world, loop back on themselves and form consistent “causal models” of the external world categories Rehder (2003, 2016). Based on these causal models a “probabilistic formal concepts” were defined that provide the idealized description of categories Vityaev et al. (2012); Vityaev and Martinovich (2015). The causal relations between actions and their outcomes describe the goal-directed activity of humans and animals in accordance with the physiological Theory of Functional Systems Vityaev (2015). For a class of rules of the form α1 ∧. . .∧αn ⇒ β , where αi and β are literals (atoms or their negations), a program system “Discovery” was developed that discovers a set of MSCRs on a sample data D for prediction of some goal property ψ , using the exact Fisher test for statistical estimation of probabilistic inequalities of definition 8 Vityaev and Kovalerchuk (2004). This system was successfully applied to solve several tasks such as financial forecasting Kovalerchuk and Vityaev (2000), medicine Kovalerchuk et al. (2001), and bio informatics Vityaev et al. (2002). Based on MSCR a more powerful program system may be developed for solving Causal AI tasks. For example, for the digital twins control systems development, where decisions based on predictions are rather important.
References Cartwright N (1979) Causal laws and effective strategies. Nous 13 Coffa A (1974) Hempel’s ambiguity. Synrhese 28 Fagin R, Halpern J, Megiddo N (1990) A logic for reasoning about probabilities. Inform Comput 80 12
Fetzer J (1981) Scientific Knowledge: Causation, Explanation, and Corroboration. Dordrecht, D. Reidel Fetzer J (1993) Philosophy of Science. Paragon House, New York Hempel C (1965) Aspects of scientific explanation. In: Aspects of Scientific Explanation and other Essays in the Philosophy of Science. The Free Press Hempel C (1968) Maximal specificity and lawlikeness in probabilistic explanation. Philosophy of Science 35:116–33 Kaddour J, Lynch A, Liu Q, et al (2022) Causal Machine Learning: A Survey and Open Problems. arXiv:2206.15475 Kovalerchuk B, Vityaev E (2000) Data Mining in Finance: Advances in Relational and Hybrid methods. Kluwer Academic Publishers Kovalerchuk B, Vityaev E, Ruiz J (2001) Consistent and complete data and ”expert” mining in medicine. In: Medical Data Mining and Knowledge Discovery. Springer Rehder B (2003) Categorization as causal reasoning. Cognitive Science 27:709–748 Rehder B (2016) Reasoning with causal cycles. Cognitive Science pp 1–59 Salmon W (1970) Statistical explanation. In: The Nature and Function of Scientific Theories, Robert Garland Colodny (ed.). Dordrecht, p 173–231 Stanford (2008) Probabilistic causation. substantive revision. In: Stanford Encyclopedia of Philosophy Vityaev EE (2015) Purposefulness as a principle of brain activity. In: Anticipation: Learning from the Past, (ed.) M. Nadin. Cognitive Systems Monographs. Springer Vityaev E (2006) The logic of prediction. In: Proceedings of the 9th Asian Logic Conference. World Scientific Publishers, p 263–276 Vityaev E, Kovalerchuk B (2004) Empirical theories discovery based on the measurement theory. Mind and Machine 14:551–573 Vityaev E, Martinovich V (2015) Probabilistic formal concepts with negation. In: PSI 2014, LNCS. 8974. Springer Vityaev E, Odintsov S (2019) How to predict consistently? In: Trends in Mathematics and Computational Intelligence. Studies in Computational Intelligence. (ed.) Maria Eugenia Cornejo. 796. Springer, p 35–41 Vityaev E, Orlov Y, Vishnevsky O, et al (2002) Computer system ”gene discovery” for promoter structure analysis. In Silico Biol 2:257–262
13
Vityaev E, Demin A, Ponomaryov D (2012) Probabilistic generalization of formal concepts. Programming and Computer Software 38
14