ConceptioArchiveNCBI Bookshelf
NCBI Bookshelfopen access

Reference Guide on Statistics and Research Methods

· ncbi_books
NCBI Bookshelf · Textbooks · License: Open Access
Open Source ↗
philosophyofmind
philosophy of mind

var ncbi_startTime = new Date(); Reference Guide on Statistics and Research Methods - Reference Manual on Scientific Evidence - NCBI Bookshelf p a.figpopup{display:inline !important} .bk_tt {font-family: monospace} .first-line-outdent .bk_ref {display: inline} .body-content h2, .body-content .h2 {border-bottom: 1px solid #97B0C8} .body-content h2.inline {border-bottom: none} a.page-toc-label , .jig-ncbismoothscroll a {text-decoration:none;border:0 !important} .temp-labeled-list .graphic {display:inline-block !important} .temp-labeled-list img{width:100%} window.name="mainwindow"; Warning: more... An official website of the United States government Here's how you know The .gov means it's official. The site is secure. https:// Log in Show account info Close Account Logged in as: username Dashboard Publications Account settings Log out Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation Bookshelf Search database Books All Databases Assembly Biocollections BioProject BioSample Books ClinVar Conserved Domains dbVar Gene Genome GEO DataSets GEO Profiles GTR Identical Protein Groups MedGen MeSH NLM Catalog Nucleotide OMIM PMC Protein Protein Clusters Protein Family Models PubChem BioAssay PubChem Compound PubChem Substance PubMed SNP SRA Structure Taxonomy ToolKit ToolKitAll ToolKitBookgh Search term Search Browse Titles Advanced Help Disclaimer NCBI Bookshelf. A service of the National Library of Medicine, National Institutes of Health. Federal Judicial Center; National Academies of Sciences, Engineering, and Medicine; Policy and Global Affairs; Committee on Science, Technology, and Law; Committee on Science for Judges—Development of the Reference Manual on Scientific Evidence, Fourth Edition. Reference Manual on Scientific Evidence: Fourth Edition. Washington (DC): National Academies Press (US); 2025 Dec 31. Reference Manual on Scientific Evidence: Fourth Edition. Show details Federal Judicial Center; National Academies of Sciences, Engineering, and Medicine; Policy and Global Affairs; Committee on Science, Technology, and Law; Committee on Science for Judges—Development of the Reference Manual on Scientific Evidence, Fourth Edition. Washington (DC): National Academies Press (US) Contents Hardcopy Version at National Academies Press Search term < Prev Next > Reference Guide on Statistics and Research Methods DAVID H. KAYE AND HAL S. STERN David H. Kaye, M.A., J.D., is Regents Professor Emeritus, Arizona State University Sandra Day O’Connor College of Law and School of Life Sciences, and Distinguished Professor of Law and Academy Professor Emeritus, Pennsylvania State University School of Law. Hal S. Stern, Ph.D., is Distinguished Professor, Department of Statistics, University of California, Irvine. Authors’ Note: Research for this reference guide was completed in 2023. Professor David A. Freedman co-authored the first three editions of this reference guide. His writing and thinking are evident throughout this edition as well. CONTENTS Introduction Admissibility and Weight of Statistical Studies Varieties and Limits of Statistical Expertise Procedures That Enhance Statistical Testimony Maintaining Professional Autonomy Disclosing Limitations and Other Analyses Disclosing Data and Analytical Methods Before Trial How Have the Data Been Collected? Is the Study Designed to Investigate Causation? Types of Studies Randomized Controlled Experiments Observational Studies Generalizing the Results Descriptive Surveys and Censuses What Method Is Used to Select the Units? A sampling frame Selection bias Of the Units Selected, Which Provide Measurements? Individual Measurements Is the Measurement Process Reliable? Is the Measurement Process Valid? Are the Measurements Recorded Correctly? What Does It Mean to Be Random? How Have the Data Been Presented? Are Rates or Percentages Properly Interpreted? How Big Is the Base of a Percentage? Have Appropriate Benchmarks Been Provided? Have the Data Collection Procedures Changed? Are the Categories Appropriate? What Comparisons Are Made? Is an Appropriate Measure of Association Used? Does a Graph Portray Data Fairly? How Are Trends Displayed? How Are Distributions Displayed? Is an Appropriate Measure Used for the Center of a Distribution? Is an Appropriate Measure of Variability Used? What Inferences Can Be Drawn from the Data? Estimation What Estimator Should Be Used? What Is the Standard Error? What Is the Confidence Interval? The normal curve and large samples Other situations How Big Should the Sample Be? What Are the Technical and Interpretive Difficulties with Confidence Intervals? p What Is the p Is a Difference Statistically Significant? Recent Emphasis on the Limitations of p Tests or Interval Estimates? Is the Sample Statistically Significant? Evaluating Hypothesis Tests What Is the Power of the Test? What About Small Samples? One Tail or Two? How Many Tests Have Been Done? What Are the Rival Hypotheses? Bayesian Statistical Methods and Posterior Probabilities Correlation and Regression Scatter Diagrams Correlation Coefficients Is the Association Linear? Do Outliers Influence the Correlation Coefficient? Does a Confounding Variable Influence the Coefficient? Regression Lines What Are the Slope and Intercept? What Is the Unit of Analysis? Statistical Models Data Science and Statistical Machine Learning What Is Data Science? What Is Machine Learning? What Statistical Questions Arise with Machine Learning Studies? Is the Dataset Appropriate and of Sufficient Quality? Is the Predictor or Classifier Robust? Is the Predictor or Classifier Too Opaque? Appendix: Conditional Probability and Bayes’ Rule What Do Probabilities Apply To? What Are Conditional Probabilities? What Is Bayes’ Rule? Glossary of Terms References on Statistics and Research Methods Nontechnical Surveys General References FIGURES 1. and 2. Manipulating the scale of a graph 3. Histogram showing how frequently various numbers of heads appeared in 50 batches of 10 tosses of a quarter 4. Confidence coefficients for CIs of ± 1, 2, and 3 standard errors of a normally distributed estimator 5. Plotting points in a scatter diagram 6. Scatter diagram for income and education: men ages 25 to 34 in Kansas 7. The correlation coefficient measures the sign of a linear association, and its strength 8. A strong nonlinear association with a correlation coefficient close to zero 9. The correlation coefficient can be distorted by outliers 10. The regression line for income on education and its estimates 11. Scatter diagram for income and education, with the regression line indicating the trend 12. Turnout rate for the white candidate plotted against the percentage of registrants who are white. Precinct-level data, 1982 Democratic primary for auditor, Lee County, South Carolina TABLES 1. Test Results for Cartridge-case Comparisons 2. Data Used by a Defendant to Refute Plaintiff’s False Advertising Claims 3. Home Pregnancy Test Results 4. Test Results for Cartridge-case Comparisons 5. Admissions by Sex 6. Admissions by Sex and College Introduction Statistical assessments are prominent in many kinds of legal cases, including antitrust, employment discrimination, toxic torts, and voting rights cases. This reference guide describes the elements of statistical reasoning. We hope the explanations will help judges and lawyers to understand statistical terminology, to see the strengths and weaknesses of statistical arguments, and to apply relevant legal doctrine. The guide is organized as follows: This introduction provides an overview of the field, discusses the admissibility of statistical studies, and offers some suggestions about procedures that encourage the best use of statistical evidence. The section titled “How Have the Data Been Collected?” addresses data collection. It explains why the design of a study is the most important determinant of its quality. The section compares experiments with observational studies and surveys with censuses, indicating when the various kinds of study are likely to provide useful results. The section titled “How Have the Data Been Presented?” discusses the art of summarizing data. This section considers the mean, median, and standard deviation. These are basic descriptive statistics, and most statistical analyses use them as building blocks. This section also discusses patterns in data that are brought out by graphs, percentages, and tables. The section titled “What Inferences Can Be Drawn from the Data?” describes the logic of statistical inference, emphasizing foundations and disclosing limitations. This section covers estimation, standard errors and confidence intervals, p The section titled “Correlation and Regression” shows how associations can be described by scatter diagrams, correlation coefficients, and regression lines. Regression is often used in attempts to infer causation from association. This section explains the technique, indicating the circumstances under which it and other statistical models are likely to succeed—or fail. Advances in computing speed and the availability of large datasets have made it possible to fit models of greater complexity than those described in the section titled “Correlation and Regression.” The section titled “Data Science and Statistical Machine Learning” describes the developing fields of “data science” and “machine learning” and statistical issues that affect the confidence one can have in the output of machine-learning techniques. An appendix discusses the scope of the theory of probability, conditional probabilities, Bayes’ rule, and perspectives on statistical inference. The glossary defines statistical terms that may be encountered in litigation, including ones that do not appear in the body of the guide. Admissibility and Weight of Statistical Studies Statistical studies suitably designed to address a material issue generally will be admissible under the Federal Rules of Evidence. The hearsay rule rarely is a serious barrier to the presentation of statistical studies, because such studies may be offered to explain the basis for an expert’s opinion or may be admissible under the learned treatise exception to the hearsay rule. 1 Daubert v. Merrell Dow Pharmaceuticals, Inc 2 3 Daubert 4 Varieties and Limits of Statistical Expertise For convenience, the field of statistics may be divided into three subfields: probability theory, theoretical statistics, and applied statistics. Probability theory is the mathematical study of outcomes that are governed, at least in part, by chance. Theoretical statistics is about understanding the properties of statistical procedures, including error rates; probability theory plays a key role in this endeavor. Applied statistics draws on both these fields to develop techniques for collecting or analyzing particular types of data. Statistical expertise is not confined to witnesses with degrees in statistics. Because statistical reasoning underlies many kinds of empirical research, scholars in a variety of fields—including biology, business, economics, epidemiology, medicine, political science, psychology, and sociology—are exposed to statistical ideas, with an emphasis on the methods most important to the discipline. The diffusion of statistical concepts and methods across so many fields raises the question of who is qualified to conduct and testify to statistical assessments—a statistical methodologist, a substantive scientist, or both? Much depends on context. If the study involves assembling and then analyzing case-specific data, the choice of which data to examine and how best to model a particular process could require both subject-matter and general statistical expertise. The two types of expertise might be combined in a single individual. A labor economist, for example, should be able to supply a definition of the relevant labor market from which an employer draws its employees. This economist also might be sufficiently educated in statistical tests for group differences to compare the race of new hires to the racial composition of the labor market. When the analysis is as straightforward as the comparison of two proportions, a single substantive expert may suffice. 5 The critical question is whether the individual has the education and experience to determine which statistical techniques are appropriate for the task at hand and to apply such techniques with an appreciation of their limitations. Experts who specialize in using statistical methods, and whose professional careers demonstrate this orientation, are most likely to use appropriate procedures and correctly interpret the results. To ascertain the extent to which an expert has this orientation and acumen, one can look to formal education, professional accomplishments, and reputation in the pertinent community of experts. Has the individual studied quantitative methods? Taught them? Used them in research? Which ones? If the expert is or was an academic professional, a university teaching portfolio with courses on statistical methods is a good sign. Ideally, the publication and research record will include studies that use or describe the same or similar methods as those applied (or applicable) to the case at bar. 6 Again, not every case involving computations of probability or statistics necessitates a highly skilled statistical specialist. Medical practitioners and forensic scientists who are not statistical specialists often make statistical assessments of evidence or rely on and present the results of statistical studies. In these situations, it is important to ensure that the witnesses do not exceed the bounds of their statistical expertise. The scientist or practitioner might lack basic information about the studies underlying their testimony. State v. Garrison 7 8 Procedures That Enhance Statistical Testimony Maintaining Professional Autonomy Ideally, experts who conduct research in the context of litigation should proceed with the same objectivity that would be required in other professional contexts. Thus, experts who testify (or who supply results used in testimony) should conduct the analysis required to address, in a professionally responsible fashion, the issues posed by the litigation. 9 Disclosing Limitations and Other Analyses Statisticians analyze data using a variety of methods. To permit a fair evaluation of the analysis that is eventually settled on, the testifying expert can be asked how that approach was developed, whether alternative approaches were considered, and, if so, what the results were. 10 11 Disclosing Data and Analytical Methods Before Trial The collection of data often is expensive and subject to errors and omissions. Moreover, careful exploration of the data can be time-consuming. To minimize debates at trial over the accuracy of data and the choice of analytical techniques, pretrial-discovery procedures should be used, particularly with respect to the quality of the data and the method of analysis. 12 before 13 How Have the Data Been Collected? The interpretation of data often depends on understanding “study design”—the plan for a statistical study and its implementation. 14 In many cases, statistical studies are used to show causation. Do food additives cause cancer? Does capital punishment deter crime? Would additional disclosures in a securities prospectus cause investors to behave differently? The design of studies to investigate causation is the first topic of this section. 15 Sample data can be used to describe a population. The population is the whole class of units that are of interest; the sample is the set of units chosen for detailed study. Inferences from the part to the whole are justified when the sample is representative. Sampling is the second topic of this section. Finally, issues associated with the reliability and accuracy of collected data will be considered. Measurement error should be assessed and the likely impact of errors considered. Data quality is the third topic of this section. All the sections concern the study of “variables.” In statistics, a variable is a characteristic of the units in a study. With a study of people, the unit of analysis is the person, and the variables describe people. Two such variables would be income (dollars per year) and educational level (years of schooling completed). With a study of school districts, the unit of analysis is the district. Typical variables include average family income of district residents and average test scores of students in the district. Variables may be related to one another in various ways. Many studies examine whether a variable or group of variables, known as independent variables, are related to an outcome or dependent variable. For example, census data can be analyzed to determine whether people who complete more years of school tend to have higher incomes later in life. Educational level would be the independent variable, and annual income the dependent variable. In a study of smoking and lung cancer, the independent variable could be smoking (perhaps measured by the number of cigarettes smoked per day), and the dependent variable could mark the presence or absence of lung cancer. Is the Study Designed to Investigate Causation? Types of Studies When causation is the issue, anecdotal evidence can be brought to bear. So can observational studies or controlled experiments. Anecdotal reports may be of value, but they are ordinarily more helpful in generating lines of inquiry than in proving causation. In medicine, evidence from clinical practice can be a good starting point for discovery of cause-and-effect relationships, but the anecdotal experience of practitioners is not definitive. 16 Anecdotal evidence usually amounts to reports that events of one kind are followed by events of another kind. Typically, the reports are not even sufficient to show association, because there is no comparison group. For example, some children who live near power lines develop leukemia. Does exposure to electrical and magnetic fields cause this disease? The anecdotal evidence is not compelling because leukemia also occurs among children without exposure. 17 The next issue is crucial: Exposed and unexposed people may differ in ways other than the exposure they have experienced. For example, children who live near power lines could come from poorer families and be more at risk from other environmental hazards. Such differences can create the appearance of a cause- and-effect relationship. Other differences can mask a real relationship. Cause-and-effect relationships often are quite subtle, and carefully designed studies are needed to draw valid conclusions. An epidemiological classic makes the point. At one time, it was thought that lung cancer was caused by fumes from tarring the roads, because many lung cancer patients lived near roads that recently had been tarred. This is anecdotal evidence. But the argument is incomplete. For one thing, most people—whether exposed to asphalt fumes or unexposed—did not develop lung cancer. A comparison of rates was needed. Epidemiologists found that exposed persons and unexposed persons suffered from lung cancer at similar rates: Tar was probably not the causal agent. Exposure to cigarette smoke, however, turned out to be strongly associated with lung cancer. This study, in combination with later ones, made a compelling case that smoking cigarettes is the main cause of lung cancer. 18 A good study design compares outcomes for subjects who are exposed to some factor (the treatment group) with outcomes for other subjects who are not exposed (the control group). With comparison groups, there is another important distinction—between controlled experiments and observational studies. In a controlled experiment, the investigators decide which subjects will be exposed and which subjects will be in the control group. In observational studies, the researchers do not determine which subjects are exposed; often, the subjects themselves choose their exposures. Because of self-selection or for other reasons, the treatment and control groups are likely to differ with respect to influential factors other than the ones of primary interest. These other factors are called lurking variables or confounding variables. 19 Confounding remains a problem even for the best observational research. For example, women with herpes are more likely to develop cervical cancer than other women. Some investigators concluded that herpes caused cancer: In other words, they thought the association was causal. Later research showed that the primary cause of cervical cancer was human papilloma virus (HPV). Herpes was a marker of sexual activity. Women who had multiple sexual partners were more likely to be exposed not only to herpes but also to HPV. The association between herpes and cervical cancer was due to other variables. 20 Randomized Controlled Experiments In randomized controlled experiments, investigators assign subjects to treatment or control groups at random. The groups are therefore likely to be comparable, except for the treatment. This minimizes the role of confounding. Minor imbalances will remain, owing to the play of random chance; the likely effect on study results can be assessed by statistical techniques. 21 The following example should help bring the discussion together. Today, we know that taking aspirin helps prevent heart attacks. But initially, there was some controversy. People who take aspirin rarely have heart attacks. This is anecdotal evidence for a protective effect, but it proves almost nothing. After all, few people have frequent heart attacks, whether or not they take aspirin regularly. A good study compares heart-attack rates for two groups: people who take aspirin (the treatment group) and people who do not (the controls). An observational study would be easy to do, but in such a study the aspirin-takers are likely to be different from the controls. Indeed, they are likely to be sicker—that is why they are taking aspirin. The study would be biased against finding a protective effect. Randomized experiments are harder to do, but they provide better evidence. 22 23 In summary, data from a treatment group without a control group generally reveal very little and can be misleading. Comparisons are essential. If subjects are assigned to treatment and control groups at random, a difference in the outcomes between the two groups can usually be accepted, within the limits of statistical error, 24 Observational Studies The bulk of the statistical studies seen in court are observational, not experimental. Take the question of whether capital punishment deters murder. To conduct a randomized controlled experiment, people would need to be assigned randomly to a treatment group or a control group. People in the treatment group would know they were subject to the death penalty for murder; the controls would know that they were exempt. Conducting such an experiment is not possible. Many studies of the deterrent effect of the death penalty have been conducted, all observational, and some have attracted judicial attention. Researchers have catalogued differences in the incidence of murder in states with and without the death penalty and have analyzed changes in homicide rates and execution rates over the years. When reporting on such observational studies, investigators may speak of “control groups” (e.g., the states without capital punishment) or claim they are “controlling for” confounding variables by statistical methods such as multiple regression. 25 When an external change in circumstances naturally creates a treatment and a control group in a manner that is comparable to random assignment by the researchers, the study may be called a “natural experiment” or a “quasi-experiment.” 26 27 28 seemingly Observational studies can be very useful even when assignments to the groups being compared do not resemble randomization imposed by an experimenter. For example, there is strong observational evidence that smoking causes lung cancer (see section titled “Types of Studies” above). Generally, observational studies provide good evidence in the following circumstances: The association is seen in studies with different designs, on different kinds of subjects, and done by different research groups. 29 The association holds when effects of confounding variables are taken into account by appropriate methods, for example, comparing smaller groups that are relatively homogeneous with respect to the confounders. 30 There is a plausible explanation for the effect of the independent variable; alternative explanations in terms of confounding should be less plausible than the proposed causal link. 31 Observational studies can produce legitimate disagreement among experts, and there is no mechanical procedure for resolving such differences of opinion. In the end, deciding whether associations are causal typically is not a matter of statistics alone, but also rests on scientific judgment. 32 There are, however, some basic questions to ask when appraising causal inferences based on empirical studies: Was there a control group? Unless comparisons can be made, the study has little to say about causation. If there was a control group, how were subjects assigned to treatment or control: through a process under the control of the investigator (a controlled experiment) or through a process outside the control of the investigator (an observational study)? If the study was a controlled experiment, was the assignment made using a chance mechanism (randomization), or did it depend on the judgment of the investigator? If the data came from an observational study or a nonrandomized controlled experiment, How did the subjects come to be in treatment or in control groups? Are the treatment and control groups comparable? If not, what adjustments were made to address confounding? Were the adjustments sensible and sufficient? Generalizing the Results In considering what conclusions can be drawn from studies, it is helpful to distinguish between internal external Any study must be conducted on certain subjects, at certain times and places, and using certain treatments. To extrapolate from the conditions of a study to more general conditions raises questions of external validity. For example, studies suggest that definitions of insanity given to jurors influence decisions in cases of incest. Would the definitions have a similar effect in cases of murder? Other studies indicate that recidivism rates for ex-convicts are not affected by providing them with temporary financial support after release. Would similar results be obtained if conditions in the labor market were different? Confidence in the appropriateness of an extrapolation cannot come from the experiment itself. It comes from knowledge about outside factors that would or would not affect the outcome. Such judgments are easiest in the physical and life sciences, but even here, there are problems. For example, it may be difficult to infer human responses to substances that affect animals. First, there are often inconsistencies across test species. A chemical may be carcinogenic in mice but not in rats. Extrapolation from rodents to humans is even more problematic. Second, to get measurable effects in animal experiments, chemicals are administered at very high doses. Results are extrapolated—using mathematical models—to the very low doses of concern in humans. However, there are many dose–response models to use and few grounds for choosing among them. Generally, different models produce radically different estimates of the “virtually safe dose” in humans. 33 Sometimes several studies, each having different limitations, all point in the same direction. This combination is why most experts believe that smoking causes lung cancer and many other diseases. So, too, a variety of studies indicate that jurors who approve of the death penalty are more likely to convict in a capital case. 34 Descriptive Surveys and Censuses We now turn to a second topic—choosing units for study. A census tries to measure some characteristic of every unit in a population. This is often impractical. Then investigators use sample surveys, which measure characteristics for only part of a population. The accuracy of the information collected in a census or survey depends on how the units are selected for study and how the measurements are made. 35 What Method Is Used to Select the Units? By definition, a census seeks to measure some characteristic of every unit in a whole population. It may fall short of this goal; in which case one must ask whether the missing data are likely to differ in some systematic way from the data that are collected. 36 A sampling frame To illustrate a sampling frame, suppose that a defendant in a criminal case seeks a change of venue. According to the defendant, popular opinion is so adverse that it would be difficult to impanel an unbiased jury. To prove the state of popular opinion, the defendant commissions a survey. The relevant population consists of everyone in the jurisdiction who might be called for jury duty. The sampling frame is the list of all potential jurors, which is maintained by court officials and is made available to the defendant. In this hypothetical case, the fit between the sampling frame and the population would be excellent. In other situations, the sampling frame is more problematic. In an obscenity case, for example, the defendant can offer a survey of community standards. 37 38 Selection bias Many surveys do not use probability methods. In commercial disputes involving trademarks or advertising, the population of all potential purchasers of a product is hard to identify. Pollsters may resort to an easily accessible subgroup of the population—for example, shoppers in a mall. Such convenience samples may be biased by the interviewer’s discretion in deciding whom to approach—a form of selection bias—and the refusal of some of those approached to participate—nonresponse bias (see section titled “Of the Units Selected, Which Provide Measurements?” below). Selection bias is acute when constituents write their representatives, listeners call into radio talk shows, interest groups collect information from their members, individuals complete available online surveys, or attorneys choose cases for trial. 39 A well-known example of selection bias is the 1936 Literary Digest Digest 40 Digest 41 There are procedures that attempt to correct for selection bias. In quota sampling, for example, the interviewer is instructed to interview so many women, so many older people, so many ethnic minorities, and the like. But quotas still leave discretion to the interviewers in selecting members of each demographic group and therefore do not solve the problem of selection bias. 42 Probability methods are designed to avoid selection bias. Once the population is reduced to a sampling frame, the units to be measured are selected by a lottery that gives each unit in the sampling frame a known, nonzero probability of being chosen. 43 44 Of the Units Selected, Which Provide Measurements? Probability sampling ensures that within the limits of chance, the sample will be representative of the sampling frame. But will all these units be measured? When documents are sampled for audit, they can all be examined, at least in principle. Human beings are less easily managed, and some will refuse to cooperate. In the 1936 Literary Digest 45 Surveys should therefore report nonresponse rates. A large nonresponse rate warns of potential bias. 46 47 In short, a good survey defines an appropriate population, uses a probability method for selecting the sample, has a high response rate, and gathers accurate information on the sample units. When these goals are met, the sample tends to be representative of the population, and data from the sample can be extrapolated to describe the characteristics of the population. Of course, surveys may be useful even if they fail to meet these criteria. But then, additional arguments are needed to justify the inferences. Individual Measurements Is the Measurement Process Reliable? Reliability and validity are two aspects of accuracy in measurement. In statistics, reliability refers to reproducibility of results. A reliable measuring instrument returns consistent measurements. A scale, for example, is perfectly reliable if it always reports the same weight for the same unchanged object. It may not be accurate—it may always report a weight that is too high or one that is too low—but the perfectly reliable scale always reports the same weight for the same object. Its errors, if any, are systematic: They tend to point in the same direction. Courts often use “reliable” to mean “that which can be relied on” for some purpose, such as establishing probable cause through a reliable informant or as part of an argument for admitting hearsay statements. 48 Daubert v. Merrell Dow Pharmaceuticals 49 50 The Supreme Court was faced with the imperfect reliability of a test for IQ scores in Hall v. Florida 51 52 53 A simpler courtroom example comes from DNA identification. An early method of identification required laboratories to determine the lengths of fragments of DNA. By making independent repeated measurements of the same fragments, laboratories determined the likelihood that two measurements differed by specified amounts. 54 55 Coding of data also can affect reliability. In many studies, descriptive information is obtained on the subjects. For statistical purposes, the information usually has to be reduced to numbers. The process of reducing information to numbers is called “coding,” and the reliability of the process should be evaluated. For example, in a meticulous study of death sentencing in Georgia, legally trained evaluators examined short summaries of cases and ranked them according to the defendant’s culpability. 56 Metrologists (specialists in measurement science) similarly distinguish between “reproducibility” and “repeatability.” The difference lies in the degree to which the conditions for making measurements are similar. With a repeatable procedure, measurements by the same examiner using the same equipment at the same place and time should be close to one another. With a reproducible procedure, measurements from different examiners even at different places and times also should be similar. 57 Is the Measurement Process Valid? Reliability is necessary but not sufficient to ensure accuracy. In addition to reliability, validity is needed. A valid measuring instrument measures what it is supposed to. Thus, a polygraph measures certain physiological variables, for example, pulse rate or blood pressure, in response to stimuli. The measurements may be reliable. Nonetheless, the polygraph is not valid as a lie detector unless the measurements are well correlated with lying. 58 When there is an established way of measuring a variable, a new measurement process can be validated by comparison with the established one. Breathalyzer readings can be validated against alcohol levels found in blood samples. LSAT or GRE scores used for law school admissions can be validated against grades earned in law school. A common measure of validity is the correlation coefficient between the predictor and the criterion (for example, test scores and later performance). 59 Employment discrimination cases illustrate some of the difficulties. Plaintiffs suing under Title VII of the Civil Rights Act may challenge an employment test that has a disparate impact on a protected group, and defendants may try to justify the use of a test as valid, reliable, and a business necessity. 60 61 A further problem is that test-takers are likely to be a select group. The ones who get the jobs are even more highly selected. Generally, selection attenuates (weakens) the correlations because differences in performance within a narrow band of highly qualified applicants do not exhibit as much variation as would be expected if a wider range of applicants were selected. 62 63 Measurements also can be made on a nominal scale, as when a chemist notes that litmus paper has turned red, or a criminalist declares that there is a “physical fit” between two fragments of glass. For binary (yes–no) classifications, validity (accuracy) of the classifier can be evaluated with experiments to estimate “sensitivity” and “specificity” (or corresponding error probabilities). Suppose that firearms examiners are given pairs of spent cartridge cases and are required to decide whether they come from the same gun or from two different guns. The experimenter, who has prepared the test pairs, knows the truth; the examiners do not. Results related to a small experiment are in Table 1 64 Table 1 Test Results for Cartridge-case Comparisons. The examiners performed flawlessly in these 75 instances. If we call the same-source condition “positive” and the “different-source” condition “negative,” there were zero false-positive reports and zero false-negative reports. To put it another way, the examiners were 100% accurate in dealing with same-source pairs—they called 30 out of 30 such pairs positives, for an observed sensitivity of 1; likewise, they were 100% accurate in dealing with different-source pairs—they called 45 out of 45 such pairs different, for an observed specificity of 1. Whether the results of this small experiment are representative of what might be seen in a larger set of experiments, and whether the experiments would be representative of the outcomes in case work are further questions. Are the Measurements Recorded Correctly? Judging the adequacy of data collection involves an examination of the process by which measurements are taken. Are responses to interviews coded correctly? Do mistakes distort the results? How much data are missing? What was done to compensate for gaps in the data? These days, data are stored in computer files. Cross-checking the files against the original sources (e.g., paper records), at least on a sample basis, can be informative. Data quality is a pervasive issue in litigation and in applied statistics more generally. 65 66 What Does It Mean to Be Random? In the law, a selection process sometimes is called “random,” provided that it does not exclude identifiable segments of the population. Statisticians use the term in a more rigorous and technical sense. For example, to choose one person at random from a population in the strict statistical sense, we would have to ensure that everybody in the population has the same probability of selection. With a randomized controlled experiment, subjects are assigned to treatment or control at random in the strict sense—by tossing coins, throwing dice, looking at tables of random numbers, or more commonly these days, by using a random number generator on a computer. The same rigorous definition applies to random sampling. Randomness in the technical sense provides assurance of unbiased estimates from a randomized controlled experiment or a probability sample. Randomness in the technical sense also justifies calculations of standard errors, confidence intervals, and p How Have the Data Been Presented? After data have been collected, they should be presented in a way that makes them intelligible and that helps reveal their implications. Data can be summarized with a few numbers or with graphical displays. However, the wrong summary can mislead. 67 Are Rates or Percentages Properly Interpreted? How Big Is the Base of a Percentage? Rates and percentages often provide effective summaries of data, but these statistics can be misinterpreted. A rate reports a comparison of one number against some other quantity, for example, the number of reported crimes per hundred thousand residents. A percentage reports a comparison between two numbers by putting them in terms of a common base (100). Expressing the ratio of the two numbers on a common base makes it easy to compare them. One application of percentages is for reporting increases or decreases in a quantity or rate by describing the percent change compared to the initial or base amount. When the base is small, however, a small change in absolute terms can generate a large percentage gain or loss. (This could lead to newspaper headlines such as “Increase in Rate of Thefts Alarming,” even when the total number of thefts is small. 68 Have Appropriate Benchmarks Been Provided? The selective presentation of numerical information is like quoting someone out of context. Is the fact that a particular actively managed fund of large-cap stocks boasted a return of 25% in 2021 indicative of outstanding management? Considering that the 500 large-cap stocks in “the benchmark S&P 500 notched a total return . . . of 28.7% [that] year,” a growth rate of 25% is less indicative of unusual financial acumen than one might have thought. 69 70 Have the Data Collection Procedures Changed? Changes in the process of collecting data can create problems of interpretation. Statistics on crime provide many examples. The number of petty larcenies reported in Chicago more than doubled one year—not because of an abrupt crime wave, but because a new police commissioner introduced an improved reporting system. 71 72 73 Are the Categories Appropriate? Misleading summaries also can be produced by the categories used for comparison. In Philip Morris, Inc. v. Loew’s Theatres, Inc. 74 R.J. Reynolds Tobacco Co. v. Loew’s Theatres, Inc. 75 Table 2 76 77 78 Table 2 Data Used by a Defendant to Refute Plaintiff’s False Advertising Claims. There was a similar distortion in claims for the accuracy of a home pregnancy test. The manufacturer advertised the test as 99.5% accurate under laboratory conditions. The underlying data are summarized in Table 3 Table 3 Home Pregnancy Test Results. The table does indicate that only one error occurred in 200 assessments, for 99.5% overall accuracy. But the table also shows that the test can make two types of errors: It can tell a pregnant woman that she is not pregnant (a false negative), and it can tell a woman who is not pregnant that she is (a false positive). The reported 99.5% accuracy rate conceals a crucial fact—the company had virtually no data with which to measure the rate of false positives. 79 The problem with combining categories into broader ones has surfaced in criminal cases as well. As noted in the section titled “Is the Measurement Process Valid?” above, criminalists examine pairs of items, such as spent cartridge cases, to decide whether they are associated with the same source. In firearms-toolmark matching, the traditional comparison between an item of known origin and the item whose origin is in question culminates in a report of “identification,” “inconclusive,” or “elimination.” In validity studies, toolmark examiners are given items to evaluate when the experimenter, but not the criminalist, knows whether the items are from the same source or from two different sources. Table 4 Table 1 Table 4 Test Results for Cartridge-case comparisons The new row reveals that the criminalists reached no definitive conclusion in most of the cases, and that all these “inconclusives” pertained to pairs of cartridge cases from two different guns. Adding in all these “inconclusives,” the overall proportion of correct decisions drops from 100% to only (30 + 45) / (30 + 45 + 157) = 32%. In one sense, this is a poor “correct decision rate.” 81 82 83 What Comparisons Are Made? Finally, there is the issue of which numbers to compare. Researchers sometimes choose among alternative comparisons. Why did they choose the one they did? Would another comparison give a different view? A government agency, for example, may want to compare the amount of service now being given with that of earlier years—but what earlier year should be the baseline? If the first year of operation is used, a large percentage increase should be expected because of startup problems. If last year is used as the base, was it also part of the trend, or was it an unusually poor year? If the base year is not representative of other years, the percentage may not portray the trend fairly. No single question can be formulated to detect such distortions, but it may help to ask for the numbers from which the percentages were obtained; asking about the base can also be helpful. 84 Is an Appropriate Measure of Association Used? Many cases involve statistical association. Does a test for employee promotion have an exclusionary effect that depends on race or sex? Does the incidence of murder vary with the rate of executions for convicted murderers? Do consumer purchases of a product depend on the presence or absence of a product warning? This section discusses tables and percentage-based statistics that are frequently presented to answer such questions. 85 Percentages often are used to describe the association between two variables. Suppose that a university is alleged to discriminate against women in admitting students, and that the university consists of only two colleges—engineering and business. The university admits 350 out of 800 male applicants; by comparison, it admits only 200 out of 600 female applicants. Such data commonly are displayed as in Table 5 86 Table 5 Admissions by Sex. The table indicates that 350/800 = 44% of the males are admitted, compared with only 200/600 = 33% of the females. One way to express the disparity is to subtract the two percentages: 44% − 33% = 11 percentage points. Although such subtraction is commonly seen in jury discrimination cases, 87 For Table 5 88 89 However, the selection ratio has its own problems. In the last example, if the selection rates are 5% and 1%, then the exclusion rates are 95% and 99%. The ratio is 99/95 = 104%, meaning that females have, on average, 104% the risk of males of being rejected. The underlying facts are the same, of course, but this formulation sounds much less disturbing. A statistic known as the odds ratio is more symmetric. If 5% of male applicants are admitted, the odds on a man being admitted are 5 to 95 = 1 to 19; the odds on a woman being admitted are 1:99. The odds ratio is (1:99)/(1:19) = 19:99. The odds ratio for rejection instead of acceptance is the same, except that the order is reversed. 90 91 Data showing disparate impact are generally obtained by aggregating—putting together—data from a variety of sources. Unless the source material is fairly homogeneous, aggregation can distort patterns in the data. We illustrate the problem with the hypothetical admission data in Table 5 Table 6 Table 6 Admissions by Sex and College. The entries in Table 6 Table 5 Table 5 Table 6 92 Does a Graph Portray Data Fairly? Graphs are useful for revealing key characteristics of a batch of numbers, trends over time, and the relationships among variables. 93 How Are Trends Displayed? Graphs that plot values over time are useful for seeing trends. However, the scales on the axes matter. In Figures 1 2 Figure 2 94 95 Figures 1 and 2 Manipulating the scale of a graph. How Are Distributions Displayed? A graph commonly used to display the distribution of data is the histogram. One axis denotes the numbers, and the other indicates how often these numbers fall within specified intervals (called “bins” or “class intervals”). For example, we flipped a quarter 10 times in a row and counted the number of heads in this “batch” of 10 tosses. Repeating this exercise until we obtained 50 batches, we recorded the following counts: 96 The histogram is shown in Figure 3 97 Figure 3 Histogram showing how frequently various numbers of heads appeared in 50 batches of 10 tosses of a quarter. Is an Appropriate Measure Used for the Center of a Distribution? Perhaps the most familiar descriptive statistic is the mean (or “arithmetic mean”). The mean can be found by adding all the numbers and dividing the total by how many numbers were added. By comparison, the median cuts the numbers into halves: half the numbers are larger than the median and half are smaller. 98 Studies of damage awards in tort cases find that the mean is larger than the median. 99 100 101 Research also has shown considerable stability in the ratio of punitive to compensatory damage awards, and the Supreme Court has placed great weight on this ratio in deciding whether punitive damages are excessive in a particular case. In Exxon Shipping Co. v. Baker 102 103 104 105 Is an Appropriate Measure of Variability Used? The location of the center of a batch of numbers reveals nothing about the variations exhibited by these numbers. The numbers 1, 2, 5, 8, 9 have 5 as their mean and median. So do the numbers 5, 5, 5, 5, 5. In the first batch, the numbers vary considerably about their mean; in the second, the numbers do not vary at all. Statistical measures of variability include the range, the interquartile range, and the standard deviation. The range is the difference between the largest number in the batch and the smallest. The range seems natural, and it indicates the maximum spread in the numbers, but the range is unstable because it depends entirely on the most extreme values. The interquartile range is the difference between the 25th and 75th percentiles. By definition, 25% of the data fall below the 25th percentile, 90% fall below the 90th percentile, and so on. The median is the 50th percentile. The interquartile range contains 50% of the numbers and is resistant to changes in extreme values. The standard deviation can be viewed as a kind of average or typical deviation from the mean. 106 2 There are no hard and fast rules about which statistic is the best. In general, the bigger the measures of spread are, the more the numbers are dispersed. 107 What Inferences Can Be Drawn from the Data? The inferences that may be drawn from a study depend on the design of the study and the quality of the data (see section titled “How Have the Data Been Collected?” above). The data might not address the issue of interest, might be systematically in error, or might be difficult to interpret because of confounding. Statisticians would group these concerns together under the rubric of “bias.” In this context, bias means systematic error, with no connotation of prejudice. We turn now to another concern, namely, the impact of random chance on study results. This contribution to total error may be called random error, sampling error, chance error, or statistical error. 108 If a pattern in the data is the result of chance, it is likely to wash out when more data are collected. By applying the laws of probability, a statistician can assess the likelihood that random error will create spurious patterns of certain kinds. Such assessments are often viewed as essential when making inferences from data. Thus, statistical inference typically involves tasks such as the following, which will be discussed in the rest of this guide. Estimation. Significance testing. p- p p- p- 109 Developing a statistical model. Computing posterior probabilities. 110 Key ideas of estimation and testing will be illustrated by courtroom examples, with some complications and mathematical details omitted for ease of presentation. 111 The first example, on estimation, concerns the Presidential Recordings and Materials Preservation Act of 1974, which impounded President Richard Nixon’s presidential papers after he resigned. 112 113 The Nixon papers were stored in 20,000 boxes at the National Archives in Alexandria, Virginia. It was plainly impossible to value this entire population of material. Appraisers for the plaintiff therefore took a random sample of 500 boxes. (From this point on, details are simplified; thus, the example becomes somewhat hypothetical.) The appraisers determined the fair market value of each sample box. The average of the 500 sample values turned out to be $2,000. The standard deviation (see section titled “Is an Appropriate Measure of Variability Used?” above) of the 500 sample values was $2,200. Many boxes had low appraised values, whereas some boxes were considered to be extremely valuable; this spread explains the large standard deviation. Estimation What Estimator Should Be Used? With the Nixon papers, it is natural to use the average value of the 500 sample boxes to estimate the average value of all 20,000 boxes comprising the population. With the average value for each box having been estimated as $2,000, the plaintiff demanded compensation in the amount of 20,000 × $2,000 = $40,000,000. In more complex problems, statisticians may have to choose among several estimators. Generally, estimators that tend to make smaller (or less costly) errors are preferred; however, “error” (or the losses that result from errors) might be quantified in more than one way. Moreover, the advantage of one estimator over another may depend on features of the population that are largely unknown, at least before the data are collected and analyzed. For complicated problems, professional skill and judgment may therefore be required when choosing a sample design and an estimator. In such cases, the choices and the rationale for them should be documented. What Is the Standard Error? An estimate based on a sample is likely to be off the mark, at least by a small amount, because of random error. The standard error gives the likely magnitude of this random error, with smaller standard errors indicating better estimates. 114 With a random sample of 500 boxes and a standard deviation of $2,200, the standard error for the sample average is estimated to be about $100. 115 116 How is the standard error to be interpreted? Just by the luck of the draw, a few too many high-value boxes may have come into the sample, in which case the estimate of $40,000,000 is too high. Or, a few too many low-value boxes may have been drawn, in which case the estimate is too low. This is random error. The net effect of random error is unknown, because data are available only on the sample, not on the full population. However, the net effect is likely to be something close to the standard error of $2,000,000. Random error throws the estimate off, one way or the other, by something close to the standard error. The role of the standard error is to gauge the likely size of the random error. The plaintiff’s argument may be open to a variety of objections, particularly regarding appraisal methods. However, the sampling plan is sound, as is the extrapolation from the sample to the population. And there is little need for a larger sample: Relative to the total claim of $40 million, a standard error of $2 million is reasonably small. What Is the Confidence Interval? Although random errors larger in magnitude than the standard error are commonplace, random errors larger in magnitude than two or three times the standard error are unusual. Confidence intervals make these ideas more precise. Usually, a confidence interval for the population average (the average of all the values in the population) is centered at the sample average; the desired confidence level is obtained by adding and subtracting a suitable multiple of the standard error. In dealing with large samples, statisticians who say that the population average falls within 1 standard error of the sample average will be correct about 68% of the time. Those who say “within 2 standard errors” will be correct about 95% of the time, and those who say “within 3 standard errors” will be correct about 99.7% of the time, and so forth. The normal curve and large samples These confidence levels correspond to areas under a famous bell-shaped curve—the normal curve. According to a fundamental theorem of statistics (the central limit theorem), if we were to draw not merely a single large sample from a much larger population, or even 500 of them as in the Nixon example, but millions upon millions of them (replacing the items from each sample before drawing the next one), and if we then sorted the sample averages into appropriate bins of a histogram, the heights of the bins would be greatest near the population average, and they would fall off symmetrically on each side as prescribed by the values of the normal curve centered at this point and having a standard deviation that is the standard error. 117 The confidence interval relies on this understanding of the distribution of sample means but goes in the other direction. Knowing how the sample statistic behaves as a function of the population parameters, we draw an interval around the sample statistic that we hope will cover the true value (the parameter). The mathematical properties of the normal distribution justify the numbers used to express the “confidence.” 118 To get a 68% confidence interval, start at the sample average, then add and subtract 1 standard error. To get a 95% confidence interval, start at the sample average, then add and subtract twice the standard error. To get a 99.7% confidence interval, start at the sample average, then add and subtract three times the standard error. With the Nixon papers, the 68% confidence interval for plaintiff’s total demand runs from $40,000,000 − $2,000,000 = $38,000,000 to $40,000,000 + $2,000,000 = $42,000,000. The 95% confidence interval runs from $40,000,000 − (2 × $2,000,000) = $36,000,000 to $40,000,000 + (2 × $2,000,000) = $44,000,000. The 99.7% confidence interval runs from $40,000,000 − (3 × $2,000,000) = $34,000,000 to $40,000,000 + (3 × $2,000,000) = $46,000,000. To write this more compactly, we abbreviate standard error as SE. Thus, 1 SE is one standard error, 2 SE is twice the standard error, and so forth. With a large sample and an estimate like the sample average, a 68% confidence interval ranges from estimate − 1 SE to estimate + 1 SE. A 95% confidence interval is ranges from estimate − 2 SE to estimate + 2 SE. A 99.7% confidence interval ranges from estimate − 3 SE to estimate + 3 SE. For a given sample size, increased confidence can be attained only by widening the interval. The 95% confidence level is the most popular, but some authors use 99%, and 90% is seen on occasion. (The corresponding multipliers on the SE are about 2, 2.6, and 1.6, respectively.) The phrase “margin of error” generally means twice the standard error. In medical journals, “confidence interval” is often abbreviated as “CI.” The relationship between width and confidence is shown in Figure 4 Figure 4 Confidence coefficients for CIs of ±1, 2, and 3 standard errors of a normally distributed estimator. The picture shows a tradeoff between precision (the width of the interval) and confidence (that the interval covers the actual population value). In sum, an estimate based on a sample will differ from the exact population value, because of random error. The standard error gives the likely size of the random error. If the standard error is small, random error probably has little effect. If the standard error is large, the estimate may be far away from the population value. Confidence intervals are a technical refinement that provides a formalism for assessing the impact of random error or chance variation on the estimate. Other situations Intervals based on the standard error, with confidence levels read off the normal curve, are appropriate for estimators that are essentially unbiased and obey the central limit theorem. They generally work for sums, averages, and rates, although much depends on the design of the sampling and other matters. CIs determined via the normal curve may not work well as estimates of extremely small quantities. Table 1 ad infinitum How should we proceed? It is easy to show that an interval from 0 to 6.5% would achieve at least 95% “confidence.” 119 n 120 In still other situations (for example, where observations are dependent), other methods can be used to obtain confidence intervals. One class of methods repeatedly resamples from the observed data to approximate what would happen in repeated samples from the full population. 121 122 How Big Should the Sample Be? There is no easy answer to this sensible question. Much depends on the level of error that is tolerable, the material being sampled, and the sampling method. Increasing the size of the sample provides no protection against bias (“nonsampling error”). Indeed, beyond some point, large samples are harder to manage and more vulnerable to nonsampling error. To reduce bias, the researcher must improve the design of the study or use a statistical model more tightly linked to the data-collection process. Larger samples generally will reduce the level of random error (“sampling error”). It rarely will be sensible to draw a probability sample with fewer than, say, two or three dozen items, and with such small samples, methods based solely on the normal curve (see section titled “What Is the Standard Error?” above) will not apply. A pilot sample is often valuable, partly to obtain an initial estimate of characteristics of the population that may be used in determining the final sample size. 123 Population size (i.e., the number of items in the population) usually has little bearing on the precision of estimates for the population average. This is surprising. On the other hand, population size has a direct bearing on estimated totals. Both points are illustrated by the Nixon papers (see section titled “What Is the Standard Error?” above). To be sure, drawing a probability sample from a large population may involve a lot of work. Samples presented in the courtroom have ranged from 5 (tiny) to 1.7 million (huge). 124 What Are the Technical and Interpretive Difficulties with Confidence Intervals? To begin with, “confidence” has an esoteric meaning in this context. The confidence level indicates the percentage of the time that intervals from repeated samples would cover the true value. The confidence level does not express the chance that repeated estimates would fall into the stated confidence interval. 125 Second, the confidence level does not give the probability that the unknown parameter lies within the confidence interval. 126 Third, this “confidence” is a statement about the entire region. One might think that values near the center of the interval are more likely to be closer to the true value than those at the extremes. Yet the statement of confidence applies equally to all the values in the interval. Moreover, that the intervals have sharp boundaries does not imply that there is an important difference between a value at the edge of the interval and one just beyond it. Fourth, for a given confidence level, a narrower interval indicates a more precise estimate, whereas a broader interval indicates less precision. 127 128 The final point to make is that standard errors and confidence intervals are often derived from statistical models for the process that generated the data. The model usually has parameters—numerical constants describing the population from which samples were drawn. When the values of the parameters are not known, the statistician must work backwards, using the sample data to make estimates. That was the case for valuing the Nixon papers. One parameter is the average value of all 20,000 boxes, and another parameter is the standard deviation of the 20,000 values. 129 If the data come from a probability sample or a randomized controlled experiment (see sections titled “Descriptive Surveys and Censuses” and “Is the Study Designed to Investigate Causation?” above), then the statistical model may be connected tightly to the actual data-collection process. In other situations, using the model may be tantamount to assuming that a sample of convenience is like a random sample, or that an observational study is like a randomized experiment. With the Nixon papers, the appraisers drew a random sample, and that justified the statistical calculations—if not the appraised values themselves. In many contexts, the choice of an appropriate statistical model is less than obvious. When a model does not fit the data-collection process, estimates and standard errors will not be probative. Standard errors and confidence intervals are designed to take account of random errors but not systematic ones such as selection bias or nonresponse bias (see sections titled “What Method Is Used to Select the Units?” and “Of the Units Selected, Which Provide Measurements?” above). For example, after reviewing studies to see whether a particular drug caused birth defects, a court observed that mothers of children with birth defects may be more likely to remember taking a drug during pregnancy than mothers with normal children. 130 131 p-values, Significance Levels, and Hypothesis Tests What Is the p-value? In 1969, Dr. Benjamin Spock came to trial in the U.S. District Court for Massachusetts. The charge was conspiracy to violate the Military Service Act. The jury was drawn from a panel of 350 persons selected by the clerk of the court. The panel included only 102 women—substantially less than 50%—although a majority of the eligible jurors in the community were female. The shortfall in women was especially poignant in this case: “Of all defendants, Dr. Spock, who had given wise and welcome advice on child-rearing to millions of mothers, would have liked women on his jury.” 132 Can the shortfall in women be explained by the mere play of random chance? To approach the problem, a statistician could formulate and test a null hypothesis. Here, the null hypothesis says that the panel is like 350 persons drawn at random from a large population that is 50% female. The expected number of women drawn would then be 50% of 350, which is 175. The observed number of women is 102. The shortfall is 175 − 102 = 73. How likely is it to find a disparity this large or larger, between observed and expected values? The probability is called the p- p The p p n θ 133 134 p 135 p Large p- p p p p- p- Because p p p- p p p 136 To recapitulate the logic of p p Computing p- 137 p- p- Is a Difference Statistically Significant? If an observed difference is in the middle of the distribution that would be expected under the null hypothesis, there is no surprise. The sample data are of the type that often would be seen when the null hypothesis is true. The observed difference (or similar statistic) does not fall into a region that is appropriate for rejecting the null hypothesis. It is not “significant,” as statisticians are wont to say. On the other hand, if the sample difference is far from the expected value—according to the null hypothesis—then the sample is unusual. The difference is significant, and the null hypothesis is rejected. Statistical significance can be determined by comparing p 138 p 139 p 140 Because the term “significant” is merely a label for a certain kind of p- p- 141 142 It is easy to mistake the p- p 143 p- Recent Emphasis on the Limitations of p-values For many of the reasons mentioned above, there has been considerable controversy about issues related to p 144 145 146 p p 147 p 148 Tests or Interval Estimates? How can a highly significant difference be practically insignificant? The reason is simple: p 149 p p- 150 p- A “significant” effect can be small. Conversely, an effect that is “not significant” can be large. By inquiring into the magnitude of an effect, courts can avoid being misled by p- Is the Sample Statistically Significant? Many a sample has been praised for its statistical significance or blamed for its lack thereof. Technically, this makes little sense. Statistical significance is about the difference between observations and expectations. Significance therefore applies to statistics computed from the sample, but not to the sample itself, and certainly not to the size of the sample. Findings can be statistically significant. Differences can be statistically significant (see section titled “Is a Difference Statistically Significant?” above). Estimates can be statistically significant (see section titled “Statistical Models” below). By contrast, samples can be representative or unrepresentative. They can be chosen well or badly (see section titled “What Method Is Used to Select the Units?” above). They can be large enough to give reliable results or too small to bother with (see section titled “What is the Confidence Interval?” above). But samples cannot be “statistically significant,” if this technical phrase is to be used as statisticians use it. Evaluating Hypothesis Tests What Is the Power of the Test? The power of a statistical study needs to be considered to avoid mistaking the absence of evidence for an effect with evidence for the absence of the effect. When a p- 151 When a study with low power fails to show a significant effect, the results may therefore be more fairly described as inconclusive rather than negative. The proof is weak because power is low. On the other hand, when studies have a good chance of detecting a meaningful association, failure to obtain significance can be persuasive evidence that there is nothing much to be found. 152 What About Small Samples? For simplicity, the examples of statistical inference discussed here (see sections titled “Estimation” and “ p p 153 1. It is hard to validate the assumptions underlying any proposed statistical approach. 2. Because approximations based on the normal curve generally cannot be used, suitable confidence intervals may be difficult to compute for parameters of interest. Likewise, p 154 3. Small samples may be unreliable, with large standard errors, broad confidence intervals, and tests having low power. One Tail or Two? Significance testing assesses the fit of the data to the null hypothesis within a given statistical model. In assessing whether the data fit the model, there is a choice to be made about whether to use a one-tailed or two-tailed p p or p either p- p- Some experts have argued for one or the other type of test, 155 p- How Many Tests Have Been Done? Repeated testing complicates the interpretation of significance levels. If enough comparisons are made, random error almost guarantees that some will yield “significant” findings, even when there is no real effect. To illustrate the point, consider the problem of deciding whether a coin is biased. The probability that a fair coin will produce 10 heads when tossed 10 times is (1/2) 10 Artifacts from multiple testing are commonplace. Because research that fails to uncover significance often is not published, reviews of the literature may produce a misleadingly large number of studies finding statistical significance. 156 157 Even a single researcher may examine so many different relationships that a few will achieve statistical significance by mere happenstance. For example, the researcher could choose a range of different outcome variables or different explanatory variables. Almost any large dataset—even pages from a table of random digits—will contain some unusual pattern that can be uncovered by diligent search. Having detected the pattern, the analyst can perform a statistical test for it, blandly ignoring the search effort. 158 159 There are statistical methods for dealing with multiple looks at the data, which permit the calculation of meaningful p- 160 161 What Are the Rival Hypotheses? The p- p- In Mapes Casino, Inc. v. Maryland Casualty Co 162 something 163 164 Bayesian Statistical Methods and Posterior Probabilities Standard errors, p- 165 10 But what of the converse probability: If the coin does land heads 10 times, what is the chance that it is fair? 166 p p 167 168 169 With the exceptions of parentage testing and so-called probabilistic genotyping software for interpreting complex DNA mixtures, 170 171 172 Correlation and Regression Regression models are used by many social scientists to infer causation from association. Such models have been offered in court to prove disparate impact in discrimination cases, to estimate damages in antitrust actions, and for many other purposes. The sections titled “Scatter Diagrams,” “Correlation Coefficients,” and “Regression Lines” cover some preliminary material, showing how scatter diagrams, correlation coefficients, and regression lines can be used to summarize relationships between variables. 173 Scatter Diagrams The relationship between two variables can be graphed in a scatter diagram (also called a scatterplot or scattergram). We begin with data on income and education for a sample of 238 men, ages 25 to 34, residing in Kansas. 174 Figure 5 Figure 5 Plotting points in a scatter diagram. Figure 6 Figure 6 Scatter diagram for income and education: men ages 25 to 34 in Kansas. Correlation Coefficients Two variables are positively correlated when their values tend to go up or down together, such as income and education in Figure 6 r Figure 7 r Figure 7 The correlation coefficient measures the sign of a linear association and its strength. A correlation coefficient of 0 indicates no linear association between the variables. The maximum value for the coefficient is +1, indicating that all the dots fall exactly on a straight line that slopes up. Sometimes, there is a negative association between two variables: Large values of one tend to go with small values of the other. The age of a car and its fuel economy in miles per gallon illustrate the idea. Negative association is indicated by negative values for r r Weak associations are the rule in the social sciences. In Figure 6 175 176 Is the Association Linear? The correlation coefficient has a number of limitations, to be considered in turn. The correlation coefficient is designed to measure linear association. Figure 8 Figure 8 A strong nonlinear association with a correlation coefficient close to zero. The correlation coefficient only measures the degree of linear association. Do Outliers Influence the Correlation Coefficient? The correlation coefficient can be distorted by outliers—a few points that are far removed from the bulk of the data. The left-hand panel in Figure 9 177 Figure 9 The correlation coefficient can be distorted by outliers. Does a Confounding Variable Influence the Coefficient? The correlation coefficient measures the association between two variables. Researchers—and the courts—are usually more interested in causation. Causation is not the same as association. The association between two variables may be driven by a lurking variable that has been omitted from the analysis (see section titled “Is the Study Designed to Investigate Causation?” above). For an easy example, there is an association between shoe size and vocabulary among schoolchildren. However, learning more words does not cause the feet to get bigger, and larger feet do not make children more articulate. In this case, the lurking variable is easy to spot—age. In more realistic examples, the lurking variable may be harder to identify. In statistics, lurking variables are called confounders or confounding variables. Association may reflect causation, but a large correlation coefficient is not enough to warrant causal inference. A large value of r 178 Regression Lines The regression line can be used to describe a linear trend in the data. The regression line for income on education in the Kansas sample is shown in Figure 10 Figure 10 The regression line for income on education and its estimates. Figure 11 Figure 6 Figure 11 Figure 11 Scatter diagram for income and education, with the regression line indicating the trend. What Are the Slope and Intercept? The regression line can be described in terms of its intercept and slope. Often, the slope is the more interesting statistic. In Figure 10 179 The slope of the regression line has the same limitations as the correlation coefficient: (1) The slope may be misleading if the relationship is strongly nonlinear; and (2) the slope may be affected by confounders and outliers. With respect to (1), the slope of $5,700 per year in Figure 10 What Is the Unit of Analysis? If association between characteristics of individuals is of interest, these characteristics should be measured on individuals. Sometimes individual-level data do not exist, but rates or averages for groups are available. “Ecological” correlations are computed from such rates or averages. These correlations generally overstate the strength of an association. For example, average income and average education can be determined for men living in each state and in Washington, D.C. The correlation coefficient for these 51 pairs of averages turns out to be 0.7. However, states do not go to school and do not earn incomes. People do. The correlation for income and education for men in the United States is only 0.4. The correlation for state averages overstates the correlation for individuals—a common tendency for ecological correlations. 180 Ecological analysis is often seen in cases claiming dilution in voting strength of minorities. In this type of voting rights case, plaintiffs must prove three things: (1) the minority group constitutes a majority in at least one district of a proposed plan; (2) the minority group is politically cohesive—that is, votes fairly solidly for its preferred candidate; and (3) the majority group votes sufficiently as a bloc to defeat the minority-preferred candidate. 181 The secrecy of the ballot box means that polarized voting cannot be directly observed. Instead, plaintiffs in voting rights cases rely on ecological regression, with scatter diagrams, correlations, and regression lines to estimate voting behavior by groups and demonstrate polarization. 182 Figure 12 183 184 Figure 12 Turnout rate for the white candidate plotted against the percentage of registrants who are white. Precinct-level data, 1982 Democratic primary for auditor, Lee County, South Carolina. Statistical Models Statistical models are widely used in the social sciences and in litigation. For example, the census suffers an undercount, more severe in certain places than others.If some statistical models are to be believed, the undercount can be corrected—moving seats in Congress and millions of dollars a year in tax5185 funds. 186 A regression model attempts to combine the values of certain variables (the independent variables) to get expected values for another variable (the dependent variable). The model can be expressed in the form of a regression equation. A simple regression equation has only one independent variable; a multiple regression equation has several independent variables. Coefficients in the equation will be interpreted as showing the effects of changing the corresponding variables. This is justified in some situations, as the next example demonstrates. Hooke’s law (named after Robert Hooke, England, 1653–1703) describes how a spring stretches in response to a load: Strain is proportional to stress. To verify Hooke’s law experimentally, a physicist will make a number of observations on a spring. For each observation, the physicist hangs a weight on the spring and measures its length. A statistician could develop a regression model for these data: length a b weight ε (1) The error term, denoted by the Greek letter ε a b 187 ε 188 Equation (1) has two parameters, a b a b a b 400 + (0.05 × 1) = 400.05. If the weight is 3, the expected length is 400 + (0.05 × 3) = 400 + 0.15 = 400.15. In either case, the actual length will differ from expected, by a random error ε In the simplest situation, the ε ε ε The parameters a b 189 â a b. â (2) Of course, no one really imagines there to be a box of tickets hidden in the spring. However, the variability of physical measurements (under many but by no means all circumstances) does seem to be remarkably like the variability in draws from a box. 190 Equation (1) is a statistical model for the data, with unknown parameters a b. ε â â a b ε These points apply to more complicated statistical models. The simple linear regression model exemplified in equation (1) can be extended in many ways. If another variable might contribute to or be associated with a response, an additional term ( constant variable d 191 192 In social science and legal applications, such advanced statistical models often are invoked—even when they lack an independent theoretical basis. Hooke’s law—incorporated in equation (1)—is relatively easy to test experimentally. For something like the salaries that an employer pays to workers, validation of a proposed relationship between the dependent variable and the independent ones would be difficult to validate experimentally. When expert testimony relies on statistical models, the court may well inquire, what are the assumptions behind the model, and why do they apply to the case at hand? The assumptions being questioned can include the choice of variables in the regression and the nature of the assumed relationship. It is important to distinguish between two situations: The nature of the relationship between the variables is known and regression is being used to make quantitative estimates of parameters in that relationship, and The nature of the relationship is largely unknown and regression is being used to determine the nature of the relationship—or indeed whether any relationship exists at all. Regression was developed to handle situations of the first type, with Hooke’s law being an example. The basis for the second type of application is analogical, and the tightness of the analogy is an issue worth exploration. 193 For most questions of legal interest, a wide variety of models can be used. This is only to be expected, because the science does not dictate specific equations and causal relationships. In a strongly contested case, each side will have its own model (series of models), presented by its own expert. The experts often reach opposite conclusions. The dialogue might continue with an exchange about which models produce admissible results, or more convincing ones. Although assumptions about the form of the model or the error component are challenged in court from time to time, arguments more commonly revolve around the choice of variables. One model may be questioned because it omits variables that arguably should be included—for example, skill levels or prior evaluations in an employment discrimination case. 194 195 196 197 The frequency with which regression models are used is no guarantee that they are the best choice for any particular problem. 198 199 Data Science and Statistical Machine Learning “Statistical science” has been described as “the discipline of learning about the world from data,” 200 201 202 203 204 205 206 What Is Data Science? “There is a wide variety of definitions and criteria for what constitutes data science,” making the term something of “a buzzword.” 207 208 209 Nonetheless, data science has emerged as a combination 210 What Is Machine Learning? “Machine learning” refers to methods for using data to classify objects or patterns into categories, and to produce predictions of future outcomes or events. The equations or procedures rely on computers. Of course, classification and prediction are tasks that humans do, either intuitively or with experience and expert training. Thus, machine learning (ML) is a subfield of “artificial intelligence” (AI), which seeks to build machines that perform tasks we associate with intelligent agents. 211 ML is closely related to statistics. Some methods that are packaged as ML or AI are nothing fundamentally new. 212 â 213 GPA a b LSAT a b ε a b a b Prediction by simple linear regression is so well established it is unlikely to be packaged as machine learning. That phrase is more likely to be encountered with procedures that depart from such explicit modeling and that are said to be “algorithmic.” 214 The regression model-based procedure for predicting GPA entails two parts that can be called algorithmic. First, there was an algorithm for finding the particular regression line. 215 â Because algorithms underlie all computation, it is not the existence of algorithms that makes ML distinctive. It is the willingness to develop algorithms for predictions or classifications without necessarily positing a statistical model of how the data arise. Statisticians historically relied on statistical models with a relatively small number of interpretable parameters (like the intercept a b 216 To convey the rough idea, let’s imagine that instead of imposing the linear regression model on the LSAT–GPA data, we used the following algorithm: (1) compute the average GPA for the students with LSATs ranging from 120–130, 130–140, and so on through 170–180; (2) plot these averages as heights at the midpoints of the intervals (125, 135, . . ., 175); (3) connect them to form a jagged line, and (4) use this jagged line to make predictions. Of course, one might ask why widths of 10 LSAT points should be used. Why not use shorter or longer segments for computing the means? An optimization procedure that reserves some of the data to try out different possibilities could help identify which width to use. 217 218 Our jagged line would not be used in practice. There are better procedures for fitting a curve to a cloud of data points. The slice-compute-and-connect algorithm is here only to illustrate the idea A number of important ML procedures do not yield any explicit equation (like the single regression line or the more complex jagged line) for making predictions or classifications. The two stages are done together, and the relevant variables are not necessarily pre-defined by the developer of the prediction model. Also, the features of the data that have the most influence on the output often are unknown. What Statistical Questions Arise with Machine Learning Studies? Although machine-learning advocates sometimes advertise that their methods make very few assumptions and provide great flexibility, they are still subject to the statistical and logical issues discussed in previous sections. Data quality, robustness, and transparency are a few of the issues that can arise. 219 220 Is the Dataset Appropriate and of Sufficient Quality? To begin with, as with traditional statistical methods, care is needed in the assembly of the data used to build and test the algorithms. Algorithms must be developed on data that are accurate and representative of the type of data to which they will be applied. For example, a study on diagnosing autism using ML excluded the data corresponding to borderline cases of autism. That made the test set unrepresentative of the general population; moreover, when a dataset that included the difficult-to-diagnose cases was used, the accuracy declined. 221 Is the Predictor or Classifier Robust? Second, robustness is also a vital issue. The goal of ML is not merely to find some complex pattern in the data; it is to discover generalizable 222 Is the Predictor or Classifier Too Opaque? Finally, ML algorithms can be difficult to interpret, especially when the patterns that are “learned” may not be represented in easily decoded forms. This can make it hard to tell when the data environment is changing in ways that will lead to poorer performance. It also can mask the fact that the algorithm has access to features that should not be part of the data. For example, when researchers rushed to apply ML methods to COVID-19 diagnosis from chest X-rays and CT scans, many of them unwittingly used a dataset that contained chest scans of children who did not have COVID as their examples of what non-COVID cases looked like. As a result, the ML diagnoses were relying on features identifying children rather than COVID. 223 224 225 Appendix: Conditional Probability and Bayes’ Rule What Do Probabilities Apply To? The mathematical theory of probability consists of theorems derived from axioms and definitions. Mathematical reasoning is seldom controversial, but there may be disagreement as to how the theory should be applied. For example, statisticians may differ on the interpretation of data in specific applications. Moreover, there are two main schools of thought about the foundations of statistics: frequentist and Bayesian (also called objectivist and subjectivist). 226 Frequentists see probabilities as empirical facts. When a fair coin is tossed, the probability of heads is taken to be 1/2; this value is justified by the argument that if the experiment is repeated a large number of times, the coin will land heads about one-half the time. If a fair die is rolled, the probability of getting an ace (one spot) is 1/6. If the die is rolled many times, an ace will turn up about one-sixth of the time. 227 228 What Are Conditional Probabilities? Conditional probability is the probability of one event given that another has occurred. For example, suppose a fair coin is tossed twice. One event is that the coin will land heads up both times (HH). Another event is that at least one H will be seen. Before the coin is tossed, there are four possible, equally likely, outcomes: HH, HT, TH, TT. So the probability of HH is 1/4. However, if we know that at least one head has been obtained, then we can rule out two tails TT. In other words, given that at least one H has been obtained, the conditional probability of TT is 0, and the first three outcomes have conditional probability 1/3 each. In particular, the conditional probability of HH is 1/3. This is usually written as P(HH|at least one H) = 1/3. More generally, the probability of an event C is denoted P(C); the conditional probability of D given C is written as P(D|C). Two events C and D are independent if the conditional probability of D given that C occurs is equal to the conditional probability of D given that C does not occur. Using ~C to denote the event that C does not occur, C and D are independent if P(D|C) = P(D|~C). If C and D are independent, then the probability that both occur is equal to the product of the probabilities: P(C and D) = P(C) × P(D) (3) This is the multiplication rule (or product rule) for independent events. If events are dependent, then conditional probabilities must be used: P(C and D) = P(C) × P(D|C) (4) This is the multiplication rule for dependent events. Statisticians caution against, but sometimes succumb to, using the multiplication rule for independent events when the events are dependent. 229 What Is Bayes’ Rule? If we use probabilities to describe uncertainty regarding hypotheses as well as to describe events, we can use the formulas for conditional probabilities to obtain an equation known as Bayes’ rule that has been proposed to guide legal factfinding. If two mutually exclusive hypotheses H 0 H 1 A 230 231 (5) This is one way of expressing Bayes’ rule. It yields the conditional probability of hypothesis H 1 A For a stylized example in a criminal case, suppose that blood from a single source is found at the scene of a crime and that the defendant in the case has type A blood. It is natural to have one hypothesis H 0 H 1 A H 0 H 0 H 0 A Type A blood occurs in 42% of the population. Assuming that the laboratory’s findings are always correct, P( A H 0 232 A A H 1 H 0 H 1 (10) Thus, the data increase the probability that the blood is the defendant’s. The probability went up from the prior value of P H 1 H 1 A Equation (5) can be rewritten as follows: (11) This form gives us a convenient way to express the idea that the data change the chances in favor of H 1 H 0 H 1 H 0 233 234 Posterior odds Likelihood ratio Prior odds (12) where the “likelihood ratio” is the probability of the data (the crime-scene blood is type A) under the hypotheses H 1 H 0 H 1 H 0 Used with Bayes’ rule, this likelihood ratio means that the data increase the prior odds by a factor of a little more than 2, from 1:1 to about 2.31:1. Odds of 2.31 to 1 are the same as a probability of 2.31/(2.31 + 1) = 0.70. Equations (5), (11), and (12) are all forms of the same fundamental relationship. In the context of Bayes’ rule, the likelihood ratio is called a Bayes’ factor. 235 change 236 The same ideas can be applied more broadly to the various statistical models that have been illustrated in this reference guide. There are Bayesian approaches to fitting regression or other statistical models to make predictions and to obtain estimates of population parameters. 237 238 Glossary of Terms The following definitions are adapted from a variety of sources, including Michael O. Finkelstein & Bruce Levin, Statistics for Lawyers Statistics The Art of Statistics: How to Learn from Data absolute value. adjust for. alpha ( ). alternative hypothesis. area sample. arithmetic mean. average. Bayes factor. Bayes’ rule. Bayesian inference. beta ( ). between-observer variability. bias. bias-variance trade-off big data bin. binary variable. binomial distribution. blind. bootstrap. categorical data; categorical variable. central limit theorem. chance error. chi-squared ( 2 ). class interval. cluster sample. coefficient of determination. R R coefficient of variation. collinearity. conditional probability. confidence coefficient. confidence interval. confidence level. confounding variable; confounder. consistent estimator. content validity. continuous variable. control for. control group. controlled experiment. convenience sample. correlation coefficient. r covariance. covariate. credible interval. criterion. data. data science. degrees of freedom. t dependence. dependent variable descriptive statistics. differential validity. discrete variable. distribution. disturbance term. double-blind experiment. dummy variable. econometrics. epidemiology. error term. estimator. expected value. which is equal to 7. experiment explanatory variable. external validity. factors. See independent variable. false negative. false positive. Fisher’s exact test. p- p fitted value. See residual. fixed significance level. p p frequency; relative frequency. frequency distribution. frequentist. Gaussian distribution. general linear model. grab sample. heteroscedastic. highly significant. p- histogram. homoscedastic. hypergeometric distribution. hypothesis. hypothesis test. p identically distributed. independence. independent variable. indicator variable. internal validity. interquartile range. interval estimate. least sq least squares estimator. level. p likelihood. likelihood ratio. linear combination. u v u v linear regression. list sample. logistic regression. loss function. lurking variable. Markov chain. mean. measurement validity. median. meta-analysis. metrology. mode. model. Monte Carlo method. multicollinearity. multiple comparison p- p multiple correlation coefficient. R R multiple regression. multistage cluster sample. multivariate methods. natural experiment. nonparametric method. nonresponse bias. nonsampling error. normal distribution. null hypothesis. observational study. observed significance level. p- odds. odds ratio. one-sided hypothesis; one-tailed hypothesis. one-sided test; one-tailed test. overfitting. outcome variable. outlier. p p- p- p p p p p- parameter. percentile. placebo. point estimate. Poisson distribution. population. population size. posterior probability power. p- practical significance. practice effects. predicted value. predictive validity. predictor. prior probability. probability. probability density. probability distribution. probability histogram. probability model. probability sample. psychometrics. qualitative variable; quantitative variable. quartile. quasi-experiment. R R 2 ). R R random error. random variable. randomization. randomized controlled experiment. range. rate. regression coefficient. regression diagnostics. regression equation. regression line. regression model. relative frequency. relative risk. reliability. repeatability. representative sample. reproducibility. resampling. residual. response variable. risk factor. robust. sample. sample size. sample weights. sampling distribution. sampling error. sampling frame. sampling interval. scatter diagram. selection bias. sensitivity. sensitivity analysis. sign test. p p significance level. p significance test. p p α p p p- p- significant. p simple random sample. n N n simple regression. size. skip factor. specificity. α α spurious correlation. standard deviation (SD). standard error (SE). standard error of regression. R standardization. standardized variable. statistic. statistical controls. statistical dependence. statistical hypothesis. statistical independence. statistical model. statistical significance. p- statistical test. stratified random sample. stratification. study validity. subjectivist. systematic error. systematic sample. k k k k t t t t t t- t t t t t t t t p test statistic. 2 t p t time series. treatment group. two-sided hypothesis; two-tailed hypothesis. two-sided test; two-tailed test. Type I error. Type II error. unbiased estimator. uniform distribution. validity. Study validity is the extent to which results from a study can be relied upon. Study validity has two aspects, internal and external. A study has high internal validity when its conclusions hold under the particular circumstances of the study. A study has high external validity when its results are generalizable. For example, a well-executed randomized controlled double-blind experiment performed on an unusual study population will have high internal validity because the design is good; but its external validity will be debatable because the study population is unusual. Validity is used also in its ordinary sense: assumptions are valid when they hold true for the situation at hand. variable. variance. weights. within-observer variability. z z z z t z t z z z z z t- z- t z z z z The z p- References on Statistics and Research Methods Nontechnical Surveys Adam Chilton & Kyle Rozema, Trial by Numbers: A Lawyer’s Guide to Statistical Evidence (2024). David Freedman et al., Statistics (4th ed. 2007). Darrell Huff, How to Lie with Statistics (1993). Gregory A. Kimble, How to Use (and Misuse) Statistics (1978). Ethan Bueno de Mesquita & Anthony Fowler, Thinking Clearly with Data: A Guide to Quantitative Reasoning and Analysis (2021). David S. Moore & William I. Notz, Statistics: Concepts and Controversies (10th ed. 2019). Michael Oakes, Statistical Inference: A Commentary for the Social and Behavioral Sciences (1986). Statistics: A Guide to the Unknown (Roxy Peck et al. eds., 4th ed. 2005). David Spiegelhalter, The Art of Statistics: Learning from Data (2019). Hans Zeisel, Say It with Figures (6th ed. 1985). General References Encyclopedia of Statistical Sciences (Samuel Kotz et al. eds., 2d ed. 2005). The Oxford Handbook of Quantitative Methods (Todd D. Little ed., 2014). Best Practices in Quantitative Methods ( Jason W. Osborne ed., 2008). Footnotes 1 See generally Id. 2 509 U.S. 579, 589–90 (1993). 3 See 4 Daubert 5 Of course, the case could be presented in the form of interlocking testimony from two experts—the labor economist followed by any expert with the appropriate expertise in the method for making the statistical comparison. The substantive knowledge may be of lesser value in selecting a statistical model. Various models might be consistent with the substantive knowledge, and there may be purely statistical criteria for choosing among them. 6 Academic experts are likely to have a long list of publications, but length alone is not the best indication of pertinent expertise. Not all publishers are equal, and every field has a relatively small number of top-tier journals. 7 585 P.2d 563 (Ariz. 1978). 8 Id The New Wigmore: A Treatise on Evidence: Expert Evidence 9 See 10 Id cf Litigation Science After the Knowledge Crisis Selective Inference: The Silent Killer of Replicability https://doi 11 ASA Comm. on Professional Ethics, Ethical Guidelines for Statistical Practice, Feb. 2022, at 3 & 4, https://perma supra 12 See reprinted in supra Improving Legal Statistics 13 Drawing on “open science” practices, commentators have proposed adaptations of practices for “registration” of academic research studies. See supra p 14 For introductory treatments of data collection, see, for example, David Freedman et al., Statistics Statistics: Concepts and Controversies Say It with Figures Prove It with Figures: Empirical Methods in Law and Litigation 15 See also Reference Guide on Epidemiology 16 Consequently, many courts have suggested that attempts to infer causation from anecdotal reports are inadmissible as unsound methodology under Daubert v. Merrell Dow Pharmaceuticals, Inc. See, e.g. ex rel. cf E.g. 17 See 18 Richard Doll & A. Bradford Hill, A Study of the Aetiology of Carcinoma of the Lung https://doi See supra 19 Epidemiologists sometimes limit “confounding” to “a bias due to the existence of a common cause of exposure and outcome” and define “selection bias” as “bias by [selecting units for study based] on common effects of otherwise unrelated variables.” Stephen R. Cole et al., Illustrating Bias Due to Conditioning on a Collider https://doi 20 For additional examples and further discussion, see David A. Freedman, From Association to Causation: Some Remarks on the History of Statistics 21 Randomization of subjects to treatment or control groups puts statistical tests of significance on a secure footing. See “ 22 One important feature of a clinical trial is “blinding” to prevent unconscious bias. In a “double blind” trial, neither the patients nor the clinicians ascertaining the outcomes know who received the treatment and who did not. This prevents the expectations of the subjects and the experimenters from systematically affecting the observed outcomes in one group relative to the other. 23 But randomized experiments also show that aspirin can cause internal bleeding, raising the practical questions of whether and when a low-dose regimen is more beneficial than harmful. See, e.g. https://perma E.g. Vitamin E and the Risk of Prostate Cancer: Results of the Selenium and Vitamin E Cancer Prevention Trial (SELECT) https://doi 24 See 25 Multiple regression is described in Daniel L. Rubinfeld & David Card, Reference Guide on Multiple Regression and Advanced Statistical Models Regression Analysis: A Constructive Critique What You Can and Can’t Properly Do with Regression https://doi Statistical Models: Theory and Practice 26 See, e.g. Conceptualising Natural and Quasi Experiments in Public Health https://doi 27 John Snow, On the Mode of Transmission of Cholera 55–98 (2d ed. 1855). 28 Id 29 For example, case-control studies are designed one way and cohort studies another, with many variations. See, e.g. Gordis Epidemiology supra 30 The idea is to control for the influence of a confounder by stratification—making comparisons separately within groups for which the confounding variable is nearly constant and therefore has little influence over the variables of primary interest. For example, smokers are more likely to get lung cancer than nonsmokers. Age, gender, social class, and region of residence are all confounders, but controlling for such variables does not materially change the relationship between smoking and cancer rates. 31 A. Bradford Hill, The Environment and Disease: Association or Causation? 32 At best, statistical analysis can help address the question of how impactful an unknown confounding variable would have to be to vitiate an inference of causation. See Smoking and Lung Cancer: Recent Evidence and a Discussion of Some Questions https://doi Are Greenland, Ioannidis and Poole Opposed to the Cornfield Conditions? A Defence of the E-value https://doi 33 David A. Freedman & Hans Zeisel, From Mouse to Man: The Quantitative Assessment of Cancer Risks https://doi Linear Low-Dose Extrapolation for Noncancer Health Effects Is the Exception, Not the Rule https://doi Pre-clinical Animal Models Are Poor Predictors of Human Toxicities in Phase 1 Oncology Clinical Trials https://doi Animal Study Translation: The Other Reproducibility Challenge https://doi Extrapolating from Animals to Humans https://doi Is It Possible to Overcome Issues of External Validity in Preclinical Animal Research? Why Most Animal Models Are Bound to Fail https://doi 34 Phoebe C. Ellsworth, Some Steps Between Attitudes and Verdicts, in Lockhart v. McCree 35 See Reference Guide on Survey Research 36 The U.S. Decennial Census does not count everyone that it should, and it counts some people who should not be counted. There is evidence that net undercount is greater in some demographic groups than others. Supplemental studies may enable statisticians to adjust for errors and omissions. See See Statistical Controversies in Census 2000 Methods for Census 2000 and Statistical Adjustments, in 37 On the admissibility of such polls, see State v. Midwest Pride IV, Inc. 38 Survey researchers may turn to other methods to reach cell phone users with area codes from the jurisdiction’s geographic locale. Kyley McGeeney & Courtney Kennedy, Pew Research Center, Advances in Telephone Survey Sampling (2015), https://perma See Sample and Respondent Provided County Comparisons Among Cellular Respondents Using Rate Center Assignments https://doi 39 In re In re . See infra 40 See supra 41 Id 42 See id 43 Many types of probability sampling have been developed. In simple random sampling, units are drawn at random without replacement. In particular, each unit has the same probability of being chosen for the sample. Id. n 44 See Statistical Proof, in 45 Maurice C. Bryson, The Literary Digest Poll: Making of a Statistical Myth supra 46 For discussions of the admissibility of surveys with extremely low response rates, see In re ConAgra Foods, Inc. United States v. H & R Block, Inc. Survey Response Rates: Rapid Literature Review https://perma In United States v. Gometz Gometz Id. cf. In re 47 See Shere Hite and the Trouble with Numbers 48 E.g. 49 509 U.S. 579, 590 n.9 (1993). 50 In practice, more refined measures of internal consistency are used to estimate a test’s reliability. See, e.g. High Stakes Test Construction and Test Use, in 51 572 U.S. 701 (2014). 52 Id 53 “Standard error” is the subject of the section titled “Is a Difference Statistically Significant?” below. The relationship between the reliability coefficient and different standard errors as well as the choice of a cut-off score that accounts for a given type of standard error is explained in David H. Kaye, Deadly Statistics: Quantifying an “Unacceptable Risk” in Capital Punishment 54 See 55 Id. Reference Guide on Human DNA Identification Evidence 56 David C. Baldus et al., Equal Justice and the Death Penalty: A Legal and Empirical Analysis 49–50 (1990). 57 Alan H. Dorfman & Richard Valliant, A Re-Analysis of Repeatability and Reproducibility in the Ames-USDOE-FBI Study https://doi Reliability and Validity of Forensic Science Evidence 58 See 59 As the discussion of the correlation coefficient in the section titled “Correlation Coefficients” below indicates, the closer the coefficient is to 1, the greater the validity. For a review of data on test reliability and validity, see Measuring Success: Testing, Grades, and the Future of College Admissions 60 See, e.g. 61 See “ Hall v. Florida 62 For an extreme example, consider a firm that provides personal tutoring for college-entrance examinations and hires only job applicants who had near-perfect scores themselves to be the tutors. It could well be that receiving a high score is associated with being an effective tutor. But if the firm hired no low-scoring applicants as tutors, it would be impossible to see that association in a study of the correlation between the scores of the exclusively high-scoring tutors and those of the students they tutor after being hired. 63 See Modeling Selection Effects, in True Score Theory: The Traditional Method, in 64 The numbers are rounded-off versions of those in Table C1 of Heike Hofmann et al., Treatment of Inconclusives in the AFTE Range of Conclusions https://doi See 65 E.g. Erratum https://doi FBI Notifies Crime Labs of Errors Used in DNA Match Calculations Since 1999 66 See, e.g. aff’d cf 67 See generally supra supra supra supra 68 Lyda Longa, Increase in Thefts Alarming 69 Karen Langley, Stock Pickers Watched the S&P 500 Pass Them By Again in 2021 70 The selection of the benchmark may merit scrutiny. Securities and Exchange Commission (SEC) Rule 33–698 requires mutual funds to display past returns alongside those of “an appropriate broad-based securities market index.” But the rule “does not prohibit funds from comparing their past returns to those of newly-chosen index(es).” Kevin Mullally & Andrea Rossi, Moving the Goalposts? Mutual Fund Benchmark Changes and Performance Manipulation https://perma Id 71 James P. Levine et al., Criminal Justice in America: Law in Action 99 (1986) (referring to a change from 1959 to 1960). 72 David Seidman & Michael Couzens, Getting the Crime Rate Down: Political Pressure and Crime Reporting 73 John A. Eterno & Eli B. Silverman, The Crime Numbers Game: Management by Manipulation (2012); Michael D. Maltz, Missing UCR Data and Divergence of the NCVS and UCR Trends in https://doi E.g. https://perma 74 511 F. Supp. 855 (S.D.N.Y. 1980). 75 511 F. Supp. 867 (S.D.N.Y. 1980). 76 Philip Morris 77 Id 78 Id 79 Only two women in the sample were not pregnant; the test gave correct results for both of them (a specificity of 100%). Although a false-positive rate of 0 is ideal, an estimate based on a sample of only two women is not. These data are reported in Arnold Barnett, How Numbers Can Trick You See Id 80 The numbers are similar to those in Hofmann et al., supra 81 See 82 Some scientists and litigants maintain that, by definition, every inconclusive result is an error (and should be counted as such) because it does not express the true association (positive or negative) within each pair. Id See, e.g. Inconclusives and Error Rates in Forensic Science: A Signal Detection Theory Approach https://doi See Inconclusive Conclusions in Forensic Science: Rejoinders to Scurich, Morrison, Sinha and Gutierrez https://doi 83 We discuss sampling error in false-positive proportions in the section titled “What Is the Standard Error?” below and describe a more complete measure of probative value in the Appendix section titled “What Is Bayes’ Rule?” 84 For assistance in coping with percentages, see Zeisel, supra 85 Correlation and regression are discussed in the section titled “Correlation and Regression” below. 86 A table of this sort is called a “cross-tab” or a “contingency table.” Table 5 is “two-by-two” because it has two rows and two columns, not counting rows or columns containing totals. 87 E.g. Statistical Evidence of Discrimination in Jury Selection, in 88 A procedure that selects candidates from the least successful group at a rate less than 80% of the rate for the most successful group “will generally be regarded by the Federal enforcement agencies as evidence of adverse impact.” EEOC Uniform Guidelines on Employee Selection Procedures, 29 C.F.R. § 1607.4(D) (1978). The rule is designed to help spot instances of substantially discriminatory practices, and the commission usually asks employers to justify any procedures that produce selection ratios of 80% or less. 89 See supra 90 For women, the odds of rejection are 99 to 1; for men, 19 to 1. The ratio of these odds is 99:19. Likewise, the odds ratio for an admitted applicant being a man as opposed to a denied applicant being a man is 99:19. 91 But see Odds Ratios as a Measure of Disproportionate Treatment: Application to Jury Venires https://doi 92 Tables 5 and 6 are hypothetical, but closely patterned on a real example. See Sex Bias in Graduate Admissions: Data from Berkeley https://doi 93 See generally 94 Florida Statistical Analysis Center, Florida Dep’t of Law Enforcement, Statewide Reported Domestic Violence Offenses in Florida, 1992–2020, https://perma 95 See 96 The coin landed heads 7 times in the first 10 tosses; by coincidence, there were also 7 heads in the next 10 tosses; there were 5 heads in the third batch of 10 tosses; and so forth. 97 In Figure 3 All the bins in Figure 3 See supra 98 Technically, at least half the numbers are at the median or larger; at least half are at the median or smaller. When the distribution is symmetric, the mean equals the median. The values diverge, however, when the distribution is asymmetric, or skewed. 99 Herbert M. Kritzer et al., An Exploration of “Noneconomic” Damages in Civil Jury Awards cf Punitive Damages in Securities Arbitration: An Empirical Study TXO Production Corp. v. Alliance Resources Corp The Supreme Court and Junk Social Science: Selective Distortion in Amicus Briefs 100 E.g. In re Educational Testing Service Praxis Principles of Learning and Teaching, Grades 7–12 Litigation 101 To get the total award, just multiply the mean by the number of awards; by contrast, the total cannot be computed from the median. (The more pertinent figure for the insurance industry is not the total of jury awards, but actual claims experience including settlements; of course, even the risk of large punitive damage awards may have considerable impact.) 102 554 U.S. 471 (2008). 103 Id 104 According to the Court, “the outlier cases subject defendants to [disproportionate] punitive damages,” id Id See Variability in Punitive Damages: Empirically Assessing Punitive Damages: From Myth to Theory 105 554 U.S. at 513. 106 When the distribution follows the normal curve, about 68% of the data will lie within plus-or-minus 1 standard deviation of the mean; about 95% will lie within 2 standard deviations of the mean. For other distributions, the proportions will be different. 107 In Exxon Shipping Co. v. Baker Id Id See The Predictability of Punitive Damages 108 Econometricians use the parallel concept of random disturbance terms. See supra p p 109 As explained in the section titled “ p p p But see On the Origins of the .05 Level of Statistical Significance By contrast, formal hypothesis testing (as developed by Jerzy Neyman and Egon Pearson) requires explicit specification of an alternative hypothesis and then development of a test procedure with pre-specified probabilities of making two types of errors—type I (rejecting the null hypothesis when it is true) and type II (failing to reject the null hypothesis when the alternative is true). The end result here is a decision rather than a p See, e.g. The Fisher, Neyman-Pearson Theories of Testing Hypotheses: One Theory or Two? https://doi 110 The theorem is named after the Reverend Thomas Bayes (England, c. 1701–1761). An elementary version of the rule is derived in the Appendix. Bayes’ essay on the subject was published after his death: An Essay Toward Solving a Problem in the Doctrine of Chances Statistical Inference: A Likelihood Paradigm Some Issues in the Foundation of Statistics reprinted in Comparing Philosophies of Statistical Inference, in What Is Bayesianism? in reprinted in 111 Some technical details appear in the appendices to David H. Kaye & David A. Freedman, Reference Guide on Statistics, in 112 44 U.S.C. § 2111. 113 Nixon v. United States, 978 F.2d 1269 (D.C. Cir. 1992); Griffin v. United States, 935 F. Supp. 1 (D.D.C. 1995). 114 We distinguish between (1) the standard deviation of the sample data, which measures the spread in the sample values, and (2) the standard error of the sample average, which measures the likely size of the random error in the sample average. Generally, the standard error of an estimator (such as the sample average) is the standard deviation of that estimator as applied to (hypothetically) repeated samples. Courts typically use the broader term “standard deviation” when referring to the standard error. E.g 115 The standard error for the sample average equals Freedman et al., supra N n N 116 We are assuming a simple random sample. Generally, the formula for the standard error must take into account the method used to draw the sample and the nature of the estimator. In fact, the Nixon appraisers used more elaborate statistical procedures. Moreover, they valued the material as of 1995, extrapolated backward to the time of taking (1974), and then added interest. After 20 years of litigation, Nixon’s estate settled with the government for $18 million. U.S. Dep’t of Just., Press Release, June 12, 2000, https://perma 117 The normal curve is the density of a normal distribution. The distribution has two parameters—the population mean (often denoted µ σ x where e µ σ 118 The area under the normal curve between x µ σ x µ σ µ σ µ σ 119 When the long-run false-positive error rate is 6.5%, the probability of 45 independent correct classifications of cartridges from different pairs of guns is (1−0.065) 45 120 E.g. Bayesian Analysis of Misclassified Binomial Data: Double-Sampling and the Zero-Numerator Problem https://doi A Look at the Rule of Three The Role of Informative Priors in Zero-Numerator Problems: Being Conservative Versus Being Candid 121 See, e.g. supra Frequentist Statistical Inference, in SAP America v. InvestPic Cybergenetics Corp. v. Inst. of Env’l Sci. & Rsch. aff’d 122 See In re supra 123 Suppose a researcher is interested in the association between alcohol intoxication and single-vehicle accidents. A small sample of accident records (along with intuitions) could suggest some guess for the population proportion of single-vehicle crashes in which the driver was found to be intoxicated. Assume the guess is 50%. Suppose further that the researcher is willing to tolerate errors of up to ±3% in the estimate from the larger sample being planned. Finally, suppose that the researcher plans to report a conventional 95% confidence interval. It can be shown that a simple random sample of approximately 1,067 single-vehicle accident records should suffice. Hence, the researcher collects 1,067 records at random and determines the proportion in which intoxication was in the accident record. The observed proportion in this sample might be 40% or 0.4, to pick a concrete number. The earlier guess of 50% no longer matters—it was just used to get a sense of the sample size for the full study, and the researcher now can report a 95% CI from the full study. The estimate becomes. 124 Lebrilla v. Farmers Grp., Inc., No. 00-CC-017185 (Cal. Super. Ct., Orange Cnty., Dec. 5, 2006) (preliminary approval of settlement). This was a class action lawsuit on behalf of plaintiffs who were insured by Farmers and had automobile accidents. Plaintiffs alleged that replacement parts recommended by Farmers did not meet specifications: Small samples were used to evaluate these allegations. At the other extreme, it was proposed to adjust Census 2000 for undercount and overcount by reviewing a sample of 1.7 million persons. See supra 125 Opinions reflecting this misinterpretation include Turpin v. Merrell Dow Pharm., Inc. Garcia v. Tyson Foods, Inc. In re Silicone Gel Breast Implants Prods. Liab. Litig. Language from another reference guide in the previous edition of this Reference Manual See, e.g. Confidence Intervals and Replication: Where Will the Next Mean Fall? https://doi Silicone Gel 126 See, e.g. supra Center for Biological Diversity v. U.S. Fish & Wildlife Serv. Cnty. of Douglas v. Nebraska Tax Equalization & Rev. Comm’n 127 See Cimino v. Raymark Industries, Inc. rev’d Id. Id 128 In Hilao v. Estate of Marcos Id. Ayyad v. Sprint Spectrum, L.P. 129 These parameters can be used to approximate the distribution of the sample average. See supra 130 Brock v. Merrell Dow Pharm., Inc., 874 F.2d 307, 311–12 (5th Cir.), modified 131 In Brock 132 Hans Zeisel, Dr. Spock and the Case of the Vanishing Women Jurors 133 In mathematical notation, the probability for the number x n f x; n, θ n x n x θ n θ n f supra θ See, e.g. Ruminations on Jurimetrics: Hypergeometric Confusion in the Fourth Circuit 134 The binomial probability of getting exactly 102 women in the sample is f x n θ 15 x 15 15 x p 15 135 See supra 136 Some opinions present a contrary view. E.g. p Instances of the transposition fallacy in criminal cases are collected in Kaye et al., supra McDaniel v. Brown See The Interpretation of DNA Evidence: A Case Study in Probabilities, in https://www 137 The binomial analysis in the Spock nθ θ infra p See 138 It is not necessary to compute the p α α α α p α p α 139 The Supreme Court implicitly referred to this practice in Castaneda v. Partida Hazelwood School District v. United States p- 140 Some opinions quote statisticians who urge giving readers p p E.g., In re p p p p p 141 Context affects judgments of what outcomes lack practical significance. Some relatively small differences can have a large practical effect. For example, a loss of half a percentage point in the revenue of a corporation that is a consequence of a competitor’s illegal conduct would be practically significant in computing damages if the total revenue was large. 142 E.g. id p cf. infra But see 143 E.g., Waisome cf p p 144 John P.A. Ioannidis, Why Most Clinical Research Is Not Useful https://doi Psychology, Science, and Knowledge Construction: Broadening Perspectives from the Replication Crisis https://doi supra 145 Valentin Amrhein et al., Retire Statistical Significance https://doi 146 In a special issue of The American Statistician Moving to a World Beyond “p < 0.05 ” https://doi Statistical Significance, P-values, and Replicability https://doi Nature supra 147 Surprisingly, when p p E.g. Replication Power and Regression to the Mean Solutions for Quantifying P-value Uncertainty and Replication Power https://doi 148 Ronald L. Wasserstein & Nicole A. Lazar, ASA Statement on Statistical Significance and P-values https://dx when properly applied and interpreted ASA President’s Task Force Statement on Statistical Significance and Replicability https://doi 149 See p- E.g. 150 Cf. Washington 151 More precisely, power is the probability of rejecting the null hypothesis when the alternative hypothesis ( see α α α Statisticians usually denote power by the Greek letter beta ( β β accepting The chance of a false negative may be computed from the power. Some commentators have claimed that the cutoff for significance should be chosen to equalize the chance of a false positive and a false negative, on the ground that this criterion corresponds to the more-probable-than-not burden of proof. The argument is fallacious, because 1– α β See p Hypothesis Testing in the Courtroom, in 152 Some formal procedures (meta-analysis) are available to aggregate results across studies. See, e.g., In re See, e.g. Statistical Assumptions as Empirical Commitments, in 153 Advocates sometimes contend that samples are “too small to allow for meaningful statistical analysis,” United States v. New York City Bd. of Educ. Id E.g. see generally 154 With large samples, approximate inferences (based on the central limit theorem, for example) may be quite adequate. These approximations will not be satisfactory for small samples unless the sampled values (not just the sample means) are approximately normally distributed. 155 See, e.g. But see supra 156 E.g. Publication Bias in Clinical Research https://doi Effect of the Statistical Significance of Results on the Time to Completion and Publication of Randomized Efficacy Trials Statistical Problems in the Reporting of Clinical Trials: A Survey of Three Medical Journals https://doi 157 See 158 The practice has been denominated HARKing. Norbert L. Kerr, HARKing: Hypothesizing After the Results Are Known https://doi 159 Searching for significance in this way has been called data mining, data dredging, data snooping, selection bias, and, more recently, p The Extent and Consequences of P-Hacking in Science https://doi 160 See, e.g. supra Points of Significance: Comparing Samples—Part II Statistical Power and Significance Testing in Large-scale Genetic Studies Karlo v. Pittsburgh Glass Works Daubert Id see Multiple Hypothesis Testing in https://perma 161 ASA Comm. on Professional Ethics, supra 162 290 F. Supp. 186 (D. Nev. 1968). 163 Id. Id. Compare with cf. “ 164 E.g p 165 Operating characteristics include the expected value and standard error of estimators, probabilities of error for statistical tests, and the like. 166 We call this a converse probability because it is of the form P( H 0 H 0 H 0 H 0 supra p- H 0 H 0 167 Bayesian procedures are sometimes defended on the ground that the beliefs of any rational observer must conform to the Bayesian rules. However, the definition of “rational” is purely formal. See The Axioms of Subjective Probability Some Issues in the Foundation of Statistics reprinted in The Laws of Probability and the Law of the Land 168 Here, confidence has the meaning ordinarily ascribed to it, rather than the technical interpretation applicable to a frequentist confidence interval. Consequently, it can be related to the burden of persuasion. See Apples and Oranges: Confidence Coefficients and the Burden of Persuasion 169 But see infra 170 See supra 171 Donald B. Rubin, Estimation in Parallel Randomized Experiments https://doi 172 See Administering Section 2 the Voting Rights Act After Forensic Statistics in the Courtroom, in supra Rounding Up the Usual Suspects: A Legal and Logical Analysis of DNA Database Trawls See, e.g. Wigmore, supra Digging into the Foundations of Evidence Law 173 The focus is on simple linear regression. See also supra 174 These data are from a public-use data file of the 2021 American Community Survey and were obtained from the data https://data 175 Law School Admissions Council, Summary of 2017, 2018, and 2019 LSAT Correlation Study Results, https://perma Id 176 G. Frederiek Estourgie-van Burk et al., Body Size in Five-year-old Twins: Heritability and Comparison to Singleton Standards https://doi The Genetic Correlation Between Height and IQ: Shared Genes or Assortative Mating? https://doi Resolving the Genetic and Environmental Sources of the Correlation Between Height and Intelligence: A Study of Nearly 2600 Norwegian Male Twin Pairs https://doi 177 Cf The Dynamics of Daubert: Methodology, Conclusions, and Fit in Statistical and Econometric Studies 178 See also supra 179 The regression line, like any straight line, has an equation of the form y a bx a y x b y x Figure 11 −$32,800 + ($5,700 per year) × 16 years = −$32,800 + $91,200 = $58,400. The slope b 180 Correlations are computed from the March 2005 Current Population Survey for men ages 25–64. Freedman et al., supra Cf supra Ecological Inference and the Ecological Fallacy, in 181 See E.g. Gingles see Racially Polarized Voting Re-Solidifying Racial Bloc Voting: Empirics and Legal Doctrine in the Melting Pot The Relegation of Polarization 182 Less readily visualized procedures also are used. See, e.g. The R Journal: eiCompare: Comparing Ecological Inference Estimates Across EI and EI:R×C https://perma See, e.g. Gingles But see supra supra 183 By definition, the turnout rate equals the number of votes for the candidate, divided by the number of registrants; the rate is computed separately for each precinct. The intercept of the line in Figure 11 184 E.g. Models, Race, and the Law supra supra Thornburg v. Gingles See The Quantitative Empirics of Redistricting Litigation: Knowledge, Threats to Knowledge, and the Need for Less Districting Ecological Inference in Voting Rights Act Disputes: Where Are We Now, and Where Do We Want to Be? Redistricting plans based predominantly on racial considerations are unconstitutional unless narrowly tailored to meet a compelling state interest. Shaw v. Reno, 509 U.S. 630 (1993). Whether compliance with the Voting Rights Act can be considered a compelling interest is an open question, but efforts to sustain racially motivated redistricting on this basis have not fared well before the Supreme Court. See but see 185 Data from James W. Loewen & Bernard Grofman, Recent Developments in Methods Used in Vote Dilution Litigation 186 See supra 187 We will assume that the physicist uses a perfectly accurate and unchanging set of standard weights. 188 In standard statistical terminology, the process of measuring length here is unbiased. 189 See a a a See a b a b See and supra 190 This is the Gauss model for measurement error. See supra 191 The model would be distance a b seconds 2 ε distance a b seconds ε 192 Such extensions and modifications of simple linear regression are introduced in the Reference Guide on Multiple Regression and Advanced Statistical Models Regression and Other Stories 193 See, e.g. Review Essay: Causality and Statistical Learning 194 E.g. In re 195 E.g see also In re 196 E.g. rev’d In re 197 E.g. aff’d 198 See, e.g., In re 199 See, e.g. Reference Guide on Multiple Regression see supra See, e.g. Graphical Models for Causation, and the Identification Problem https://doi 200 Spiegelhalter, supra 201 This is the subtitle of the Penguin Books edition of Speigelhalter, supra 202 The Harvard Data Science Review https://perma 203 U.S. News & World Rep., Best Undergraduate Data Science Programs, https://www 204 The Bureau of Labor Statistics has an occupational category for “data scientists” that excludes statisticians. U.S. Bureau of Labor Statistics, Occupational Employment and Wages, May 2021, https://www 205 Applications of particular legal interest include predicting criminal behavior at the individual level ( see, e.g. Statistical Procedures for Forecasting Criminal Behavior: A Comparative Assessment https://doi see, e.g. see, e.g. Assessing the Admissibility of a New Generation of Forensic Voice Comparison Testimony Speaker Recognition Based on Deep Learning: An Overview https://doi 206 For a general discussion of AI and related societal implications, see James E. Baker & Laurie N. Hobert, Reference Guide on Artificial Intelligence 207 Vijay Kotu & Bala Deshpande, Data Science: Concepts and Practice 1 (2d ed. 2019). 208 Id see also 209 See Quomoto Dicitur “Data Science” Latine 210 Initially the increased amount of data was heralded as the dawn of a Big Data era involving Big Data specialists. See What’s the Big Idea? “Big Data” and Its Origins 211 AI is extremely broad in its scope and tools, with no single, commonly accepted definition. See, e.g. Artificial Intelligence, Predictive Policing, and Risk Assessment for Law Enforcement See, e.g. The Rise and Fall of the Legal Expert System https://perma supra 212 Susan Athey & Guido W. Imbens, Machine Learning Methods that Economists Should Know About https://doi supra 213 Additional predictors and other functional forms could be explored with the statistical procedures mentioned in the Reference Guide on Multiple Regression and Advanced Statistical Models 214 E.g. supra supra 215 The equations for estimating a b â 216 For discussion of how such procedures relate to more traditional statistical methods, see Special Issue: Commentaries on Breimen’s Two Cultures Paper 217 We could reserve a random 10% of the data and develop the jagged line on the remaining 90% using a width of, say, two LSAT points. Then we check how well the fitted line works on the reserved 10% by finding the correlation between the fitted line’s predictions and the reserved data points. We do this repeatedly, say 10 times, to get an average correlation for the jagged line with the two-point width. Then we return to the full training set and repeat this 10-fold train-test process for other values of the width. Finally, we pick the width with the highest correlation (“cross-validity”) and apply it to the full training set (100% of the data). That gives the prediction line. 218 Moritz Hardt & Benjamin Recht, Patterns, Predictions, and Actions: Foundations of Machine Learning 146 (2020) (“Methodologically, much of modern machine learning practice rests on a variant of trial and error train-test paradigm 219 For reviews of ML studies identifying frequent mistakes or departures from good practice, see Sayash Kapoor & Arvind Narayanan, Leakage and the Reproducibility Crisis in Machine-learning-based Science https://doi Common Pitfalls and Recommendations for Using Machine Learning to Detect and Prognosticate for COVID-19 Using Chest Radiographs and CT Scans Prediction Models for Diagnosis and Prognosis of Covid-19: Systematic Review and Critical Appraisal https://doi See, e.g. Prediction of Hypertension Using Traditional Regression and Machine Learning Models: A Systematic Review and Meta-Analysis https://doi Comparison of Machine Learning and Logistic Regression Models in Predicting Acute Kidney Injury: A Systematic Review and Meta-Analysis https://doi 220 The choice of statistics to indicate the accuracy of a predictor or classifier is another important topic. 221 Daniel Bone et al., Applying Machine Learning to Facilitate Autism Diagnostics: Pitfalls and Promises https://doi 222 This is a very rough description of cross-validation with data available when developing the algorithm. For discussion of data-splitting and resampling techniques, see, for example, Max Kuhn & Kjell Johnson, Applied Predictive Modeling 223 Id 224 Berk, supra 225 E.g. supra Establishment of Best Practices for Evidence for Prediction: A Review https://doi supra 226 The extent to which statisticians using Bayesian methods are overtly subjective in their analyses varies. “Objective Bayesians” use Bayes’ rule without eliciting prior probabilities from subjective beliefs. One strategy is to use preliminary data to estimate the prior probabilities and then apply Bayes’ rule to that empirical distribution. This “empirical Bayes” procedure avoids the charge of subjectivism at the cost of departing from a fully Bayesian framework. With ample data, it can be effective, and the estimates or inferences can be understood in frequentist terms. Another “objective” approach is to use “noninformative” priors that are supposed to be independent of all data and prior beliefs. However, the choice of such priors can be questioned, and the approach has been contested by frequentists and subjective Bayesians. E.g. Is “Objective Bayesian Analysis” Objective, Bayesian, or Wise? https://perma Philosophies of Probability, in 227 Probabilities may be estimated from relative frequencies, but probability itself is a subtler idea. For example, suppose a computer prints out a sequence of ten letters H and T (for heads and tails), which alternate between the two possibilities H and T as follows: H T H T H T H T H T. The relative frequency of heads is 5/10 or 50%, but it is not at all obvious that the chance of an H at the next position is 50%. There are difficulties in both the subjectivist and objectivist positions. See supra 228 “Subjective” is not necessarily the same as arbitrary. The degrees of belief must obey the same mathematical rules (axioms and theorems) as relative frequencies do, and the personal choices for particular numbers can be subjected to interpersonal standards for acceptance by other individuals. 229 E.g. Miracles and Statistics: The Casual Assumption of Independence https://doi People v. Collins infra 230 “Mutually exclusive” means that either one hypothesis or the other—but not both—is true. 231 We use the multiplication rule (4) to find that P( A H 0 A H 0 H 0 (6) and P( A H 1 A H 1 H 1 (7) Moreover, if A can occur only when either H 0 H 1 P( A A H 0 A H 1 (8) The multiplication rule (4) also shows that P( H 1 A A H 1 A (9) Using (7) to evaluate P( A H 1 A 232 Not all statisticians would accept the identification of a population frequency with P( A H 0 H 0 233 If the odds in favor of an event or a hypothesis are j k j j k 234 Unlike equation (5), equations (11) and (12) hold even when other hypotheses could account for A 235 See supra cf supra 236 For cases in which this mistake has been made, see Reference Guide on Human DNA Identification Evidence 237 Modern treatments of Bayesian methods include Donald A. Berry, Statistics: A Bayesian Perspective Bayesian Data Analysis Bayesian Methods: A Social and Behavioral Sciences Approach Doing Bayesian Data Analysis: A Tutorial with R, JAGS, and Stan Statistical Rethinking: A Bayesian Course with Examples in R and STAN 238 For problematic assumptions of independence in litigation, see, for example, Wilson v. State supra supra Copyright Bookshelf ID: NBK621595 Contents < Prev Next > Share Views PubReader Print View Cite this Page Federal Judicial Center; National Academies of Sciences, Engineering, and Medicine; Policy and Global Affairs; Committee on Science, Technology, and Law; Committee on Science for Judges—Development of the Reference Manual on Scientific Evidence, Fourth Edition. Reference Manual on Scientific Evidence: Fourth Edition. Washington (DC): National Academies Press (US); 2025 Dec 31. Reference Guide on Statistics and Research Methods. PDF version of this title In this Page CONTENTS FIGURES TABLES Introduction How Have the Data Been Collected? How Have the Data Been Presented? What Inferences Can Be Drawn from the Data? Correlation and Regression Data Science and Statistical Machine Learning Appendix: Conditional Probability and Bayes’ Rule Glossary of Terms References on Statistics and Research Methods Recent Activity Clear Turn Off Turn On Reference Guide on Statistics and Research Methods - Reference Manual on Scienti... Reference Guide on Statistics and Research Methods - Reference Manual on Scientific Evidence Your browsing activity is empty. Activity recording is turned off. Turn recording back on See more... (function($){ $('.skiplink').each(function(i, item){ var href = $($(item).attr('href')); href.attr('tabindex', '-1').addClass('skiptarget'); // ensure the target can receive focus $(item).on('click', function(event){ event.preventDefault(); $.scrollTo(href, 0, { onAfter: function(){ href.focus(); } }); }); }); })(jQuery); Follow NCBI .cls-11 { fill: #737373; } Twitter Facebook LinkedIn .cls-11, .cls-12 { fill: #737373; } .cls-11 { fill-rule: evenodd; } GitHub .cls-1{fill:#737373;} NCBI Insights Blog Connect with NLM .st20 { fill: #FFFFFF; }

.st30 { fill: none; stroke: #FFFFFF; stroke-width: 8; stroke-miterlimit: 10; } Twitter .st10 { fill: #FFFFFF; }

.st110 { fill: none; stroke: #FFFFFF; stroke-width: 8; stroke-miterlimit: 10; } Facebook Youtube .st4 { fill: none; stroke: #FFFFFF; stroke-width: 8; stroke-miterlimit: 10; }

.st5 { fill: #FFFFFF; } National Library of Medicine 8600 Rockville Pike Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov

Record · ID 1522 · SHA-256 b28a9e882c5be0aa
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.