Showing posts with label Philosophy Comp. Show all posts
Showing posts with label Philosophy Comp. Show all posts

Thursday, May 19, 2011

A Clarification of (C)

Birnbaum proves in his (1962) that (S) and (C) jointly entail (L).  He claims that (C) entails (S), which would mean that (C) entails (L) by itself, but he does not prove it.  In his (1964), he clarifies that he can only prove that (C) entails (S) by using a principle (M) that is strictly weaker than (S) but does not follow from (C).  Thus, the strongest result he can claim is that (C) and (M) jointly entail (L).  In their (1986), however, Evans et al. prove that in fact (C) alone does entail (L).  In this post I begin the task of reconstructing the Evans et al. proof.  First I need to clarify a point of obscurity in Birnbaum’s 1962 formulation of (C).

In 1962, Birnbaum formulates (C) as follows:

The Principle of Conditionality (C): If an experiment E is (mathematically equivalent to) a mixture G of components {Eh}, with possible outcomes (Eh, xh), then Ev(E,(Eh, xh)) = Ev(Eh, xh).

The parenthetical phrase “mathematically equivalent to” here turns out to be essential.  (C) applies to any experiment that contains an ancillary statistic.  In a mixture experiment, the outcome of the random process that determines which component experiment to perform is an ancillary statistic.  However, non-mixture experiments can have ancillary statistics as well.  These experiments are “mathematically equivalent to” mixture experiments, but they do not involve an actual two-stage process that consists of first using a random process to choose a component experiment and then performing that component experiment.

In his (1972), Birnbaum formulates the notion of an ancillary statistic as follows:

h = h(x) is called an ancillary statistic if it admits the factored form f(x;θ) = g(h) f(x|h; θ) where g = g(h) = Prob (h(X)=h) is independent of θ.  

In other words, an ancillary statistic is a statistic independent of the parameters whose probability distribution that can be factored out of the likelihood function f(x;θ) to yield that conditional likelihood function f(x|h; θ).  A generic mixture experiment yields a simple example.  Suppose you flip a coin to decide whether to perform experiment E1 or E2, where information about the bias of that coin is not informative about the data-generating process in E1 or E2.  There is an unconditional likelihood function f(x;θ) for this mixture experiment as a whole.  However, if a frequentist knows which way the flip turns out and thus which of E1 or E2 is performed he or she will typically take that information into account and use the conditional likelihood function f(x|h; θ) that takes this information into account.  He or she will thus neglect the mixture structure of the experiment, acting as if it had been known all along that the experiment which is actually performed would be performed.

In his (1972), Birnbaum formulates (C) in terms of the notion of an ancillary statistic.  He first defines some notation:

(Eh,x) denotes a model of evidence determined by an outcome x of the experiment Eh: (Ω,Sh,fh)  where S={x: h(x)=h}.  E may be called a mixture experiment, with components Eh having respective probabilities g(h).

He then reformulates (C):

Conditionality (C): If h(x) is an ancillary statistic, then Ev(E,x)=Ev(Eh,x), where h=h(x).

This formulation is not different in substance from Birnbaum’s 1962 formulation; it is merely more explicit that being mathematically equivalent to a mixture experiment means having an ancillary statistic.

Tuesday, May 17, 2011

A Well-Motivated Frequentist Response to Birnbaum's Theorem

I’m giving a “Works in Progress” talk on Friday to explain my current position on frequentist responses to Birnbaum’s proof.  Here is my abstract:
Frequentists appear to be committed to the sufficiency principle (S) and the conditionality principle (C).  However, Birnbaum (1962) proved that (S) and (C) entail the likelihood principle (L), which frequentist methods violate.  To respond adequately to Birnbaum’s theorem, frequentists must place restrictions on (S) and/or (C) that block Birnbaum’s proof and argue that those restrictions are well motivated.  Restricting (C) alone will not suffice, because (S) by itself implies too much of the content of (L) for frequentists to accept it.  Specifically, frequentists need to restrict (S) so that it does not apply to mixture experiments some of whose components have respective outcomes with the same likelihood function.  Berger and Wolpert (1988, p. 46) claim that such a restriction would be artificial, but in fact it has a strong frequentist motivation: reduction to the minimal sufficient statistic in such an experiment throws away information about what sampling distribution is appropriate for frequentist inference. 
I’ll try to explain the basic argument here.  Start with the claim that (S) by itself implies too much of the content of (L) for frequentists to accept it.  Kalbfleisch makes this point in his (1975) as a criticism of Durbin’s proposal to restrict (C) rather than (S).  Consider two experiments E1 and E2.  E1 involves flipping a coin five times and reporting the number of heads.  E2 involves flipping a coin until it comes up heads and reporting the number of flips required.  Suppose that in both experiments the flips are i.i.d. Bernoulli.  Imagine an instance of E1 and E2 in which one gets one head in E1, and five flips in E2, so that both E1 and E2 consist of flipping a coin five times and getting heads once.  The likelihood principle says that these two outcomes have the same evidential meaning.

The sufficiency principle does not imply that those outcomes of E1 and E2 have the same evidential meaning.  However, it does say that they would have had the same evidential meaning if they had been two outcomes of one experiment rather than outcomes of two different experiments.  So consider a mixture experiment E* that involves first flipping a coin two decide whether to perform E1 or E2 and then performing the selected experiment.  (The coin used to decide which experiment to perform should be distinct from the coin used in E1 or E2 with independent bias.)  According to the sufficiency principle, the outcome that consists of performing experiment E1 and getting one head has the same evidential meaning as the outcome that consists of performing experiment E2 and getting five tosses when each is performed as part of the mixture experiment E*.  To get the result that they also have the same evidential meaning when performed outside of E*, one needs to appeal to something like (C).  (This is essentially how Birnbaum proves that (S) and (C) entail (L).)  However, to go so far but no farther seems rather unreasonable.  As Kalbfleisch puts it, “In order to reject (L) and accept (S) one must attach great importance to the possibility of choosing randomly between E1 and E2” (p. 252).  To avoid adopting this strange position, someone who rejects (L) should reject (S) as well.

The minimal restriction on (S) that blocks Birnbaum’s proof is to modify (S) so that it does not apply to mixture experiments some of whose components have respective outcomes with the same likelihood function.  Berger and Wolpert say that this restriction “seems artificial, there being no intuitive reason to restrict sufficiency to certain types of experiments” (1988, p. 46).  One can flesh out an argument along these lines as follows.  The following argument for (S) is compelling and completely general:

Conditional on the value of a sufficient statistic, which outcome occurs is independent of the parameters of the experimental model.   Independent variables do not contain information about one another.  Thus, conditional on a sufficient statistic, which outcome occurs does not contain any information about the parameters of the experimental model.  Therefore, the evidential meaning of an experimental outcome is the same as the evidential meaning of an outcome corresponding to the same value of the sufficient statistic. 

Because this argument makes no assumptions about whether the experiment in question is pure or mixed, it would be artificial to restrict (S) to non-mixture experiments.

The problem with this argument (from a frequentist perspective) is that it assumes that the experimental model appropriate for frequentist inference is fixed in advance, regardless of which outcome occurs.  But in a mixture experiment, (C) implies that which experimental model is appropriate for frequentist inference depends on which component experiment is performed.  When the components of the mixture experiment have respective outcomes with the same likelihood function, reduction to the minimal sufficient statistic throws away the information about which component experiment was performed.  Thus, applying (S) to a mixture experiment some of whose components have respective outcomes with the same likelihood function is inappropriate from a frequentist perspective.

I think this is quite a good frequentist response to Berger and Wolpert’s objection.  The challenge for a frequentist is to make the needed restriction on (S) precise in a defensible way.  Berger and Wolpert claim that the distinction between mixture and non-mixture experiments is difficult if not impossible to characterize clearly, suggesting that this challenge will not be easy to meet.  As long as the distinction appears to be real, however, a frequentist need not be bothered too much by difficulties in formulating it precisely.

I have argued that frequentists need to restrict (S).  Fortunately for the frequentist, the needed restriction is well-motivated from a frequentist perspective.  I should note that frequentists may also need to restrict (C).  They certainly do need to do so if Evans, Fraser, and Monette (1986) are correct in their claim that (C) alone implies (L).  I have just begun looking at their paper.  They point out that the fact that seemingly innocuous principles (S) and (C) imply the highly controversial principle (L) should be a clue that there is more to (S) and (C) than meets the eye.  They seem to think that Birnbaum’s way of characterizing experimental models is too simple and that with a more adequate approach (L) would no longer follow from appropriately modified versions of (S) and (C).  It looks like their paper will take some time to digest but will be well worth the effort.  

Wednesday, April 27, 2011

Term Paper on Birnbaum's Proof

Here's a pdf of the term paper described in the previous post.  It has many loose ends, but I think it's a good start on an exciting project.

Friday, April 22, 2011

Birnbaum's Proof Part II: The Details

The notation needed to explain Birnbaum's proof in detail outruns the capabilities of Blogger, so I wrote it up as a pdf.  My conclusions are essentially unchanged from my rough sketch of the argument: Birnbaum's proof is valid, but I am suspicious that he has not formulated the principles of conditionality and sufficiency properly.  If you interpret the principle of conditionality not as a statement about evidential equivalence but instead as a directive about how to analyze experimental results (which seems appropriate to me at this time), then it is incompatible with the principle of sufficiency as formulated by Birnbaum, and indeed with any principle that can do the work that the principle of sufficiency does in Birnbaum's proof.  Another way to undermine Birnbaum's proof would be to insist, as Durbin (1970) does, that a conditional analysis can only condition on a variable that is part of the minimal sufficient statistic, although that move seems to me less appropriate to me at this time.

I did realize in examining Birnbaum's proof that it is only appropriate for experiments with discrete sample spaces.  However, I do not think that this limitation is serious because the fact that no measurement is completely precise means that all real experiments have discrete sample spaces, the idea of a continuous sample space being only a useful idealization.

Saturday, April 16, 2011

Birnbaum's Proof Part 1: The Rough Idea

The centerpiece of Birnbaum's 1962 paper is his proof that the conditionality and sufficiency principles (as he formulates them) entail the likelihood principle (as he formulates it).  This proof is significant, again, because frequentists generally accept conditionality and sufficiency but do not accept the likelihood principle, which follows from Bayesianism and implies many of the consequences of the Bayesian position that frequestists find objectionable.  In a future post, I will delve into the details of Birnbaum's proof; in this post, I just want to display its overall structure and introduce the objections it has received. 

Birnbaum considers two experiments that have pairs of respective outcomes—call them “star pairs”—that determine proportional likelihood functions.  He then constructs a hypothetical mixture experiment with these two experiments as its components.  By the conditionality principle, an outcome of either component experiment has the same as the evidential meaning of the corresponding outcome of the mixture experiment.  Now, there is a sufficient statistic that lumps together outcomes of the mixture experiment corresponding to star pair outcomes of the component experiments.  By the sufficiency principle, then, these outcomes of the mixture experiment have the same evidential meaning.  Given that an outcome of the mixture experiment has the same evidential meaning as a corresponding outcome of a component experiment, and that outcomes of the mixture experiment corresponding to star pair outcomes of the two component experiments have the same evidential meaning as one another, it follows that star pair outcomes of the two component experiments have the same evidential meaning as one another.  That is just what the likelihood principle asserts.

Call the two component experiments E and E’, respectively; call the mixture experiment E*; and let (x*, y*) be a "stair pair," with x* an outcome of E and y* an outcome of E’.  Then the following diagram depicts the structure of Birnbaum’s proof, using lines to indicate evidential equivalence and denoting above each line which principle is invoked to establish equivalence:



Several objections to this proof have appeared in the statistics literature (e.g. Durbin 1970, Cox and Hinkley 1974, Kalbfleisch 1974, Joshi 1990) and at least one in the philosophy literature (Mayo 2011), but it is still widely accepted.  I am suspicious of Birnbaum's proof, but I do not yet have confidence in any precise diagnosis of where it goes wrong.

One objection to Birnbaum's proof toward which I am sympathetic says that the conditionality princple should be understood not as a claim about evidential equivalence, but as a directive about how to analyze experimental results: thus, the conditionality principle says, "analyze experimental results conditional on which experiment was actually performed."  Understood in this way, the conditionality principle prohibits the use of the sufficient statistic that lumps together results from experiment E and experiment E', blocking Birnbaum's proof.  (Kalbleisch develops a version of this idea in his 1974, but the specific way in which he develops it may be problematic.) Birnbaum denies that the conditionality principle is to be understood as a directive (1962, p. 281 and elsewhere), but it is not clear to me that he has good reasons for doing so.

Friday, April 15, 2011

Birnbaum's Likelihood Principle

In this post, I present Birnbaum’s formulation of the likelihood principle and explain why frequentists reject and Bayesians accept this principle. Again, the big picture: frequentists typically accept conditionality and sufficiency principles while rejecting the likelihood principle. The likelihood principle is a central tenant of Bayesianism that follows directly from using Bayes’ theorem as an update rule. Many of the features of Bayesianism that frequentists find objectionable follow from the likelihood principle alone, so it is a short step from accepting the likelihood principle to becoming a Bayesian. Birnbaum argues that the conditionality and sufficiency principles are equivalent to the likelihood principle, putting significant pressure on frequentists to justify their position.


You should not be surprised to learn that the likelihood principle appeals to the notion of a likelihood; or, more precisely, a likelihood function. Birnbaum models an experiment as having a well-defined joint probability density f(x, θ) for all x in its sample space and all θ in its parameter space. This joint density implies a conditional density f_X|Θ(x|θ) for each θ. The likelihood function is this conditional density considered as a function of θ rather than x, defined up to an arbitrary multiplicative constant c: L( θ|x)=cf_X|Θ(x|θ). Roughly speaking, the likelihood function tells you how probable the model makes the data as a function of that model’s parameters.

Most frequentists are happy to use the likelihood function of an experiment in certain specific ways, such as maximum likelihood estimation and likelihood ratio testing. However, they do not accept the likelihood principle, which says, roughly, that all of the information about θ an experiment provides is contained in the likelihood function of θ. As Birnbaum formulates it, the likelihood principle is (like the conditionality and sufficiency principles) a claim about evidential equivalence. Specifically, it asserts that if the two experiments E and E’ with a common parameter space produce respective outcomes x and y that determine proportional likelihood functions, then Ev(E, x)=Ev(E’, y).

Consider two possible experiments. In the first experiment, you decide to spin a coin 12 times, and it comes up heads 3 times. Assuming that the spins are independent and identically distributed Bernoulli trials with the probability of heads on a given trial equal to p, the likelihood function of this outcome (up to an arbitrary multiplicative constant) is L(p|x=3) = (12 choose 3) 3^p 9^(1-p). In the second experiment, you decide to spin the coin until you obtain 3 heads. As it turns out, heads comes up for the third time on the 12th spin. Assuming that again the spins are independent and identically distributed Bernoulli trials with the probability of heads on a given trial equal to p, the likelihood function of this outcome (up to an arbitrary multiplicative constant) is L(p|x=12) = (11 choose 2) 3^p 9^(1-p). The likelihood functions for these two experiments are proportional, so the likelihood principle says that they have the same evidential meaning.

Standard frequentist methods say, contrary to the likelihood principle, that these two experiments do not have the same evidential meaning. (In fact, the second experiment but not the first allows one to reject the null hypothesis p=.5 at the .05 level in a one-sided test.) Frequentist methods are based on P values, where a P value is, roughly, the probability of a result at least as extreme as the observed result. The likelihood principle implies that only the outcome actually obtained in an experiment is relevant to the evidential interpretation of that experiment. Because P values refer to results other than the result that actually occurred (namely, the unrealized results that are at least as extreme as the observed result), they are incompatible with the likelihood principle.

This conflict between the likelihood principle becomes particularly stark when one considers “try and try again” stopping rules, which direct one to continue sampling until one achieves a particular result, such as a specific P value or posterior probability. For instance, a possible stopping rule is to continue collecting data until one achieves a nominally .05 significant result; that is, a result that appears to be significant at the P=.05 level if one analyzes the data as if a fixed-sample size stopping rule had been used. A frequentist would insist that data gathered according to this stopping rule does not have the same evidential meaning as the same data gathered according to a fixed-sample stopping rule. After all, the experiment with the “try and try again” stopping rule is guaranteed to generate a nominally .05 significant result, so its real P value is not .05, but 1.

Bayesians argue on the contrary that it is absurd to make the evidential meaning of an experiment sensitive to the stopping rule used. Why should the evidential meaning of a result depend on an experimenter’s intentions, which are after all “inside his or her head?”

There is much more that could be said about frequentist-Bayesian disputes about the relevance of stopping rules to inference. For present purposes, it is enough to note that the relevance of stopping rules for frequentist tests violates the likelihood principle.

The likelihood principle, while unacceptable to frequentists, is a simple consequence of the use of Bayes’ theorem as an update rule. According to Bayes’ theorem, the posterior probability of a hypothesis is equal to its prior probability times its likelihood, divided by the average of the prior probabilities of all hypotheses in the hypothesis space weighted by their likelihoods. Thus, given a prior distribution over the hypothesis space, the posterior probability of a hypothesis depends only on the likelihood function. (The arbitrary multiplicative constant included in the likelihood can be factored out of both the top and the bottom of Bayes’ theorem, so it cancels out.) Thus, Bayesianism implies the likelihood principle.

Thursday, April 14, 2011

Birnbaum's Sufficiency Principle

In my previous post, I gave an example that motivates the conditionality principle and presented Birnbaum’s formulation of that principle.  In this post, I do likewise for the sufficiency principle.  To recap the big picture: frequentists typically accept conditionality and sufficiency principles while rejecting the likelihood principle.  The likelihood principle is a central tenant of Bayesianism that follows directly from using Bayes’ theorem as an update rule.  Many of the features of Bayesianism that frequentists find objectionable follow from the likelihood principle alone, so it is a short step from accepting the likelihood principle to accepting Bayesianism.  Birnbaum argues that the conditionality and sufficiency principles are equivalent to the likelihood principle, putting significant pressure on frequentists to justify their position.

The sufficiency principle appeals to the notion of a sufficient statistic.  Roughly speaking, a sufficient statistic lumps together some outcomes of an experiment that have the following property: given that some outcome in the lumped-together set occurred, which one of those outcomes occurred is independent of the parameters of the experiment.  For instance, consider the outcome of a series of two coin tosses, where the tosses are assumed to be independent and identically distributed Bernoulli trials with probability p of heads.  The outcome space for this experiment is the set of possible sequences of outcomes of two coin tosses: HH, HT, TH, TT.  A sufficient statistic for this experiment is the number of heads.  This statistic is sufficient because, given the number of heads, the exact sequence of heads and tails is independent of p.

The sufficiency principle says, roughly, that a sufficient statistic summarizes the results of an experiment with no loss of information.  In other words, given the value t(x) of a statistic T(X) that is sufficient for θ, you don’t learn any more about θ by learning x.  This claim is very widely accepted and appears to be well-motivated.  x is independent of θ conditional on t(x), and it’s hard to see how one quantity could provide information about another quantity of which it is independent.  For instance, the sufficiency principle says that, given the number of heads obtained in a sequence of n tosses, you don’t learn any more about p by learning the exact sequence of heads and tails.  (This application of the likelihood principle requires the assumption that the tosses are independent and identically distributed Bernoulli trials; if the possibility that the tosses are non-independent were on the table, for instance, then information about sequence would be relevant and the number of heads would not be a sufficient statistic.)

Birnbaum formulates the likelihood principle, like the sufficiency principle, as a claim about evidential equivalence.  Take an experiment E with outcome x and a derived experiment E’ with outcome t=t(x), where T(X) is a sufficient statistic; then Ev(E, x)=Ev(E’,t).  In other words, reduction to a sufficient statistic does not change the evidential meaning of an experiment.

My comments at the end of the previous post are also appropriate here: it is a good idea to be wary of general principles insofar as they are motivated merely by the fact that they seem to capture the intuitions at play in simple examples.  However, it is worth keeping in mind that the sufficiency principle is not motivated only (or even, I think, primarily) by its intuitive appeal in simple cases, but also by the general claim that one quantity cannot provide information about another quantity of which it is independent.  Similarly, the conditionality principle is motivated by the general claim that experiments that were not performed are irrelevant to the interpretation of the experiment that was performed.  However, the conditionality principle is formulated to apply to experiments that are “mathematically equivalent” to mixture experiments, so it is not clear that this general claim is general enough to warrant the principle.

Birnbaum's Conditionality Principle

Allan Birnbaum’s 1962 paper “On the Foundations of Statistical Inference” purports to prove that the conditionality and sufficiency principles—which frequentists typically accept—jointly entail the likelihood principle—which frequentists typically reject.  The likelihood principle is an important consequence of Bayesianism.  Moreover, many of the consequences of Bayesianism that frequentists typically find objectionable (e.g., the stopping rule principle) follow from the likelihood principle alone.  Thus, once one accepts the likelihood principle and its consequences, there is little to stop one from becoming a Bayesian.  The prominent Bayesian L. J. Savage said that he began to take  Bayesian seriously “only through the recognition of the likelihood principle.”  As a result, he called the initial presentation of Birnbaum’s paper “really a historic occasion.”

In this post, I will discuss why both frequentists and Bayesians find the conditionality principle attractive, and I will provide Birnbaum’s formulation of that principle.

In rough intuitive terms, the conditionality principle says that only the experiment that was actually performed is relevant for interpreting that experiment’s results.  Stated this way, the principle seems rather obvious.  For instance, suppose your lab contains two thermometers, one of which is more precise than the other.  You share your lab with another researcher, and you both want to use the more precise thermometer for today’s experiments.  You decide to resolve your dispute by tossing a fair coin.  Once you have received your thermometer and run your experiment, there are two kinds of methods you could use to analyze your results.  One kind of method is an unconditional approach, which assigns margins of error in light of the fact that you had a 50/50 chance of using either thermometer, without taking into account which thermometer you actually used.  The other kind is an conditional approach, which assigns margins of error to measurements in light of the thermometer you actually used, ignoring the fact that you might have used the other thermometer.  Most statisticians regard as highly counterintuitive the idea that you should take into account the fact that you might have used a thermometer other than the one you actually used in interpreting your data.  Thus, most statisticians favor the conditional approach in this case.  The conditionality principle is designed to capture this intuition.

To express the conditionality principle precisely, it will be necessary to introduce some notation.  Birnbaum models an experiment as having a parameter space Ω of vectors θ, a sample space S of vectors x, and a joint probability distribution f(x, θ) defined for all x and θ.  He writes the outcome x of experiment E as (E, x), and the “evidential meaning” of that outcome as Ev(E, x).  He does not attempt to characterize the notion of evidential meaning beyond the constraints given by the conditionaly, sufficiency, and likelihood principles.  Each of those principles states conditions in which two outcomes of experiments have the same evidential meaning.

Birnbaum expresses the conditionality principle in terms of the notion of a mixture experiment.  A mixture experiment E involves first choosing which of a number of possible “component experiments” to perform by observing the value of a random variable h with a known distribution independent of θ, and then taking an observation xh from the selected component experiment Eh.  One can then represent the outcome of this experiment as either (h, xh) or, equivalently, as (Eh, xh).  The conditionality principle says that Ev(E, (Eh, xh))=Ev(Eh, xh).  In words, the evidential meaning of the outcome of a mixture experiment is the same as the evidential meaning of the corresponding outcome of the component experiment that was actually performed.

The above discussion of the conditionality principle follows a pattern that is common in philosophy: start with an intuition-pumping example, then state a principle that seems to capture the source of the intuition at work in that example.  It takes only a little experience with philosophical disputes to become suspicious of this pattern of reasoning.  There are always many general principles that can be used to license judgments about particular cases, and there are typically counterexamples to whatever happens to be the most “obvious” or “natural” general principle.  Take, for instance, theories of causation.  It is easy to give examples to motivate, say a David Lewis-style counterfactual analysis of causation.  For instance, the Titanic sank because it struck an iceberg.  Analysis: the Titanic struck an iceberg and sank, and if it hadn’t struck that iceberg then it wouldn’t have sunk.  In general: c causes e if and only if c and e both occur, and if c hadn’t occurred then e wouldn’t have occurred.  This analysis seems to capture what’s going on in the Titanic example, but counterexamples abound.  For instance, suppose that (counterfactually, so far as I know) there had been a terrorist on board the Titanic who would have sabotaged it and caused it to sink the next day if it hadn’t struck the iceberg and sunk.  Presumably, one still wants to say in this scenario that the iceberg caused the Titanic to sink.  Nevertheless, the Titanic would have sunk even if it hadn’t struck the iceberg.  Typically in philosophical debates, a counterexample like this one leads to a revision of the original analysis that blocks the counterexample; that revised analysis is then subjected to another counterexample, which leads to further revision; and this counterexample-revision-counterexample cycle iterates until the analysis becomes so complex that the core idea that motivated the original analysis starts to seem hopeless.  That idea is abandoned, a new idea is proposed, and the process is repeated with that new idea.

In short, my training in philosophy inclines me to be suspicious of Birnbaum’s conditionality principle, even though it seems to capture what’s going on in the simple thermometer example.  Because of the conditionality principle’s technical and specialized nature, however, it is not as easy to think of potential counterexamples.  I will table this concern for now; in future posts, I will discuss counterexamples to the conditionality principle that statisticians have proposed, and revisions to that principle they have suggested.

Revised Topics

My philosophy comp topic has evolved gradually, while my history comp topic has changed drastically.


My current philosophy comp project begins with a 1962 paper in which Allan Birnbaum argues that two principles frequentist statisticians typically accept—the conditionality and sufficiency principles—imply a principle they typically reject—the likelihood principle. The likelihood principle is a consequence of Bayes’ theorem, and Bayes’ theorem provides perhaps the simplest way to implement the likelihood principle in statistical inference, so Birnbaum’s argument tends to push frequentists toward Bayesianism.

Birnbaum’s argument is famous among those interested in the philosophy of statistics, but it has been criticized. Several statisticians have argued that Birnbaum’s formulation of either the conditionality principle or the sufficiency principle is too strong, and that replacing it with a suitably weakened principle would not allow Birnbaum’s argument to go through. However, these statisticians disagree among themselves about how Birnbaum’s principles should be weakened, and their specific proposals have been criticized. Joshi and Mayo have raised stronger objections to Birnbaum’s argument, arguing that there is a flaw in Birnbaum’s logic rather than in his premises.

I do not yet know what to say about Birnbaum’s argument, but I think that with enough work I am bound to find something interesting. Whether Birnbaum is right or not, there is work to be done in pinpointing exactly where either he or his critics go wrong, and the results of such an analysis are likely to have significant implications for the foundations of statistics.

I am abandoning my history comp project based on the Millikan oil-drop experiment. Millikan’s notebooks have already received careful scrutiny, and after some preliminary work it is not clear that an experimental approach will yield any significant new insights in time for the comp deadline. Moreover, it has come to my attention that there is a researcher in Germany who is way ahead of me in tracking down and investigating extant versions of Millikan’s apparatus.

Instead, I am planning to write my history comp on a puzzling passage in Darwin’s Origin of Species. I wrote a paper on this topic last year and received encouraging comments on it and suggestions for expanding it. In particular, I am planning to investigate how this passage changed through subsequent editions of the Origin and to look for evidence that might indicate why Darwin made the particular changes he did.

Friday, February 25, 2011

Is Spectrum Bias a Problem for Error Statistics?

A phenomenon called spectrum bias might help my argument that advocates of error statistics should take the positive predictive value (PPV) and negative predictive value (NPV) of their tests seriously. Spectrum bias is typically discussed in the context of medical diagnostic tests. Such tests are characterized by their sensitivity and specificity, where a test’s sensitivity is the probability that it yields a positive result if the condition in question is present, and its specificity is the probability that it yields a negative result if the condition in question is absent. PPV and NPV are more clinically relevant than sensitivity and specificity. However, sensitivity and specificity are more popular measures of a test’s performance because, unlike PPV and NPV, they are generally taken to be intrinsic properties of the test, independent of the prevalence of the condition in the population.
Spectrum bias is the phenomenon that sensitivity and specificity are not, in fact, intrinsic properties of medical tests. Like PPV and NPV, they vary when they are applied to different populations. There are both theoretical and empirical studies supporting the claim that spectrum bias exists. At least one study I have looked at purports to show that sensitivity and specificity vary with features of the population almost as much as PPV and NPV. One part at least of the explanation for this phenomenon is that medical conditions typically are not truly dichotomous; they can be present to varying extents. Misclassification is more likely for individuals who are close to the classification cutoff. As a result, sensitivity and specificity are lower for populations in which many individuals are close to the cutoff than they are for populations without this feature.

If spectrum bias afflicts error statistical tests generally, then an advocate of error statistics cannot deny the relevance of PPV and NPV on the grounds that they are not intrinsic properties of tests without also impugning their preferred error rates α and β.
I need to find out more about spectrum bias and its prevalence and severity before I can be confident that this argument is a good one. However, it does seem promising and is not likely to have been considered before within the philosophy of science, where spectrum bias seems to be largely unknown.

Wednesday, February 23, 2011

A Refinement of Ioannidis' Argument

I've briefly written up in Word an idea for a refinement of Ioannidis' argument that would yield results that are relevant to error statistics and the base-rate fallacy.  You can download the file here.  (Unfortunately, the figures don't show up properly in the Google docs viewer that the link brings up--you need to download the file and view it in Word.)

Sunday, February 6, 2011

Ioannidis' Argument

John Ioannidis is a “meta-researcher” who has written several well-known papers about the reliability of research that uses frequentist hypothesis testing.  (The Atlantic published a reasonably good article about Ioannidis and his work in November 2010.)  One of his most-cited papers, called “Why Most Published Research Findings Are False,” presents a more general version of the example I presented in my last post.  Steven Goodman and Sander Greenland wrote a response arguing that Ioannidis’ analysis is overly pessimistic because it does not take into account the observed significance value of research findings.  Ioannidis then responded to defend his original position.  I’m planning to work through this exchange, with an eye toward the question whether an argument like Ioannidis’ could be used to present the base-rate fallacy objection to error statistics in a way that is more consistent with frequentist scruples than Howson’s original presentation.

Ioannidis’ argument generalizes the example I gave in my previous post: it uses the same kind of reasoning, but with variables instead of constants.  Ioannidis idealizes science as consisting of well-defined fields i=1, …, n, each with a characteristic ratio Ri of “true relationships” (for which the null hypothesis is false) to “no relationships” (the null is true) among those relationships it investigates, and with characteristic Type I and Type II error rates αi and βi.  Under these conditions, he shows, the probability that a statistically significant result in field i reflects a true relationship is (1- βi)Ri/(Ri- βiRi- αi).

It’s somewhat difficult to see where this expression comes from in Ioannidis’ presentation.  This blog post by Alex Tabarrok presents Ioannidis’ argument in a way that’s easier to follow, including the following diagram:



Here’s the idea.  You start with, say, 1000 hypotheses in a given field, at the top of the diagram.  The ratio of true hypotheses to false hypotheses in the field is R, so a simple algebraic manipulation (omitting subscripts) shows that the fraction of all hypotheses that are true hypotheses is R(1+R), while the fraction that are false is 1- R(1+R).  That brings us to the second row from the top in the diagram: if R is, say ¼ (so that there is one true hypothesis investigated for every four false hypotheses investigated), then on average out of 1000 hypotheses 200 will be true and 800 false.  Of the 200 true hypotheses, some will generate statistically significant results, while others will not.  The probability that an investigation of a true relationship in this field yields a statistically significant result is, by hypothesis, β.  If β is, say, .6, then on average there will be 120 positive results out of the 200 true hypotheses investigated.  Similarly, of the 800 false hypotheses, some will generate statistically significant results; the probability that any one hypothesis will do so is α.  Letting α=.05, then, on average there will be 40 positive results for the 800 false hypotheses investigated.  Thus, the process of hypothesis testing in this field yields on average 40 false positives for every 120 true positives, giving it a PPV of .75.  Running through the example with variables instead of numbers, one arrives at a PPV of (1- β)R/(R- βR- α).  This result implies that PPV is greater than .5 if and only if (1- β)R > α.

Ioannidis goes on to model the effects of bias on PPV and to develop a number of “corollaries” about factors that affect the probability that a given research funding is true (e.g. “The hotter a scientific field… the less likely the research findings are to be true”).  He argues that for most study designs in most scientific fields, the PPV of a published positive result is less than .5.  He then makes some suggestions for raising this value.  These elaborations are quite interesting, but for the moment I would like to slow down and examine the idealizations in Ioannidis’ argument.

The idea that science consists of well-defined “fields” with uniform Type I and Type II error rates across all experiments is certainly an idealization, but a benign one so far as I can tell.  The assumption that each field has an (often rather small) characteristic ratio R of “true relationships” to “no relationships” is more problematic.  First, what do we mean by a “true relationship?”  One of the most basic kinds of relationship that researchers investigate is simple probabilistic dependence: there is a “true relationship” of this kind between X and Y if and only if P(X & Y) ≠ P(X)*P(Y).  However, if probabilistic dependence is representative of the relationships that scientists investigate, then one might reasonably claim that the ratio of “true relationships” to “no relationships” in a given field is always quite high, because nearly everything is probabilistically relevant to nearly everything else, if only very slightly.  In fact, a common objection to null hypothesis testing is that (point) null hypotheses are essentially always false, so that testing them serves no useful purpose.

One could avoid this objection by replacing “no relationship” with, say, “negligible relationship” and “true relationship” with “non-negligible relationship.”  However, for the argument to go through one would then have to reconceive of α as the probability of rejecting the null given that the true discrepancy from the null is non-negligible, and of β as the probability of failing to reject the null given that the true discrepancy from the null is negligible.  Fortunately, these probabilities would generally be nearly the same as the nominal α and β.

The assumption that each field has a characteristic R is more problematic than the assumption that each field has a characteristic α and β for a second reason as well: ascribing a value for R to a test requires choosing a reference class for that test.  The assumption that each field has a characteristic α and β is simply a computational convenience; α and β are defined for a particular test even though this assumption is false.  By contrast, the assumption that each field has a characteristic R is more than a convenience: something like it must be at least approximately true for R to be well defined in a particular case.  This point seems to me the Achilles’ heel of Ioannidis’ argument, and of attempts to persuade frequentists to treat PPV as a test operating characteristic on par with α and β.  A frequentist could reasonably object that there is no principled basis for choosing a particular reference class to use in a particular case in order to estimate R.  And even with a particular reference class, there are significant challenges to obtaining a reasonable estimate for R.

Thursday, February 3, 2011

Positive Predictive Value as an Operating Characteristic

The Mayo and Howson papers I examined in my last few posts came out of a symposium at the 1996 meeting of the Philosophy of Science Association.  In this post, I turn my attention to the third paper that came out of that symposium, this one by Ronald Giere.

Giere takes a somewhat neutral, third-party stance on the debate between Mayo and Howson, although his sympathies seem to lie more with error statistics.  He contrasts Mayo and Howson’s views as follows: Howson attempts to offer a logic of scientific inference, analogous to deductive logic, whereas Mayo aims to describe scientific methods with desirable operating characteristics.

This distinction does seem to capture how Howson and Mayo think about what they are doing.  However, it does not make me any more sympathetic to error statistics, because it seems to me a mistake to try to separate method from logic.  The operating characteristics of a scientific method are desirable to the extent that they allow one to draw reliable inferences, and the extent to which they allow one to draw reliable inferences depends on logical considerations.

Nevertheless, the logic/method distinction is useful for understanding the perspective of frequentists such as Mayo.  In fact, one may be able to use the insight this distinction provides to recast the base-rate fallacy objection in a way that will strike closer to home for a frequentist.  The key is to present the objection in terms of the positive predictive value of a test and to argue that positive predictive value (PPV) is an operating characteristic on par with Type I and Type II error rates.  In fact, a test can have low rates of Type I and Type II error (low α and β), but still have low positive predictive value.  Consider the following (oversimplified) example:

Suppose that in a particular field of research, 9/10 of the null hypotheses tested are true.  For simplicity, I will assume that all of the tests in this field use the same α and β levels: the conventional α=.05, and the lousy but fairly common β=.5.  The following 2x2 table displays the most probable set of outcomes out of 1000 tests:

Test rejects H0
Test fails to reject H­0

H0 is true
45
855
900
H0 is false
50
50
100

95
905


Intuitively, PPV is the probability that a positive result is genuine.  In more frequentist terms, it is the frequency of false nulls among cases in which the test rejects the null.  In this example, PPV is 45/95 = .53.  As the example shows, a test can have low α and β (desirable) without having high PPV, if the base rate of false nulls is sufficiently low.  (Thanks to Elizabeth Silver for providing me with this example.)

Superficially, this way of presenting the base-rate objection appears to be closer to the frequentist framework than Howson’s way of presenting it.  However, a frequentist might object that PPV is not an operating characteristic of a test in the same way that Type I and Type II error rates are.  Type I error rates, one might think, are genuine operating characteristics because they do not depend on any features of the subject matter to which the test is applied.  One simply stipulates a Type I error rate and chooses acceptance and rejections for one’s test statistic that yield that error rate.  By contrast, to calculate PPV one has to take into account the fraction of true nulls within the subject area in question.  Thus, PPV is not an intrinsic feature characteristic of a test, but an extrinsic feature of the test relative to a subject area.

This objection ignores the fact that Type I error rates are calculated on the basis of assumptions about the subject matter under test—most often, assumptions of normality.  As a result, Type I error rates are not intrinsic features of tests either, but of tests as applied to subject areas in which (typically) things are approximately normal.  Normality assumptions may be more widely applicable and more robust than assumptions about base rates, but they are nonetheless features of the subject matter rather than features of the test itself.  Type II error rates are even more obviously features of the test relative to a subject matter, because they are typically calculated for a particular alternative hypothesis that is taken to be plausible or relevant to the case at hand.

A frequentist could respond simply by conceding the point: PPV is an operating characteristic of a test that is relevant to whether one can conclude that the null is false on the basis of a positive result.  To do so, however, would be to abandon the severity requirement and to move closer to the Bayesian camp.

The example given above uses the same kind of reasoning that John Ioannidis uses in his paper “Why Most Published Research Findings are False.”  It might be useful to move next to that paper and the responses it received.


Before moving on, I'd like to note a couple other interesting other moves Giere makes in his paper.  First, he characterizes Bayesianism and error statistics as extensions of the rival research programs that Carnap and Reichenbach were developing around 1950, but without those programs' foundationalist ambitions.  Second, Giere emphasizes a point that I think is very important: Bayesianism (as it is typically understood in the philosophy of science) is concerned with the probability that propositions are true.  It is not concerned (at least in the first instance) with how close to the truth any false propositions may be.  Yet, in many (if not all) scientific applications, the truth is not an attainable goal.  Even staunch scientific realists admit that our best scientific theories are very probably false.  Where they differ from anti-realists is that they claim that our best theories are close to and/or tending toward the truth.  One might think that the emphasis in real scientific cases on approximate truth rather than probable truth favors error statistics over Bayesianism.  However, when one moves into real scientific cases one should also move into real Bayesian methods, which include Bayesian methods of model building, which are Bayesian in that involve conditioning on priors but are not like the Bayesian methods that philosophers tend to focus on because they aim to produce models that are approximately true rather than models that have a high posterior probability.  Unlike Bayesian philosophers, perhaps, Bayesian statisticians have develop a variety of methods that can handle the notion of approximate truth just as well as error-statistical methods.

Saturday, January 29, 2011

Mayo’s Reasons for Rejecting Premise 4 of Howson’s Argument

I realize now that I misunderstood Mayo’s use of J in place of ~H.  ~H says that breast cancer is absent, whereas J says that breast disease (inclusive of breast cancer) is absent.  The idea seems to be that a non-cancerous breast disease is likely to trigger a false-positive result in a test for breast cancer, and that this possibility makes it the case that a positive test result does not pass H severely.

This point does not affect the upshot of my analysis.  Howson can simply stipulate a hypothetical case in which there is no state corresponding to J in which a false positive test is likely.  That is enough to show that the severity requirement is unsound in principle.

Moreover, Mayo grants that ~J (which says that breast disease is present) does pass a severe test despite having (we can assume) a low posterior probability.  Thus, she allows that a hypothesis can meet the severity requirement despite having a low posterior, effectively granting premise 2 of Howson’s argument, and turns her attention to premise 4.

Here is Mayo’s reconstruction of Howson’s argument modified to reflect the fact that Mayo denies that H passes a severe test but allows that ~J does so:

  1. An abnormal result is taken as failing to reject H (i.e., as “accepting H”); while rejecting J, that no breast disease exists
  2. ~J passes a severe test and thus ~J is indicated according to (*). (Modified)
  3.  But the disease is so rare in the population (from which the patient was randomly sampled) that the posterior probability of ~J given e is still very low (and that of ~H is still very high).  (Modified)
  4. Therefore, “intuitively,” ~J is not indicated but rather ~H is.  (Modified)
  5. Therefore, (*) is unsound.
Mayo’s argument against premise 4 is interesting, but an orthodox Bayesian has an easy response.  She points out that both error-statistical and Bayesian tests involve probabilistic calculations that are themselves deductive.  The error-statistical framework only becomes ampliative with the introduction of the severity requirement (*), which goes beyond those calculations to make an assertion about which claims are well supported by tests.  She demands that (*) be compared not against the deductive probabilistic calculations that Bayesians perform, but against a truly ampliative Bayesian rule.  What she is demanding, in effect, is a rule of detachment, which tells a Bayesian when to infer from a statement of the form “the probability of H is p” to the statement “H.”

A Bayesian has at least two possible responses to this maneuver.  First, it is not clear that Bayesian updating is a deductive method of inference.  It uses a rule—Bayes’ theorem—that follows from the axioms of probability, but those axioms are not dictated by classical logic, nor is the normative claim that the right way to update one’s degree of belief in H when one has an experience the only direct epistemic import of which is to change one’s degree of belief in E to 1 is by conditioning on E, for all propositions H and E.  Second, Mayo has not given any reason why Bayesians should adopt a rule of detachment, rather than being strict probabilists.  The demand that orange-selling Bayesians provide an apple to compare with her apple would be unfair if part of the Bayesian position were that oranges can do everything apples can do at least as well apples do it.  (A Bayesian can still approximate high-probability beliefs as full beliefs as a useful heuristic when doing so is not likely to lead to trouble.)

Having demanded (unfairly) that Bayesians adopt a rule of detachment, Mayo ascribes to Howson the following implicit rule:

·        There is a good indication or strong evidence for the correctness of hypothesis H just to the extent that it has a high posterior probability.

She then turns Howson’s example against him, ascribing to him this rule of detachment.  She notes that for a woman in her forties, the posterior probability of breast cancer given an abnormal mammogram is about 2.5%, which makes it very close to Howson’s hypothetical example.  Under an error-statistics approach, the hypothesis that such a woman does not have breast cancer does not pass a severe test with a positive result; nor does the hypothesis that such a woman does have breast cancer.  To provide strong evidence one way or the other, follow-up tests are needed.  Under a Bayesian approach with a rule of detachment, the fact that the posterior probability that the woman has breast cancer is small provides a good indication that breast cancer is absent, “so the follow-up that discovered these cancers would not have been warranted.”

This argument is grossly unfair to the Bayesian position.  For an orthodox Bayesian, whether or not follow-up tests are warranted (and whether or not the initial test was warranted) for a given individual depends on that individuals expected utilities.  Rounding down probability 2.5% to 0% in an expected utility calculation is likely to lead to errors when the utility of the unlikely event is extremely high or extremely low, as in this case.  This fact speaks not against Bayesianism, but against simple rules of detachment.

In summary, Mayo has shown that error statistics is sometimes more sensible than Bayesianism with a simple-minded rule of detachment.  But a sensible Bayesian would not use such a rule of detachment, so this conclusion has no force against Bayesianism.

Mayo’s Reasons for Rejecting Premise 2 of Howson’s Argument

In my previous post, I presented Mayo’s reconstruction of Howson’s argument against error statistics:
  1. An abnormal result is taken as failing to reject H (i.e., as “accepting H”); while rejecting J, that no breast disease exists.
  2. H passes a severe test and thus H is indicated according to (*).
  3. But the disease is so rare in the population (from which the patient was randomly sampled) that the posterior probability of H given e is still very low (and that of J is still very high).
  4. Therefore, “intuitively,” H is not indicated but rather J is.
  5. Therefore (*) is unsound.
In this post, I will examine the reasons Mayo gives for rejecting premise 2.

Again, in the paper I am presently considering* Mayo expresses her severity requirement as follows:

  • (*): e is a good indication of H to the extent that H has passed a severe test with e.
where a test is “severe” with respect to H if and only if that test has a very low probability of passing H if H is false.

Howson gives an example of a medical test in which the hypothesis that a given patient has the disease in question (which in Mayo’s version of the example is breast cancer) appears to pass a severe test with a positive result, yet the posterior probability of the hypothesis that the patient has the disease conditional on the positive result is low.  He takes this case to be a counterexample which shows that Mayo’s severity requirement is unsound.

Mayo responds in part by denying that the severity requirement has been met in this case.  That is, she rejects premise 2 in her reconstruction of Howson’s.  What reasons does she give for doing so?

First, after protesting a bit about the idealized nature of Howson’s example (which seems to me irrelevant to the point at issue), Mayo says that she will try to apply her severity requirement to it.  She does so as follows (p. S208):
a.       An abnormal result e is a poor indication of the presence of disease more extensive than d if such an abnormal result is probable even with the presence of disease no more extensive than d.
b.      An abnormal result e is a good indication of the presence of  disease as extensive as d if it is very improbable that such an abnormal result would have occurred if a lesser extent of disease were present.

Mayo has in mind a more realistic case than Howson’s, in which a disease can be present to varying extents.  However, if her account aspires to provide a general theory of evidence, then it should apply to binary cases as well.  Thus, in the context of this debate it seems unfair of Mayo to change the example.  Sticking with Howson’s actual example, Mayo’s (a) and (b) reduce to the claim that a positive test result indicates the presence of disease to the extent that a positive result is improbable if the disease is absent.

At times, Mayo seems to be claiming something weaker for her severity requirement than that it provides a general theory of evidence—for instance, that it provides a reasonable guide for inductive inference when we do not have a strong evidential basis for assigning prior probabilities.  Moreover, she claims that this kind of situation is very common in science, which makes understanding her severity requirement and the error-statistical techniques that conform to it quite important for understanding scientific practice.  It seems to me that Mayo is on firm ground here.  Moreover, I suspect that there is a lot of room here for reconciling her approach with Bayesianism by showing, for example, that frequentist techniques provide reasonably good approximations to Bayesian methods when priors are not known with precision but are known to be not too extreme. 

However, Mayo goes further by presenting her account as a rival to Bayesianism, and by arguing not only that Bayesian techniques are hard to apply in many cases, but also that they are vitiated by their frequent dependencies on epistemic probabilities.  As Clark Glymour points out in his paper “Instrumental Probability,” the claim that Bayesianism and the use of epistemic probabilities are “too subjective” is often motivated by confusing justification with content.  An epistemic probability is “subjective” in the sense that it is a property of an individual’s (idealized) belief state (content), but it may nevertheless have a strong “objective” evidential basis (justification).  When the “objective” justification for an epistemic probability is strong, I see no reason to object to it and the grounds that its content is subjective.

Returning to Howson’s example, it certainly appears that a positive result in Howson’s example satisfies Mayo’s severity requirement, given that a positive result is quite improbable if the disease is absent (P~H(+)=.05 in Howson’s example, but that number can be made as small as one likes as long as the incidence rate/prior probability is adjusted downward to compensate).  But Mayo denies that a positive result satisfies the severity requirement.

In support of this claim, Mayo points out that, unlike, the Neyman-Pearson framework for statistical tests, error statistics does not use automatic accept/reject rules; for instance, within the error-statistical framework one generally would not infer that a point null hypothesis is true from the fact that one fails to reject that null hypothesis at a pre-specified alpha level.  The reason for this restraint is clear within the error-statistical framework: significance tests are unlikely to reject the null if the true value of parameter of interest is close to the null value relative to the power of the test.  As a result, one cannot infer with severity from a failure to reject a point null hypothesis that that null hypothesis is true; one can at most infer with severity that the true value of the parameter is close to the null.  (For instance, one might estimate a (1-alpha)% confidence interval for the parameter value, which would contain the null value.)

This move of Mayo’s seems to me a significant improvement in the Neyman-Pearson framework.  However, it does not help in the case at hand, in which we are not considering a point null hypothesis about a continuous variable, but rather a hypothesis about a binary variable.  On the other hand, understanding this aspect of Mayo’s account does help in understanding Mayo’s next move: Mayo claims that a failure to reject the hypothesis that the patient has breast cancer given a positive test result does not indicate that the patient has breast cancer so long as there are alternatives to this hypothesis that would very often produce the positive result.

Notice the analogy with a test of a hypothesized value for a parameter, which is presumably motivating Mayo’s claim here: one generally cannot conclude with severity that a point null hypothesis is true, because slight deviations from that point null are effectively indistinguishable from the null in a hypothesis test.  In the same way, one cannot conclude with severity that a patient has breast cancer if there are other possible situations that would make a positive test outcome likely.

This requirement is surely too strong.  Suppose that there is a non-diseased condition that mimics whatever sign or symptom of breast cancer the test picks up, generating false positives, but these condition is extremely rare—as rare as one likes.  Then there would be an alternative to the hypothesis that the patient in question has breast cancer that would very often produce a positive result, but that there is no great need to investigate in concluding that the patient has breast cancer.  Obviously one needs to rule out all possible alternatives to make a deductive inference, but one does not need to rule out incredibly rare/improbable alternatives to make a solid inductive inference.

There is a sound motivation behind Mayo’s claim that failure to reject H with a particular result does not indicate that H is true as long as there are alternatives to H that would make that result probable: in order to be telling in favor of a hypothesis, evidence must not only agree with that hypothesis but also speak against alternative hypotheses.  However, the requirement goes too far.  It seems that the sensible approach is to bring in prior probabilities and to require that any alternative hypotheses that would make the test outcome probable be themselves sufficiently improbable that the probability that any one of them is true is very small.  Bayesianism implements this approach in a precise and well-motivated way, but in situations in which Bayes’ theorem is difficult to apply one could combine informal considerations of prior probability with the severity requirement to approximate Bayesian reasoning fairly well.

Getting back to the main point, does Mayo have a good argument against premise 2?  I think not.  In a realistic case, there could be many ways in which the hypothesis that a given person has a given disease could be false, some of which might make it probable that that the person would test positive for the disease despite not having it.  Mayo would require ruling out such possibilities before declaring that the claim that the person has the disease is well supported.  As long as we're considering her account as a theory of evidence, however, we need not be constrained by what would happen in most realistic cases.  We can simply stipulate a hypothetical case in which the probability that someone who does not have the disease gets a positive result is .95, and that there are no further facts that would allow us to partition the set of people who do not have the disease into some who would get a positive result with high probability and some who would not.  There are certainly many realistic cases in which we do not know any such facts, even if they exist, so this case in not so far removed from practice to be wholly uninteresting.  In such a case, I do not see how Mayo's objections have any force against premise 2.

*The article I am considering is Deborah  Mayo’s 1997 “Error Statistics and Learning from Error: Making a Virtue of Necessity.”  It appeared in Philosophy of Science  Vol. 64, Supplement.  Proceedings of the 1996 Biennial Meeting of the Philosophy of Science Association.  Part II: Symposia Papers (Dec., 1997), pp. S195-212.  It is a response to Colin Howson’s “Error Statistics in Error,” pp. S185-S194 in the same issue.