Assignment
This course is all about establishing a solid theoretical foundation of empirical work. In this project, however, we will contrast this understanding with some, hopefully, thought provoking writing that has appeared in economic and statistical journals over the past decades. The idea here is to complement the curriculum with another dimension and make you reflect on how statistical methods are actually used, misused, debated and reformed in practice.
The CLT tells you that the sample mean converges to the normal distribution under some conditions. It does not tell you where these conditions come from and whether they are credibly satisfied in a given situation, what you should do when defensible choices lead to different conclusions, or how to communicate uncertainty honestly to a reader who will act on your conclusions. These are questions that every empirical researcher faces, often without much guidance from the textbook, and the purpose of this reading list is to address these issues from a few different angles.
The task is that you (and your partner if you have one) chooses one of the papers listed below. You will write a short report (3-5 pages), where about 1 page is a summary describing the main idea of the article and the rest is a deeper analysis of some aspect or idea that relates the paper to the topics of our course in particular, and the practice of empirical work more generally. You will then prepare a 10-15 minute presentation to the class based on the report.
Here are some additional provisions:
- I have supplied a little context and some questions/prompts to each paper that you can use as a starting point, but feel free to steer the project in a different direction, for example towards your own specialty, if you want.
- There are some natural groupings among the papers which I indicate in the end. If you are two people, then maybe you have the capacity to read two and then analyse the relation between them? Or can you coordinate with another person/group and discuss the relationship between your projects?
- Several articles in the list are discussed in the same journal issue as they appear.
- Feel free to suggest other articles that are closer to your own field of study.
On the use of AI: The intended outcome of this project is to make you reflect critically on advanced subjects. It is the quality of your reflections that will make this work successful, and not necessarily that the report and presentation are super-professional. You therefore have two options for AI use and you must make a conscious choice between these options before you do anything:
- No use of generative AI whatsoever, in any part of the project. Not for reading, not for writing, not for discussion.
- Using generative AI as a partner in understanding the ideas only.
The second option requires additional reporting. You will read the paper yourself and write the first draft of the report without use of AI. You can then discuss the ideas in the paper freely with a chatbot in a single session. When you are done, you will revise your report and presentation, again without AI. You must hand in: 1. The first draft, 2. The entire chat log (in English please) which will be evaluated for critical thinking and discussion, and 3. The finished report.
Discussions with other humans allowed and encouraged at any stage of the project.
I don’t want to be that guy, but I have to write this out formally: Any departure from this procedure will be considered as cheating, and reported as such..
The Papers
1. Wild and Pfannkuch (1999) – Statistical Thinking in Empirical Enquiry (with discussion)
Wild, C. J., & Pfannkuch, M. (1999). Statistical Thinking in Empirical Enquiry. International Statistical Review 67(3), 223–248.
Building on interviews with professional statisticians and observations of students, Wild and Pfannkuch propose a four-dimensional model of statistical thinking: (i) the investigative cycle (Problem, Plan, Data, Analysis, Conclusions — PPDAC); (ii) types of thinking, including recognition of variation, transnumeration (re-representing data to create new understanding), reasoning with statistical models, and integration of statistical with contextual knowledge; (iii) an interrogative cycle (generate, seek, interpret, criticize, judge); and (iv) dispositions such as skepticism, curiosity, and openness. The paper gives us additional vocabulary beyond what we have otherwise seen in the course. Reflecting on this structure is useful when transitioning from theory to real empirical work.
- Have you started to think about your own PhD project? Where are you currently in the PPDAC cycle? Do you have what you need to navigate this landscape?
- The authors give transnumeration a central role. Do you know of any cases in your field where re-representing the data (a different unit of observation, log-values, ratios, visualizations) have changed the interpretation of the problem in an important way?
- Is there a large gap between the theory we have learned so far in the course and what the paper identifies as statistical thinking? Where is the gap, and can we close it without sacrificing technical rigour?
- How does the interrogative cycle relate to formal hypothesis testing? Is testing a subset of it or a parallel activity?
- The paper is now almost 30 years old. Would it change if written today, given the intense focus on predictive statistics, AI and machine learning?
2. Ziliak and McCloskey (2004) – Size matters: the standard error of regressions in the American Economic Review (with discussion)
Ziliak, S. T., & McCloskey, D. N. (2004). Size Matters: The Standard Error of Regressions in the American Economic Review. Journal of Socio-Economics 33(5), 527–546.
See also the accompanying comments and replies in the same special issue and Gigerenzer’s “Mindless statistics” (next item) which appears in the same issue and attacks significance testing from a different but complimentary angle.
Ziliak and McCloskey audit every full-length empirical paper in the American Economic Review in the 1980s and 1990s and report that roughly four out of five conflate statistical significance with economic (substantive) significance. The whole special issue, including comments and rebuttals, is a pedagogical gem (see also the next item). The central message, that an estimated effect can be statistically significant and economically trivial, or statistically insignificant and substantively decisive, is now formulated in the famous American Statistical Association (ASA) Statement on p-values and the 2019 ASA editorial (see item 9 below).
- State precisely what Ziliak and McCloskey mean by “oomph”. Why does a test of \(H_0: \theta = 0\) fail, by itself, to answer the question they are asking?
- Read Wooldridge’s comment in particular. Which points of his land a blow, and where, if anywhere, does he miss?
- Audit three recent empirical papers in your own field against the Ziliak–McCloskey checklist. Report the results honestly, without editorial softening.
- Here is an exception to the rule about AI use in this task. Set up a bot that finds the empirical papers in AER from 2000 until today and repeat the questionnaire to all of them. What are the recent trends?
3. Gigerenzer (2004) – Mindless Statistics
Gigerenzer, G. (2004). Mindless Statistics. Journal of Socio-Economics 33(5), 587–606.
Gigerenzer argues that empirical social science has replaced statistical thinking with a statistical ritual, what he calls the “null ritual”: (1) set up a null hypothesis of no difference of no effect, without specifying your own hypothesis or any alternative, (2) use the 5% significance level mechanically, and (3) always perform this procedure regardless of the research question. The ritual, he shows, is not Fisher’s theory, not Neyman-Pearson theory, and not any other coherent framework, but rather an incoherent hybrid of incompatible elements, institutionalized through textbooks, editorial policies and disciplinary norms. The result of this mash-up is a procedure that neither Fisher nor Neyman/Pearson would endorse yet is what most empirical researchers (in particular in psychology which is the setting of this paper) do. The paper sits in the same special issue as Ziliak and McCloskey and shares their target (the mechanical use of significance testing), but attacks it from a different angle. Where Ziliak and McCloskey emphasize the neglect of effect magnitudes, Gigerenzer emphasizes the corruption of the logic of inference itself. The combination is powerful because the two diagnoses reinforce each other. A researcher who confuses Fisher with Neyman–Pearson is also unlikely to stop and ask whether a significant coefficient is large enough to matter.
- Chapter 8 and 9 of our textbook (not part of the curriculum per se) present testing as a unified framework. How does Gigerenzer’s historical account of the Fisher vs. Neyman-Pearson dispute complicate that presentation? Where does the textbook lean Fisher, where Neyman-Pearson, and where does it (quietly) hybridize?
- Gigerenzer claims that textbook authors propagate the null ritual by presenting a hybrid theory as if it were a single coherent framework. Examine two statistics or econometrics textbooks you have used and evaluate this claim.
- Read Gigerenzer alongside Ziliak & McCloskey (item 2). Both appear in the same special issue. Where do their critiques overlap, and where do they diagnose different pathologies? Could one be solved without the other?
- Gigerenzer draws an analogy between the null ritual and compulsive behavior. Is the analogy fair or inflammatory? What institutional mechanisms sustain the ritual, and what would it take to break them?
- Is the Wasserstein et al. (2019) ASA editorial (item 9) the institutional response Gigerenzer was calling for, or does it leave the deeper confusion about Fisher vs Neyman–Pearson untouched?
- This paper mainly deals with the field of psychology? To what degree does it apply to modern empirical economics-, business-, and finance research?
4. Breiman (2001) – The Two Cultures (with discussion)
Breiman, L. (2001). Statistical Modeling: The Two Cultures. Statistical Science 16(3), 199–231 (with discussion by Cox, Efron, Hoadley, and Parzen, and rejoinder).
Breiman partitions the field into a data-modeling culture (which assumes a parametric data-generating process and then does inference about its parameters) and an algorithmic-modeling culture (which treats the data-generating process as unknown and evaluates procedures by out-of-sample predictive accuracy). He argues that 98% of academic statisticians and nearly all econometricians live in the first culture, and that this has made the field less accurate and less relevant. The published discussion is essential. Cox (who is one of the really big names in modern statistics, and who was 77 years old in 2001) defends interpretable models on epistemic grounds. Efron (another giant) pushes back with precision, by trying to tone down the state of the “emergency”, while Hoadley and Parzen bring applied perspectives. If you are already half-trained in machine learning, this paper offers an early clarification of the distinctive features of inference/explanation and predictive statistics (see also item 14 by Shmueli).
- Where, concretely, do data-modeling commitments appear in our textbook (CB)?. Which results genuinely require them, and which survive a weakening of the assumptions?
- Breiman emphasizes the “Rashomon effect”, which is that many models may have indistinguishable predictive accuracy but different internal structure. What does this mean for interpretation, causal inference, and scientific communication?
- Read Cox’s and Efron’s discussion pieces. Whose position has aged better and why?
- Do modern hybrid methods (double/debiased machine learning, causal forests, etc) reconcile the two cultures or establish a third?
- Take a published paper that uses linear regression for inference and sketch the “algorithmic culture” version of the same paper. What is gained and what is lost?
5. Leamer (1983) – Let’s Take the Con out of Econometrics
Leamer, E. E. (1983). Let’s Take the Con Out of Econometrics. American Economic Review 73(1), 31–43.
Leamer considers the gap between the textbook view of econometrics (clean specifications, pre-specified hypotheses), and the reality of iterated specification searches, in which the researcher explores dozens or hundreds of models and reports the one whose signs and t-statistics look most respectable. Rather than recommending abandonment, he proposes extreme bound analysis, which means to report the range of coefficient estimates across all plausible specifications, and call “sturdy” only those inferences that survive. The idea of looking at the map from the assumption space to the inference space is particularly interesting. The paper launched a literature on model uncertainty, sensitivity analysis, and reporting norms, and sits upstream of both the credibility revolution (see item 10 below) and the garden-of-forking-paths literature (see item 8 below). Its diagnosis of how applied econometrics actually gets done remains uncomfortably current.
- Formalize Leamer’s concern. If a researcher runs \(K\) specifications and reports the one with the smallest \(p\)-value, bound the true size of the nominal 5% test as a function of \(K\) and the correlation structure of the estimates.
- Leamer distinguishes doubtful, fragile, and sturdy inferences. Apply the taxonomy to one empirical regression that you have run or read carefully.
- Extreme bounds analysis has itself been criticized (see, e.g., Sala-i-Martin, 1997, “I Just Ran Two Million Regressions”). Read one critique and one defense and summarize where the debate has settled.
- Draft a reporting standard for applied empirical work in your field that you believe would meet Leamer’s objections without paralyzing research.
6. Freedman (1991) – Statistical Models and Shoe Leather (with discussion)
Freedman, D. A. (1991). Statistical Models and Shoe Leather. Sociological Methodology 21, 291–313
Freedman uses John Snow’s 1854 investigation of the Broad Street cholera outbreak as the positive case, with careful accumulation of contextual knowledge, clever natural comparisons, and the use of shoe leather, producing a causal conclusion that survives without any formal inferential machinery. He contrasts this with the modern (in 1991) reflex to reach for a regression with demographic controls as soon as a question is posed, and argues that when the assumptions behind the model are not credible, the model converts ignorance into false precision rather than resolving uncertainty. The paper is short, provocative, and generous with concrete examples. Its follow-ups (From Association to Causation published in Statistical Science in 1999 and the book Statistical Models and Causal Inference) extend the argument and should be consulted before writing a first thesis regression.
- Reconstruct Snow’s inferential argument without any formulas. Which pieces of evidence do the most work? Which background assumptions are you relying on?
- Freedman can be read as anti-modeling. Is he? Distinguish his target from uses of regression he would endorse.
- Does “shoe leather” scale? In an era of administrative datasets, digital traces, big and messy data, and automated pipelines, what is the analog of walking the streets of Soho?
- Read one discussant who disagrees with Freedman. What is the strongest response to the shoe-leather view, and where does it leave model-based inference?
7. Ioannidis (2005) – Why Most Published Research Findings Are False
Ioannidis, J. P. A. (2005). Why Most Published Research Findings Are False. PLoS Medicine 2(8), e124.
This is the paper that pushed the replication crisis into the public debate. Using Bayes’ rule, elementary probability and stylized assumptions about power, bias, and base rates, Ioannidis derives the positive predictive value (PPV) of a “significant” finding and argues that in many research fields it is below 50%. The formal model is simple enough, but the impact of this work likely comes from the transparency of the derivation combined with the somewhat dark conclusion. The paper has spawned a large follow-up literature (including critiques: Jager & Leek 2014, Goodman & Greenland 2007), the rise of meta-science as a discipline and institutional reforms such as pre-registration and registered reports; see the wikipedia page for this paper for links to further reading: https://en.wikipedia.org/wiki/Why_Most_Published_Research_Findings_Are_False.
- Ioannidis assumes independent studies. Relax this. What happens under correlated biases, shared datasets, or a common research protocol?
- Which of Ioannidis’s six corollaries apply to empirical research in business economics? Which to OR or operations management?
- If Ioannidis is approximately right, which institutional reforms offer the largest expected improvement, and which are cosmetic?
- Read Jager & Leek (2014) or Goodman & Greenland (2007), articulate their critique and discuss where it lands. Does it undermine the headline conclusion?
8. Gelman and Loken (2014) – The Statistical Crisis in Science
Gelman, A., & Loken, E. (2014). The Statistical Crisis in Science. American Scientist 102(6), 460–465. (The longer 2013 technical working paper, “The garden of forking paths,” is also worth consulting.)
The authors clarify a distinction that earlier p-hacking debates has blurred: a researcher with no intention to cheat, looking at a single dataset and running what feels like a single analysis, is nonetheless making many conditional choices (which variables to include, which subgroup to emphasize, which outliers to drop, etc) that are themselves data-dependent. The resulting “garden of forking paths” inflates Type I error in ways that standard multiple comparison corrections cannot catch, because the unexplored branches are counterfactual rather than actual. The longer technical paper gives a more careful analysis of the problem.
- Distinguish carefully among (a) classical multiple testing, (b) active p-hacking, and (c) the garden of forking paths. Do any tools from the course so far address one or more of these failure modes.
- Why do Bonferroni and false-discovery rate (FDR) corrections fail for (c). What would actually solve it, and at what cost?
- Construct a plausible forking-paths story for a study in your field, listing the (unchosen) branches the researcher might have taken at each decision point.
- Pre-registration is the canonical fix. What does it cost? For which kind of research is the cost too high?
- Is honest exploratory data analysis compatible with garden-of-forking-paths concerns? Draft rules for a student who must explore the data before committing to a pre-registered analysis.
9. Wasserstein, Schirm & Lazar (2019) – Moving to a world beyond “p < 0.05”
Wasserstein, R. L., Schirm, A. L., & Lazar, N. A. (2019). Moving to a World Beyond “p < 0.05.” The American Statistician 73(sup1), 1–19. Editorial introducing the special issue Statistical Inference in the 21st Century, which contains 43 further papers on alternatives.
The editorial takes the unusual step of recommending against the use of the phrase “statistically significant”. It does not recommend abandoning p-values as a whole, but it does unconditionally reject the automatic dichotomy of p \(<\) 0.05 vs. p \(\geq\) 0.05, which should be fairly obvious to any scientist. The special issue surveys proposed alternatives and the editorial argues for pluralism, transparency, and above all: thoughtfulness. Together with the 2016 ASA Statement, this is the current mainstream-statistical position on significance testing, and a direct descendant of the McCloskey–Ziliak criticism discussed above.
- What is the ASA not saying? Read the editorial carefully and distinguish a ban on thresholds from a ban on p-values.
- For each of the main alternatives (estimation, Bayes factors, second-generation p-values, decision-theoretic thresholds) identify one situation in your field where it would be an improvement. Is it also possible to find a situation in which it would be worse?
- How does the recommendations in this editorial stand in relation to the Ziliak-McCloskey paper (item 2)?
- Simulate the writing of (the meat of) one recent paper in your area under the ASA recommendation. What changes substantively, and what is cosmetic?
- Is there a defensible role for a bright-line threshold in regulated decisions (drug approval, engineering standards, audit certification)? If so, what distinguishes those contexts?
10. Angrist & Pischke (2010) – The Credibility Revolution in Empirical Economics
Angrist, J. D., & Pischke, J.-S. (2010). The Credibility Revolution in Empirical Economics: How Better Research Design Is Taking the Con out of Econometrics. Journal of Economic Perspectives 24(2), 3–30. Read with at least one of the companion critiques in the same issue.
A self-conscious answer to Leamer (item 5), written after two decades in which applied microeconomics reorganized itself around credible identification using randomized experiments, instrumental variables, diff-in-diff, and regression discontinuity. The authors document shifts in method and in reporting style across labor, development, and public economics, and claim substantial gains in the credibility of causal statements, while noting that similar shifts have not occurred in macro and much of finance. The piece is the optimistic insider’s narrative and should not be read alone: the same JEP issue contains critiques, for instance from Leamer himself, that attack the methodological and epistemological basis of the claim. Together they make a good short introduction to the contemporary methodological divide within empirical economics.
- Compare the Angrist–Pischke narrative to Heckman’s and Deaton’s critiques. Where is the disagreement empirical, where philosophical, and where rhetorical?
- How far has the credibility revolution spread outside labor and development, for instance into finance, IO, macro, marketing, operations management?
- Does the credibility revolution answer Leamer’s original concern, or merely displace the specification search to a different location (choice of instrument, bandwidth, cluster level, event window)?
- Sketch a research agenda in your own field that would meet both Leamer’s and Angrist-Pischke’s standards simultaneously.
11. Manski (2011) – Policy Analysis with Incredible Certitude
Manski, C. F. (2011). Policy Analysis with Incredible Certitude. The Economic Journal 121(554), F261–F289.
Manski surveys the tendency of economists, forecasters, and policy bodies to issue point predictions with an air of certainty not supported by the underlying evidence. He catalogues five types of “incredible certitude” – conventional, dueling, conflating, wishful, illogical – and argues that partial identification, that is, reporting honest bounds derived only from credible assumptions, would be more informative and more honest. The paper shifts attention from sampling uncertainty to identification uncertainty, which can be much larger and invisible compared to standard confidence intervals. This paper serves as a useful counterweight to a course like ours where we always assume that we can identify estimators.
- Work through a simple partial-identification problem (such as bounds on a mean when outcomes are missing not at random, or Manski’s bounds on a treatment effect) and contrast with the point-identified version.
- Which of Manski’s five types of certitude do you most often (expect to) encounter in your own field? Give examples.
- How should a policy audience be presented with bounds rather than point estimates? What institutional, rhetorical, and political barriers stand in the way?
- Does Casella-Berger implicitly assume point identification? What would Chapter 9 look like rewritten with bounds as the default object of inference?
- Identify a current policy claim in your area and attempt to bound the underlying quantity using only assumptions your are willing to defend. Report what you learn from the exercise.
12. Box (1976) – Science and Statistics
Box, G. E. P. (1976). Science and Statistics. Journal of the American Statistical Association 71(356), 791–799.
The first signs of perhaps the most quoted sentence in all of statistics (“All models are wrong, but some are useful”, it appeared in this exact form a few years later in another paper by Box) can be found in this paper, but it is far richer than this. Box argues that science advances by an iterative loop of conjecture, estimation and model criticism, and that the role of a statistical model is not to be “true” but to be a useful approximation that helps this loop turn. He distinguishes sharply between estimation (fitting parameters within a model assumed adequate) and model criticism (checking whether the model itself is adequate), and argues that the second is at least as important as the first but receives far less formal attention in statistics courses. For a student deep into the world of Casella and Berger, this paper offers an interesting escape. The tools you are learning are real and important, but they live inside a larger scientific process of model building, criticism and revision the gives them their meaning. It answers the question “why am I doing this?” in a way the textbook does not attempt. A large part of the (not very long) paper is a historical account of the early work by RA Fisher, which is used to illustrate the points.
- Box distinguishes estimation from model criticism. Where in CB does each activity appear? Which receives more formal attention, and, if there is an imbalance, is this a problem?
- What does “all models are wrong but some are useful” mean operationally? How do you decide, in a concrete application, when a model is useful enough – and useful enough for what, exactly?
- Box describes science as an iteration between conjecture and criticism. Map a research project in your own field onto this loop. Where does formal statistical inference enter, and where does it need help from other sources of knowledge?
- How does Box’s philosophy relate to Breiman’s two cultures (item 4)? Does the algorithmic culture have an analog of the model-criticism step, or does it bypass the loop entirely? (Note the remarks on “model-free tests” in the paper).
- Is Box’s iterative view compatible with pre-registration? If a researcher commits to a model in advance, has the loop stopped turning, and is that a gain or a loss?
13. Kass (2011) – Statistical Inference: The Big Picture
Kass provides his view of the entire business of statistical inference, explaining how frequentist, Bayesian, and likelihood-based approaches relate to each other and to the practical problems they are meant to solve. He argues that inference is fundamentally about calibrating the strength of evidence for or against claims about the world, and that the competing frameworks are best understood as different strategies for achieving that calibration, each with characteristic strengths and blind spots. He shows where the frameworks agree (more often than their partisans admit), where they disagree for real, and why the choice of model usually matters more than the choice of framework. The paper is short and easy, and explicitly does not choose side – but rather formulates a common view of statistical inference and in particular how this view can improve teaching of basic statistics.
- Kass argues that the choice between frequentist and Bayesian matters less than the choice of model. Do you agree? Under what conditions would the framework matter more?
- How does Kass’s “big picture” relate to the specific critiques in the significance-testing literature (see other items in this list)? Does seeing the whole map make the critiques more urgent or less?
- Write a one-page “map of inference” for a fellow student in your program, using Kass as a starting point but adding the specific methods and debates relevant to your own field.
14. Shmueli (2010) – To explain or to predict?
Shmueli, G. (2010). To Explain or to Predict? Statistical Science 25(3), 289–310.
Shmueli argues that explanation (testing causal hypotheses derived from theory) and prediction (forecasting new observations accurately) are distinct scientific activities that call for different modeling choices at every stage. Conflating them (which she claims is the default in much of social science, including economics and business) produces models that are suboptimal for both. The paper lists activities where choices diverge: an explanatory model should be interpretable and theoretically motivated, even at some cost of fit; a predictive model should maximise out-of-sample accuracy, even at the cost of interpretability. The bias-variance trade-offs facing the two goals pull in different directions, so a variable correctly retained for explanation may be correctly dropped for prediction, and vice versa. Where Breiman (item 4, explicitly cited in this paper) frames the field as divided between two cultures of modeling, Shmueli reframes the issue as a division of purposes that cuts across both cultures: a regression can be used explanatorily or predictively, and likewise for a random forest. It is the design decisions that matter, and they are formulated in the same way, but answered differently, under both regimes.
- Shmueli classifies statisticians as either considering prediction as the main purpose of statistical modeling or unacademic. Is this dichotomy helpful, or maybe just false?
- Pay close attention to the constructive tone under point 6 at the end of Section 1.4: Good or bad predictive power of an explanatory model is informative towards the potential for new advances. Compare this to the harder statement (very generously paraphrased from e.g. Clarke & Clarke (2018): Predictive Statistics) that a model that does not predict well is essential a circular argument that only serves to encode the existing opinions of the empirical researcher.
- Pick a published paper in your field and classify it as explanatory, predictive, or confused. What would you change to align its methods with its stated goal?
- Compare Shmueli’s “explain vs predict” to Breiman’s “two cultures” (item 4). Are they describing the same distinction with different vocabulary, or real different partitions of the field? Construct an example that the two schemes classify differently.
- Take a question from your own area of study. Draft two research designs, one explanatory and one predictive, and report where the two research designs diverge.