Friday rabbit hole caveats be damned healed, a reference to p-values in my readings this week struck a chord for ready digression.
I interviewed a statistics researcher [podcast versions] a few years ago who served as an co-editor on special journal issue dedicated to the (mis)use of statistical significance in research, particularly psychology research and specifically the phenomenon know as p-hacking:
a little more explanation of the null hypothesis and how it relates to other flaws in the typical way psychology researchers have handled statistical analysis over the past several decades. Roughly, our research often comes down to finding differences between groups of participants on some variable of interest (Are boys better at math than girls?) or correlations among two or more variables (Does suicide risk rise as poverty increases?). The null hypothesis is the assumption we then make that there are no such differences or correlations in the populations from which our data were drawn, and we compute the probability that we would have found data with differences or correlations as large as we actually did if the null hypothesis were true. If that probability (called a “p-value”) is less than some specified value, typically .05, then we reject the null hypothesis and conclude that the differences or correlations we found are probably real. This is known as a statistically “significant” result. If the p-value is greater than .05, then we fail to reject the null hypothesis, for lack of convincing evidence otherwise.
Research that fails to reject the null hypothesis rarely gets published in scientific journals. Likewise, research that merely replicates rejections that were found in earlier research rarely gets published. As a consequence, there is a high professional premium on testing questions that have not been previously published, and on rejecting the null hypothesis.The widespread adoption of these customs has unintentionally encouraged a whole series of inadvisable research practices, which have led to trouble. For instance, because most studies that do not find significant effects go unpublished, we don’t know whether the findings of the published research are real or just the mistakes that our chosen p-value allows to slip through (the “file-drawer problem”). This problem might be attenuated if we published efforts to replicate previous findings, but that kind of research is rarely accepted by journals.
Note the generosity embedded in the term ‘inadvisable.’ Hmm.
