A PhD is all about statistics
We, as a collective, are chasing that p-value high (or low?).
A quick background, statistics, at its basis, is comparing groups and seeing if the difference between them is random (not-significant) or not random (significant). And the cut off for that is a p-value of less than 0.05.
There is a lot of misconceptions about statistics. I think one of them is that if you include a percentage, it tells the whole story but saying something is ‘statistically significant’ has nuance. There are so many ways to categorise your data, or do data collection, or just generally explain your data that statistics can be really misleading if not reported correctly and transparently. For example, take the below acorn population. You are comparing the number of black and white acorns. You sample 12 trees for each site and count the number of black and white acorns. If you report the total number acorns across all sites, you’d have 40% of black, 60% of white.

But say you live in Site C, so you only look at this site. Then you’d report back with 100% of black acorns.
Or say you only counted the acorns on the ground. But if you included the acorns in the trees…

…it would be 60% of black, 40% of white.
Each example gives a different percentage of black to white acorns. No example is wrong, per say but they do tell different stories about the acorn population. This bias in data happens in all industries; government, science, news, anything that reports numbers. It’s not inherently wrong, but it is good to be aware and look for how the data was collected and what are they actually reporting. Are they reporting the colour of acorns at only one location? As we saw in the above example, this isn’t representative of the acorn population as a whole.
Or maybe a company is reporting its return on investment in one year. This might not be reflective of the company’s total earnings across all of their products.
And in all honestly, that 0.05 p-value is an arbitrary number that some guy way back when decided was ‘good enough’ (https://medium.com/@pritulp/the-true-history-of-p-values-and-hypothesis-testing-what-every-a-b-testing-practitioner-should-a82ffe35dec9).
But back to PhD’s and stats. We all must do it and there is a learning curve. Because just as there is a million different ways to answer a hypothesis, there is a million and one statistical tests out there. And selecting the right test is crucial. It can feel like a lot of pressure in making sure you do the right one. This month, Hannah and Sanna are going to talk about their experience with stats and give some insights on how they’ve approached this math section of their biological PhD.

Comments