top of page

Stats/Data Processing Tips

hannahkish6
Sep 16
4 min read

By Sanna Eriksson


Data processing and statistics are one part of science that a lot of people don’t like, and a lot of people find really intimidating. Luckily for me, I loved maths throughout high school, and my brain naturally likes things like coding and calculations. For me, analysing data -- from cleaning up the raw numbers a machine gives you, to running stats, all the way to visualising the results in a nice figure -- is a super satisfying process. It takes the hard work you’ve done in the lab or in the field and actually gives you a result, and something to show yourself and others what you found. So, for today’s blog, I thought I would take you through some tips I have to making this process enjoyable and productive!

Working out how you want to visualise your data can be tricky, especially if you are working with big datasets and have multiple variables that could explain the patterns in your data. The first thing I do when I finish collecting data (or even before!), is to draw out on paper what sorts of graphs/figures I think would best show the answers to the questions I’m trying to ask. This is a really helpful task to do before you’ve even started the experiment or sampling, because it can help you to understand what you can actually get out of the data you’re planning to collect, and whether you might need to add something in, or change something around. I find it’s easier to do this on paper and then think about how to make the figure on the computer. Then I’ll take these ideas, and try to make them come to life in RStudio.


Examples of some early ideas for how I might visualise my data!


This will depend on where you’re doing your research, and what your supervisors’ preference is, but my go-to stats software is R (using RStudio). It’s a super useful, open access software, where people can add their own packages and functions for others to use, and there are sooo many different analyses you can do on there. I would definitely recommend doing an intro to R course (maybe your university offers one, otherwise there are plenty online), to get an understanding of the basics first. Once you are familiar with this, the opportunities are endless, and there is a crazy amount of information online to help with specific analyses, as well as more and more AI resources joining the mix. R is great for every step of the data analysis, from initial data cleaning, to running statistical analyses, to visualising the data. It will feel like a big learning curve, but once you get the hang of it, it should make your stats life so much easier.

My biggest piece of advice for people starting out with R, or even those who already have the basics down, is to document every step of the code as you go. In your code (see photo below for an example), you can add lines that begin with a hash (#) and these lines don’t run as code, they are recognised as notes/comments. In my code files, I include comments for every new function I use (and sometimes those that I have used before but feel I’m still learning), reminding my future self of what each input in the function does, and how it actually works. I also leave comments explaining WHY I used a certain function, or why I’ve analysed something in a particular way. This might feel silly at first, but trust me, when you come to look at the code in a few months’ time, or honestly sometimes even days after you wrote it, you will be so grateful for these comments.


An example (admittedly an extreme one) of the notes I leave for myself explaining how a new function works, so that next time I look at this code, and maybe want to make adjustments, I can remember what’s going on.

 

With all of that said, I think it’s important to remember that the process of taking your raw data and turning it into a stats result and a nice figure is a long process, with lots of hidden steps. As much as there are some aspects with very specific solutions, like choosing the right stats test for your data, there are also many parts that can actually look very different for different people. Often there are many different ways that you could visualise your data for example, and finding the one that you think best represents the information you’re trying to get across, is something that can take a lot of trial and error. Hopefully some of the tips I’ve shared here will help and make the process more enjoyable – to me, there’s something so satisfying about finally landing on a figure that you’re proud of!


Two of the figures from my first PhD chapter that took a lot of planning and trial and error to get to, but now I look at them and see all of that hard work paid off!

 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page