Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Monday, August 4, 2008

A Simple Question

So here’s a statistical question for you. Now, mind you, this is a question that ground my graduate-level experimental design class to a halt for at least half an hour, and that also brought my boyfriend and I into something resembling not a discussion, but an actual argument. Pretty heavy stuff, for statistics.

Here’s the situation: You are a botanist and you want to study the effect of two different light regimes on petunia growth—let’s say 12h light and 12 dark, and 18h light and 6h dark. You have 40 little petunia seeds planted in pots waiting for you, and your university has 2 environmental chambers for you to use. You put 20 pots in each chamber, set the light timers, and start the experiment. After the prescribed number of days on this program, you measure all your petunias, and begin to analyze the data.

Now here’s the question, and I’ll phrase it a couple different ways: How many experimental units do you have? In other words, how many independent data points? How many degrees of freedom will you get in an analysis of these data?*

Answer: 2 experimental units, 2 independent data points. 0 degrees of freedom.

“What???!!” you may splutter. “But there were 40 petunias!!” You thought there were going to be 40 experimental units and 19 degrees of freedom, didn’t you?

Well, what did happen to those pots of petunias? The problem stems from when they were all put into only 2 environmental chambers. Once in an environmental chamber, the light turned on and shone into the entire room. The light was applied to all the pots together, as a group. If the lightbulb, say, started flickering and going out in one room, it would be flickering over all those plants together. In other words, the treatment (the light) was applied to one unit, the room. Therefore the environmental chamber as a whole becomes the unit of experimentation, not the individual plants. If the experimenter were to ignore this, he would be committing the mortal, yet frighteningly common sin of pseudoreplication.

Pseudoreplication occurs when there is a lack of independence between supposed experimental units, when the treatment is applied collectively, not individually. What this means, practically, is that each of your little units you are assuming are independent are actually irrevocably linked to each other, in a way that can mask the effect that you’re actually trying to see.

Now let’s suppose that there’s a problem with one of the lights, the one in the chamber on the 12/12 regime. That light tends to flicker when the MRI machine next door gets turned on. It’s new, and people don’t generally tend to hang out in environmental chambers and read the Times and have a coffee, so it hasn’t been noticed yet. But every single petunia in that chamber collectively feels all of those light flickers. In fact it happens frequently enough that it negatively affects their growth—all of them, together. So when the data are collected, the plants in the 12/12 room are just a little bit shorter than they might have been otherwise. Their growth was stunted just enough that those plants’ heights are less than those of the plants in the 18/6 room. When the experimenter analyzes those data (not realizing yet that the experiment is pseudoreplicated), he finds a significant difference between the two and concludes that an 18/6 light regime for petunias helps them grow taller. What he doesn’t realize is that he hasn’t detected a difference due to light regime, he’s detected a difference due to faulty wiring—not at all helpful. Incidentally, even if he did finally recognize the pseudoreplication, he wouldn’t be able to analyze the results. With only one (true) experimental unit in each light regime, he wouldn’t be able to take an average and compute the variation around that average—there’s no variation because with only one data point, there’s nothing to vary. Without that, he can’t figure out if his two treatments truly are different from each other outside of the range of normal background variation. No conclusions can be made, and the entire study is wasted.

It’s easy to imagine other situations in which the pseudoreplicated nature of this study could screw up the results: a careless undergraduate props the door to one of the chambers open for a minute, forgets about it when his girlfriend calls, and then goes to lunch. In the meantime that room loses all its humidity through the open door. Or one of the lightbulbs burns out and nobody notices it for 8 hours. Et cetera.

Experimental units are independent when treatments are applied to each one individually. If this study used little light lamps for each plant, they would truly be the independent experimental units, because each one would be receiving an independent treatment. If one of the bulbs flickered and screwed up the growth of that one plant, the results overall may not be affected much, because there’s still 19 other independent data points in each that will all be averaged with the screwed-up one. Not ideal, but not the end of the world. Doing it this way sounds like a lot more work, but sometimes correct experimental design calls for a little more creativity and effort in order to get it right, and get valid results.

You roll your eyes and tell me that I’m being entirely impractical and unrealistic. “OK let’s assume that this is a well-funded university that can afford a decent electrician. Everything in the rooms has been tested and checked out. They’re fine. They’re completely monitored in every way so that if something goes wrong it’ll be noticed immediately and fixed. Stop being such a curmudgeon.” Yes, probably everything will be fine. But what if there is some variation that you don’t know about yet? You can’t monitor something you don’t know of. You have to design your study well enough, and with all precautions in place, to take care even of the most unforeseen circumstances. Only then can you get results that prove what you say they prove, with as much confidence as you think.

Pseudoreplication is everywhere. For example, a major study in my thesis area is pseudoreplicated, and sometimes I wonder if I’m the only one who’s noticed. (A developmental hormone was applied to some insects. The experimenters squirted hormone onto filter paper in the bottom of a Petri dish, and let groups of insects walk around on it and absorb it. Thus, the experimental unit here was not the insect, It was the Petri dish. But you can bet that each insect was treated as independent in the statistical analysis.) I know that sometimes I tend to lazily skim over the methods sections in papers to get to the conclusions. It’s a temptation, and a strong one, too, when there’s so much to read and so much else to do. But so much can go wrong in those dry methods sections. If we biologists can’t be trusted to always remember the lessons of our statistics classes way back in grad school, then all of us have to be on guard to catch our colleagues’ mistakes, before those unnoticed mistakes become accepted and cited in future research, even though they may well be completely erroneous.











*”data” is plural. “These data.” Not “this data.” Really. Don’t be That Guy**

**In normal situations “That Guy” might refer to the dude at the bar with his shirt tucked into his underwear who can’t figure out why all the girls are shooting him down. In nerd circles, it refers to the person who uses “data” as a singular noun. Hopefully it’s not the same person who also has his shirt tucked into his underwear, or he’ll never get a date.

Monday, July 28, 2008

Stats 101 for Journalists--Correlation vs. Causation

While perusing your favorite newspaper, you may have run across an all-caps, bold-print headline with a title something like this: EATING SPINACH EVERY DAY WILL PREVENT CANCER, DOCS SAY. Generally, these articles will include speculation from researchers on how exactly this miracle food will keep you cancer-free; perhaps its those antioxidants, perhaps its high fiber content. Whatever the reasoning, the implication is that you should run out immediately to the grocery store and commence a daily diet reminiscent of a rabbit’s.

Not that there’s anything wrong with spinach. Your mom and popeye were right, spinach is very good for you, in fact. And hey, maybe it is a key part of a diet that will aid in preventing cancer.

The problem is, in fact, a statistical one. A common statistical error that perennially causes stats profs to tear out their hair in frustration, or perhaps, if they’re old and jaded, to merely roll their eyes and shrug.

The problem is the inference of causality from correlation.

Most likely, the study had a design something like this: hundreds of people were followed throughout a number of years, and periodic surveys were sent to them asking them about their diets. They filled out the form stating how many times a week they ate certain foods, and sent it into the study center. Or maybe they got phone calls from research assistants, asking the same questions. But regardless, it was not an experimental study—that is, nobody put these hundreds of people in cages and gave them different kinds of diets, each with different amounts of certain foods. It was observational, meaning that the researchers worked with what they could get—the pre-existing diets of their study volunteers, over which they had no control. Instead of creating and administering different conditions to them, they scientists just watched the subjects and saw what happened. Their data allowed them to correlate a factor with an outcome, but not prove causation.

This may seem like an academic difference but it has far-reaching implications. In experimental conditions—say, working with mice in a laboratory—all of the conditions are carefully controlled, in order that any effects can be attributed exactly to a cause. Say that ethics regulations allowed scientists to put people in cages and experiment on them to see the effects of spinach on cancer. Every person would receive the exact same cage conditions: exact same lighting, medical treatment, air temperature, amount and type of exercise, etc etc etc. And they would receive the exact same diet—except for one key difference. Half of the caged experimental humans would receive a diet that had more spinach than the other group’s. Then, after many years of monitoring under these same conditions, if there were any difference in cancer rates between the two groups, this could be attributed exactly to the one difference that existed between the groups—that of spinach consumption. The only way to infer causality is through experimentation—manipulating conditions in a controlled manner to see what affects these differences have between groups.

Obviously, because of ethical and monetary restrictions, this kind of study design with humans is impossible. So why can’t you infer causality from observational studies—the type of survey study that was carried out to create the flashy newspaper headline? The problem is that nothing is controlled in the research subjects—you don’t know if they have the same conditions at home, the same income level, the same amount of exercise, the same anything. What if it is not the spinach that is causing some people to have lower rates of cancer, but something else, that happens to be associated somehow, coincidentally, with spinach consumption? Fresh vegetables are expensive. They also require more time, generally to prepare—to wash, cut, etc. What if its not the fact that the cancer-less people are eating spinach, it’s that they can afford to have more fresh vegetables in their diet because they are wealthier, and maybe their extra wealth allows them to see the doctor more frequently? Or what if the extra little bit of time they have in their day that allows them the time to prepare fresh vegetables like spinach also happens to be enough extra time to go jogging as well? Any number of other, hidden, things could be the actual cause, or one of many causes of the lowered cancer rates in these people. The spinach may have nothing to do with it; it may just have been associated somehow with the actual, unrecorded cause.

Experimental studies, which allow true inference of causality, are impossible in many cases when the study animals are human beings. The best correlational studies looking at human habits and disease outcome over many years have huge numbers of people and try to get as much information about their participants as possible—background health info, income, marital status, exercise habits, etc, in order to take all these factors into account. And they often find very interesting and useful results, linking certain types of diets, lifestyles, or exercise habits to long-term rates of disease. But no matter how well these studies are designed and carried out, no newspaper can ever report on their findings using the word “cause.” Even if they record as many different variables as they can think of from their study participants—exercise, religious beliefs, geneology, length of their little toe, etc, it’s impossible to know whether or not they recorded any information about the factor that is truly causing the differences seen in the study. The conditions and the participants themselves are just too variable. To talk about causation in this context is simply incorrect, and perhaps even false.

What’s then the use of these large-scale observational survey studies? These studies are useful in finding links to diseases, which can then be studied directly in an controlled experiment using mice—which are 80-some-percent genetically related to us. Once this same connection is found in a controlled, experimental environment, one can finally come to some conclusion about causation.

I had a stats prof who had written his master’s thesis on biologists’ understanding of statistics. He found that over 70% of research published over several years in a peer-reviewed biological journal had statistical errors. It’s no surprise then, that newspaper writers are prone to the same kinds of statistical mistakes. It’s then up to the discerning reader to look beyond the headline, dig a little deeper, and figure out if the research was carried out in such a way as to merit the flashy headline. Dramatic words like “causes” and “leads to” and even just “will” sell newspapers. But they may not be statistically and scientifically accurate—be smart and judge for yourself.