Showing posts with label hypothesis testing. Show all posts
Showing posts with label hypothesis testing. Show all posts

Tuesday, June 16, 2009

Hypothesis Testing

Why does this work, testing variables to see if they come from separate populations?

We get probabilities that enable us to say that our findings are significant because we believe that all things fall out, for the most part, normally. If I take an infinite number of samples from a population and test for any one thing, then make a frequency distribution of all of the findings, that distribution will look like a normal curve.

This is called a sampling distribution, and it is usually of the parameters, any parameter (mean, median, mode, SD, variance).

Test results of infinite samples will show that half of the samples will fall out on one side of the average, that center line that bisects the top of the normal curve and the horizontal line, (see why you once had to learn geometry, so you could know the word bisect) and the other half will be on the other side of that center line.

So every time you test a sample, you can compare your results by comparing them to what would be the theoretical average, the norm, the normal curve. You're comparing two means, or two averages, to see if the differences are so great that they probably don't belong in the same sampling distribution.

You either took a sample and divided it up into two groups to intervene in one and not the other and compare, or you are merely testing one group to see if variables within that group are associated with one another).

Your null hypothesis is that the difference between these means is minimal, they are so alike that they must come from the same population. What you did, or what you found, is no big deal, 95-99% of the population is this way.

If you find huge differences, however, that perhaps only 5% of the sampling distribution would have the trait you are testing (attitude, eye color, anger) associated with another, then this five percent is from another distribution, one in which 95-99% of the people in it are just like this.

You've established a new population, basically. What you did mattered, it is significant that two things are associated. Either what you did made the difference, or these things really do hang together in some normal distribution. Either way, the likelihood is greater, that under certain conditions, your findings are not due to chance.

On the other hand, assuming there is a normal distribution for what you found, in every population there is still a 1-5% chance that chance is the reason for your findings.

Now how confusing is that? Very confusing. Do your best to let it sink in.

Thus we test the null, that what is likely to be found in the universe, based upon infinite samples that theoretically fall into normal distributions, to see if we can reject it, the idea that the averages, the means of the variables we're testing are really the same in two groups.

But if they are different, then we can reject that hypothesis, the null hypothesis, which supports our research hypothesis.

Our research hypothesis is that the difference, actually, between the groups will be great because what we did, our intervention, made a big difference. There is only a 5% (or less) likelihood that there would be such a big difference in the parameter we've tested (usually a mean) were it not for the intervention. You've ruled out chance with a significance test.

The significance test compared your results with what should happen in the universe if there were no differences between the two groups, if they came from the same sampling distribution. You've tested the independence of the means.

The normal sampling distribution has one statistic, an average statistic. Your sample is compared to that. If your results vary significantly, then there must be something you did that increased the likelihood that this would happen.

You knew that, which is why you theorized there would be some association between your intervention, your independent variable, and the dependent variable. You had a theory about why this would happen when you compared groups or variables in the first place.