Dulranga's Notes
Semester 3MathematicsApplied Statistics

Hypothesis Testing

Hypothesis testing the mathematical way of saying "We have enough statistical evidence to say this statement is true".

At its core, hypothesis testing is a statistical method used to determine whether there is enough evidence in a sample of data to infer that a certain condition holds true for an entire population.

Step by step process

1. Setting up the hypothesis

Every test require two hypothesis to be defined.

  1. Null Hypothesis (H0H_0) This hypothesis says nothing changed, no effect happened, everything is as it was before.
  2. Alternative Hypothesis (H1H_1 or HAH_A) This is the actual hypothesis you trying to prove. This says the thing has changed, there is a effect, it is different from what it was before.

Example: You did a marketing campaign for a restaurant.

  • H0H_0: The campaign did not increase sales (μcampaign sample=μrestaurant\mu_{\text{campaign sample}} = \mu_{restaurant}).
  • HAH_A: The campaign increased sales (μcampaign sample>μrestaurant\mu_{\text{campaign sample}} > \mu_{restaurant}).
Note

Testing the Alternative hypothesis (HAH_A) is Difficult. But testing the Null Hypothesis (H0H_0) is not difficult comparing to Alternative hypothesis. So instead of testing HAH_A, we test H0H_0 and use P(X)=1−P(X′)P(X) = 1 - P(X') to support that if hypothesis is correct.

2. Declaring Rejection threshold

This threshold is known as the Significance level (α\alpha).

Before looking at the data, we decide how strong the evidence must be to reject H0H_0. This threshold is called the **significance level (α\alpha)

  • Standard practice sets α=0.05\alpha = 0.05 (5%5\%).
  • This means you are willing to accept a 5%5\% risk of wrongly rejecting H0H_0 when it was actually true.

Why set it before hand? This is done for authenticity. Imagine you get a α≈0.049\alpha \approx 0.049 and you say its under 0.050.05 so the hypothesis is valid. But someone can argue that α\alpha is reached by P-Hacking since its so close to 0.050.05 Also difference fields require different threshold levels due to the importance of the hypothesis if it is valid. Ex: A medical research hypothesis related to "A cure for cancer" may require the α\alpha to be <0.01< 0.01

p-value.png

3: Collect Data & Calculate a Test Statistic

You collect sample data and calculate a test statistic (like a zz-score, tt-score, or FF-statistic). This measures how far your sample results fall from what H0H_0 predicted.

4: Find the pp-value

The pp-value is the probability of getting your sample results (or more extreme) if the Null Hypothesis (H0H_0) were true.

  • Small pp-value: The observed data is extremely unlikely under H0H_0.
  • Large pp-value: The observed data easily happens by random chance under H0H_0.

5: Make a Decision

Compare your pp-value to your significance level (α\alpha):

If p≤α⟶Reject H0 (Result is statistically significant)\text{If } p \le \alpha \longrightarrow \textbf{Reject } H_0 \text{ (Result is statistically significant)} If p>α⟶Fail to Reject H0 (Not enough evidence)\text{If } p > \alpha \longrightarrow \textbf{Fail to Reject } H_0 \text{ (Not enough evidence)}
Important

Since we use the null hypothesis is extremely unlikely as the evidence, We cannot ever 100%100\% Prove a hypothesis is correct using statistics. We can simply deduce that given the data, since the Null Hypothesis is extremely unlikely, the Alternative Hypothesis is supported by the evidence. But the actual reason we get the data might be different than the hypothesis.

Example: Imagine you have a hypothesis "Using an study app will help to score higher grades in exam" If your new study app resulted in higher test scores, it could be that:

  1. Your Hypothesis: The app actually helps students learn.
  2. Confounding Variable: The students using the app were just more motivated to begin with.
  3. The Hawthorn Effect: Students performed better simply because they knew they were part of an experiment. There are different ways to get even more evidence to deduce the hypothesis is true.
[ Hypothesis ] ──> [ Controlled Experiment ] ──> [ Replication ] ──> [ Consensus ]

6. Errors in hypothesis Testing

Since we are not 100%100\% certain the Null Hypothesis is wrong/correct, (i.e. Probability of Null hypothesis is not 0%0\% or 100%100\%). There is a α\alpha probability the null hypothesis will be true. In that case, the test fails with error.

hypothesis-errors.png

Type I Error (α\alpha)

Rejected Null Hypothesis, but Null hypothesis is true. (False Positive) This is the probability that our hypothesis is actually wrong.

The probability of Type I Error is α\alpha, which is usually 0.05%0.05\% which is also known as the "Significance level".

Type II Error (β\beta)

Failed to reject Null hypothesis, but Null hypothesis is actually wrong. This is the probability that our hypothesis will be correct even though we deduced its not. There was a difference made but we think it didn't.

The probability of Type II Error is β\beta

The Statistical Power is known as (1−β1- \beta). This is the probability of correctly rejecting the null hypothesis when its actually wrong.

Relationship of α\alpha and β\beta

Assuming the sample size (n) is constantα∝1β\begin{align*} \text{Assuming the sample size (n) is constant} \\ \alpha \propto \frac{1}{\beta} \end{align*}

To reduce both, the best option is to increase the sample size (nn).

Why Alpha and Beta Inverse in Hypothesis errors

If Reality Is...Test MetricMeaning
There IS a real effectPower (80%80\%)You have an 80%80\% chance of correctly catching it.
There IS NO real effectAlpha (α=5%\alpha = 5\%)You have a 5%5\% chance of mistakenly calling a false alarm.

On this page