Semester 3 Mathematics Applied Statistics Distributions Hyper Geometric Distribution This distribution is defined to work with experiments where the one after the previous one are dependent on previous ones. Unlike Geometric Distribution , Bernoulli Distribution where Events must be independent of one another, this does not need that property.
In other words, this works when the trials are done without replacement .
We take n n n samples from N N N population without replacement. Then we observe k k k successes in all the samples we took. On the N N N population, there are K K K successes initially.
We write this as:
X ∼ Hypergeometric ( N , K , n ) X \sim \text{Hypergeometric}(N, K, n) X ∼ Hypergeometric ( N , K , n )
Where:
N N N = Total population size
K K K = Total number of success states in the population
N − K N - K N − K = Total number of failure states in the population
n n n = Number of items drawn (sample size)
X X X = Number of observed successes in the sample
P ( X = k ) = ( K k ) ( N − K n − k ) ( N n ) P(X = k) = \frac{\binom{K}{k} \binom{N - K}{n - k}}{\binom{N}{n}} P ( X = k ) = ( n N ) ( k K ) ( n − k N − K )
( K k ) \binom{K}{k} ( k K ) : How many ways to choose k k k successes out of the K K K available.
( N − K n − k ) \binom{N - K}{n - k} ( n − k N − K ) : How many ways to choose the remaining n − k n - k n − k failures out of the N − K N - K N − K available.
( N n ) \binom{N}{n} ( n N ) : Total ways to choose any n n n items out of the total N N N items.
Mean (Expected Value)
E [ X ] = n ⋅ K N E[X] = n \cdot \frac{K}{N} E [ X ] = n ⋅ N K
Variance
Var ( X ) = n ⋅ K N ⋅ ( 1 − K N ) ⋅ ( N − n N − 1 ) \text{Var}(X) = n \cdot \frac{K}{N} \cdot \left(1 - \frac{K}{N}\right) \cdot \left(\frac{N - n}{N - 1}\right) Var ( X ) = n ⋅ N K ⋅ ( 1 − N K ) ⋅ ( N − 1 N − n )
Binomial Distribution is an approximation to Hypergeometric when the population size N N N is huge relative to sample size n n n
If N > > > > n N >>>> n N >>>> n ,
Hypergeometric ( N , K , n ) ≈ Binomial ( n , p = K N ) \text{Hypergeometric}(N, K, n) \approx \text{Binomial}\left(n, p = \frac{K}{N}\right) Hypergeometric ( N , K , n ) ≈ Binomial ( n , p = N K )
p = Number of Successes Total Population = K N p = \frac{\text{Number of Successes}}{\text{Total Population}} = \frac{K}{N} p = Total Population Number of Successes = N K
if we treat, p = K N p=\frac{K}{N} p = N K
Metric Binomial Distribution(With Replacement) Hypergeometric Distribution(Without Replacement) Relationship Mean E [ X ] \mathbb{E}[X] E [ X ] n ⋅ p n \cdot p n ⋅ p n ⋅ p n \cdot p n ⋅ p Identical Variance Var ( X ) \text{Var}(X) Var ( X ) n ⋅ p ( 1 − p ) n \cdot p (1 - p) n ⋅ p ( 1 − p ) n ⋅ p ( 1 − p ) ⋅ ( N − n N − 1 ) n \cdot p (1 - p) \cdot \mathbf{\left(\frac{N - n}{N - 1}\right)} n ⋅ p ( 1 − p ) ⋅ ( N − 1 N − n ) Hypergeometric Variance is Smaller
Var ( X Hypergeometric ) = n p ( 1 − p ) ⏟ Binomial Variance × ( N − n N − 1 ) ⏟ Finite Population Correction (FPC) \text{Var}(X_{\text{Hypergeometric}}) = \underbrace{n p (1 - p)}_{\text{Binomial Variance}} \times \underbrace{\left( \frac{N - n}{N - 1} \right)}_{\text{Finite Population Correction (FPC)}} Var ( X Hypergeometric ) = Binomial Variance n p ( 1 − p ) × Finite Population Correction (FPC) ( N − 1 N − n )
This comes from hypergeometric series.
in a Geometric sequence, the next item is a "Constant " multiplication (r r r ) of previous one.
a , a r , a r 2 , a r 3 , … a, \; ar, \; ar^2, \; ar^3, \; \dots a , a r , a r 2 , a r 3 , …
In a "Hypergeometric " (Beyond geometric) sequence, the next one is multiplied by a "Rational Function "
Term k + 1 Term k = Polynomial ( k ) Polynomial ( k ) \frac{\text{Term}_{k+1}}{\text{Term}_k} = \frac{\text{Polynomial}(k)}{\text{Polynomial}(k)} Term k Term k + 1 = Polynomial ( k ) Polynomial ( k )
When you compute the ratio between consecutive probabilities P ( X = k + 1 ) P(X = k+1) P ( X = k + 1 ) and P ( X = k ) P(X = k) P ( X = k ) in this distribution:
P ( X = k + 1 ) P ( X = k ) = ( K − k ) ( n − k ) ( k + 1 ) ( N − K − n + k + 1 ) \frac{P(X = k+1)}{P(X = k)} = \frac{(K - k)(n - k)}{(k + 1)(N - K - n + k + 1)} P ( X = k ) P ( X = k + 1 ) = ( k + 1 ) ( N − K − n + k + 1 ) ( K − k ) ( n − k )