Statistics - Statistics
Unit 1
6 min read
p{7.2cm}X} Typical Task Goal & Key Triggers & Phrasings & Question Type & Section Link
Classify a statistical term or variable; identify its measurement scale. & ``State the level of measurement of [X]''
``Identify whether the bolded term is a population, sample, attribute, or attribute value''
nominal'' / ordinal'' / ``metric'' &
Terminology & Scales of Measure
Section 2.1
Find class widths/midpoints; calculate histogram bar heights or densities. & ``Calculate the bar height for the class [X]''
relative density'' / absolute density''
class length'' or class midpoint''
3D distortion'' / bar chart quality'' &
Histograms & Graphical Pitfalls
Section 2.2
Compute summary statistics (mean, variance, boxplots) for raw numbers. & ``Calculate arithmetic mean, variance, standard deviation'' (for ungrouped list)
Tukey's hinges'' / boxplot outliers''
``linear transformation ''
``geometric mean'' (growth rates)
``harmonic mean'' (speeds / margins) & Ungrouped Summary Statistics
Section 2.3
Calculate mean/variance from frequency table; find quantile percentiles. & ``Calculate arithmetic mean / variance from a frequency table''
``Below/above which value are X% of the distribution?''
``linear quantile interpolation'' & Grouped Data Analysis
Section 2.4
Quantify linear relationship; find regression line or rank correlation. & ``Quantify the strength of the linear relationship''
``Spearman's rank correlation'' (with ties)
; ``Determine the linear regression line''
``Interpret the coefficient of determination ''
``Predict the value of Y for a given X'' & Bivariate Data (Correlation & Regression)
Section 2.5
\subsection*{Part 2: Combinatorics, Probability & Distributions (Blocks 1 & 2 - Sections 3.1 to 3.4)}
Key concepts
Unit 2
22 min read
\subsubsection*{Diagnostic Guide: How to Identify This Question Type}
Look for questions asking you to:
- State the level of measurement'' or scale of measure'' of a given attribute.
- Classify whether a bolded term refers to a population, attribute, or attribute value.
- Identify a statistically appropriate empirical population for a given survey.
\subsubsection*{Theory: Descriptive vs. Inferential Statistics}
In statistics, we divide our methodologies into two main areas:
- Descriptive Statistics: Focuses on summarizing, organizing, and describing the characteristics of a collected dataset (either a sample or a population). It uses measures of location, variability, tables, and charts to present data clearly without making claims beyond the observed group.
- Inferential (Inductive) Statistics: Uses probability theory to draw conclusions, make predictions, and generalize about a larger population based on the results observed in a representative sample. It includes parameter estimation, confidence intervals, and hypothesis testing.
\subsubsection*{Theory & Slide Definitions}
In descriptive statistics, we define and observe specific characteristics of a group:
- Empirical Population (Target Population): The entire, well-defined set of objects (observational units) that are the subject of a statistical investigation. To be statistically useful, a population must be strictly bounded in three dimensions:
- Factual: Who or what is being observed (e.g., JKU students, light bulbs).
- Spatial: Where the observation is located (e.g., at JKU Linz, manufactured at Factory X).
- Temporal: When the observation is made (e.g., in the Summer Semester 2026, on June 23rd).
- Observational Unit (Case / Subject / Object): The individual entity from the population whose characteristics are being measured (e.g., one specific student, one single battery).
- Attribute (Variable): The specific characteristic or trait of the observational units that is being investigated (e.g., commute distance, eye color, daily screen time).
- Attribute Value (Value / Realization): The actual concrete outcome, measurement, or category recorded for a specific observational unit (e.g., km, ``blue'', hours).
\subsubsection*{Theory: Levels of Measurement (Scales)}
Attributes are classified by their mathematical properties:
- Nominal Scale: Qualitative categories with no natural order or ranking (e.g., car color, gender, ticket type, blood type). Only equality ( or ) can be compared.
- Ordinal Scale: Qualitative categories with a natural, objective order/rank, but the distance between values is not measurable or equal (e.g., grades: 1--5, satisfaction ratings: 1 = high to 5 = low, language ability: beginner/intermediate/fluent, pain level).
- Metric (Interval/Ratio) Scale: Quantitative values with measurable, objective distances between them.
- Metric Discrete: Countable values, typically integers (e.g., number of children, number of website visits, number of trips).
- Metric Continuous: Uncountable values that can theoretically be measured with infinite precision (e.g., commute time, weight, height, volume, reaction time, income).
\subsubsection*{Example (Inspired by Assignment 1)}
State the level of measurement for:
- Commute time (in minutes): Metric continuous (time can be divided infinitely).
- Satisfaction (1 = high, 5 = low): Ordinal (has order, but distances between points are subjective).
- Number of public transit trips: Metric discrete (countable integer trips).
- Primary ticket type (single, day pass): Nominal (categories with no ranking).
\subsection{Subtopic 1.2: Charts, Histograms & Graphical Pitfalls}
Key concepts
Unit 3
29 min read
\subsubsection*{Diagnostic Guide: How to Identify This Question Type}
Look for questions asking you to:
- Find the number of ways to arrange a set of objects (arrangements'', permutations'', straight line'', in a row'').
- Find the number of ways to select a subset of objects (choose a committee'', lottery tips'', draw packs''). - Deal with constraints like objects of the same group must stand together'' or ``committee must be either all from group A or all from group B''.
\subsubsection*{Theory: The Six Fundamental Counting Rules}
Let be the total number of available objects, and be the number of objects selected.
Counting Rule & Formula & Key Characteristic / Slide Example
Permutation & & Arranging distinct objects.
(Without Repetition) & & Slide Example: Arranging 8 books in a closet.
\addlinespace Permutation & & Arranging objects with identical groups of size .
(With Repetition) & & Slide Example: Arranging the letters in AAA BB ().
\addlinespace Combination & & Unordered selection of from without replacement.
(Without Repetition) & & Slide Example: Lotto 6 out of 45.
\addlinespace Combination & & Unordered selection of from with replacement.
(With Repetition) & & Slide Example: Choosing 3 cigarette packs of 15 brands.
\addlinespace Variation & & Ordered selection of from without replacement.
(Without Repetition) & & Slide Example: Forming numbers from digits without reuse.
\addlinespace Variation & & Ordered selection of from with replacement.
(With Repetition) & & Slide Example: Bicycle lock with 5 rings of digits 0--9.
Ask yourself two questions to identify the correct counting rule: - Does the order matter? - Yes: It is a Variation (or a Permutation if using all elements). E.g., passcodes, lottery drawings in sequence, forming numbers. - No: It is a Combination. E.g., committees, hands of cards, lottery jackpots where order of draw is ignored. - Is replacement allowed (can elements be reused)? - Yes: Use the formulas with Repetition ( or ). - No: Use the standard formulas ( or ). \subsubsection*{Exam-Specific Patterns & Solution Recipes} \paragraph{Pattern A: Straight-line arrangements where a specific subgroup must stand together}
Key concepts
Unit 4
22 min read
\subsubsection*{Diagnostic Guide: How to Identify This Question Type} Look for questions asking you to: - Identify appropriate point estimators for the population mean () or population variance () given sample data. - Evaluate whether an estimator is unbiased (conceptual T/F questions, or showing that ). - Compare two estimators to determine which is more efficient (comparing their variances). - Assess whether an estimator is consistent (checking if it converges to the true parameter as ). - Distinguish between a point estimate (a single numerical value with high precision but 0 \subsubsection*{Theory & JKU-Specific Slide Definitions} In inferential (inductive) statistics, we use sample data to draw conclusions about unknown population parameters (such as the true mean or true variance ): - Point Estimator (\hat{\theta**):} A random variable (formula) used to estimate a population parameter from sample data. - Point Estimate: The concrete numerical value obtained by evaluating the estimator formula on a specific sample dataset (e.g., g). - Standard Estimators: - For the Population Mean (): The sample mean () is the standard point estimator: - For the Population Variance (): The sample variance () is the standard point estimator: Note on the Denominator: We divide by (Bessel's correction) rather than to make the estimator unbiased. Dividing by would systematically underestimate the true population variance. \subsubsection{Theory: The Three Fundamental Parameter Properties} To assess the quality of an estimator , JKU slides define three core properties (frequently tested in True/False exam questions): - Unbiasedness (German: Erwartungstreue): An estimator is unbiased if its expected value is exactly equal to the true population parameter : If the bias is zero, the estimator does not systematically overestimate or underestimate the parameter. Key JKU Fact: The sample mean is an unbiased estimator of , and the sample variance (with denominator) is an unbiased estimator of . - Efficiency (German: Effizienz): If we have two different unbiased estimators and for the same parameter , the one with the smaller variance is more efficient: An efficient estimator fluctuates less from sample to sample, providing higher precision. - Consistency (German: Konsistenz): An estimator is consistent if, as the sample size approaches infinity, the estimator converges in probability to the true parameter : Practically, this means that as the sample size grows, the estimation error shrinks to zero. Consistency is guaranteed by the Law of Large Numbers (LLN). \subsubsection{Worked JKU Example (Point Estimation & Bias Proof)} Fritz wants to estimate the average bottled quantity of beer. He takes a random sample of bottles, yielding the measurements (in ml): , , , , . 1. Calculate the point estimate for the population mean . 2. Calculate the point estimate for the population variance . 3. Mathematically prove that the sample mean estimator is an unbiased estimator of , assuming the observations are independent and identically distributed (i.i.d.) with . \paragraph{Step 1: Calculate the Point Estimate for } The point estimate is the sample mean : \paragraph{Step 2: Calculate the Point Estimate for } The point estimate is the sample variance , using the unbiased denominator: $ s^2 &= \frac{1}{4} \left[ (502-500)^2 + (498-500)^2 + (503-500)^2 + (497-500)^2 + (500-500)^2 \right]
&= \frac{1}{4} \left[ 2^2 + (-2)^2 + 3^2 + (-3)^2 + 0^2 \right]
&= \frac{1}{4} \left[ 4 + 4 + 9 + 9 + 0 \right] = \frac{26}{4} = 6.500 \text{ ml}^2
s = \sqrt{6.5} \approx 2.550E[\bar{X}] = \mu E[\bar{X}] &= E\left[ \frac{1}{n} \sum_{i=1}^n X_i \right]
&= \frac{1}{n} E\left[ \sum_{i=1}^n X_i \right] \quad \text{(by linearity of expectation)}
&= \frac{1}{n} \sum_{i=1}^n E[X_i]
&= \frac{1}{n} \sum_{i=1}^n \mu \quad \text{(since } E[X_i] = \mu \text{ for all } i)
&= \frac{1}{n} \cdot (n \cdot \mu) = \mu
E[\bar{X}] = \mu\bar{X}\mu. \subsection{Subtopic 3.2: Confidence Intervals for the Mean (\mu$) [EXAM HEAVY]}
Key concepts
Unit 5
31 min read
\subsubsection*{Diagnostic Guide: How to Identify This Question Type}
Look for questions asking you to:
- Formulate null () and alternative () hypotheses for a given research question.
- Identify or define Type I error () and Type II error () in a concrete scenario.
- Define or calculate the Power () of a test.
- Explain the relationship between and , or interpret the meaning of a p-value (concrete Type I error of a data situation).
- Determine whether to reject based on a test statistic and critical value, or based on a p-value.
- Answer conceptual True/False questions about the limits of significance tests (e.g., why we cannot prove'' $H_0$, or the difference between significance and relevance). \subsubsection*{Theory: The JKU Hypothesis Framework} Every hypothesis test starts with a research question. We translate this question into two competing, mutually exclusive hypotheses about a population parameter: - **Null Hypothesis ($H_0$):** The status quo, representing no difference,'' no effect,'' or equality.'' We assume is true unless the data provides strong evidence against it.
- Alternative Hypothesis ( or ): The research hypothesis, representing the difference,'' effect,'' or ``change'' that the researcher wants to verify or prove.
Critical JKU Rule: The hypothesis that we want to prove is always formulated as the alternative hypothesis . Because we control the Type I error (the error of falsely rejecting ), rejecting gives us strong, controlled evidence in favor of .
\subsubsection*{Theory: The Four Possible Test Outcomes and Errors}
When making a decision based on a sample, we compare the statistical decision to the unknown reality in the population. This leads to four possible situations:
\renewcommand{\arraystretch}{1.3}
Decision (Sample) & Reality: is True & Reality: is True
Do Not Reject & Correct Decision (True Negative) & Type II Error () (False Negative)
& Probability: & Probability:
Reject (Accept ) & Type I Error () (False Positive) & Correct Decision (True Positive)
& Probability: (Significance Level) & Probability: (Power)
JKU Error Matrix for Hypothesis Testing
- Type I Error (, German: -Fehler): We reject even though is actually true. This is a false positive (e.g., stating a new feed causes different weight gain when it actually does not). We directly control this probability by setting the significance level (usually or ) before running the test.
- Type II Error (, German: -Fehler): We fail to reject even though is actually true. This is a false negative (e.g., failing to detect that a feed causes different weight gain). This probability is not directly controlled by a standard significance test.
- Power (): The probability of correctly rejecting when is true. This represents the probability of detecting a real effect when one exists.
JKU exam questions frequently test the following properties of and :
- Inverse Relationship: and are inversely related. If you make the test more conservative by lowering the significance level (e.g., from to ), you make it harder to reject . This naturally increases the probability of a Type II error () and reduces the power ().
- Do Not Add up to 100%: Although they are inversely related, and do not add up to (). They are conditional probabilities defined under different assumptions (Type I error is conditional on being true; Type II error is conditional on being true).
- Control of Type II Error: Standard significance tests do not control or calculate directly. If we fail to reject , we cannot say we accept $H_0$.'' We can only say we fail to reject '' or ``there is no significant evidence against ,'' because the Type II error could be very large (up to ).
- Sample Size Effect: The only way to decrease both and simultaneously is to increase the sample size . A larger sample provides more information, which shrinks the standard error and allows for a more powerful test with the same significance level.
\subsubsection*{Theory: The p-Value --- The Concrete Type I Error}
Thomas Forstner's slides define the p-value as:
The p-value is the concrete Type I error of a data situation.
Specifically, it is the probability under the null hypothesis of obtaining a test statistic at least as extreme as the one observed in the sample.
- Decision Rule using p-value:
p < \alpha &\implies \text{**Reject**H_0$ (Result is statistically significant)}
p \ge \alpha &\implies \text{**Do Not Reject** $H_0$ (Result is not statistically significant)}
$
- **Relation to Two-Sided vs. One-Sided Tests:**
For a symmetric distribution, the p-value of a two-sided test is exactly **twice** the p-value of a one-sided test:
$
p_{\text{two-sided}} = 2 \cdot p_{\text{one-sided}}
$
Because of this, choosing a one-sided test just to get a significant result is cheating (``fishing for significance''). To prevent this, JKU slides note that if a one-sided test is used, its significance level should be halved: $\alpha_{\text{one-sided}} = \alpha_{\text{two-sided}} / 2$ (e.g., using $\alpha = 2.5%$ for a one-sided test instead of $5%$).
\subsubsection*{Theory: Swapping the Hypotheses} Why can't we just set to prove a quantity is exactly Euro? - Continuous Precision: A continuous variable can always be measured more precisely (e.g., instead of ). Proving a point hypothesis makes no mathematical sense. - Confidence Interval Connection: In a significance test, is rejected if and only if the confidence interval does not cover the hypothesized value. If we swapped them to make and , we could only reject if the confidence interval collapsed to a single point containing only . A single-point confidence interval has a confidence level of (and thus a Type I error of ). - Equivalence Testing: To prove that a parameter is practically equal to a value, we must define an equivalence region (e.g., ) a-priori and test whether the parameter lies within this region. \subsubsection*{Theory: Significance vs. Relevance} - Significance (p-value): Tells us if a difference exists that is unlikely to have occurred by chance. It depends heavily on the sample size . If is extremely large, even a tiny, practically meaningless difference (e.g., a blood pressure reduction of mmHg) will be highly statistically significant (). - Relevance (Confidence Interval): Tells us if the size of the difference is practically meaningful in the real world. We measure relevance by comparing the confidence interval of the difference to a pre-defined relevance limit (). JKU Exam Tip: Always interpret statistical significance in light of real-life relevance. Use the confidence interval to see if the entire interval lies above the relevance limit (significant and relevant), or if it lies below it (significant but not relevant). \subsubsection*{Worked JKU Examples (Step-by-Step)} \paragraph{Example 1: Hypothesis Formulation & Identifying Errors (Dr. X's Rat Protein Study)} Dr. X postulates that more protein in rat feed causes a different weight gain in rats than less protein. Formulate the hypotheses and define the errors in this context. - Formulate Hypotheses: The research statement is that more protein causes a different weight gain (two-sided alternative): H_0 &: \text{Weight gain with more protein is equal to weight gain with less protein (\mu_{\text{more}} = \mu_{\text{less}}$)}
H_1 &: \text{Weight gain with more protein is different from weight gain with less protein ($\mu_{\text{more}} \neq \mu_{\text{less}}$)}
$
- **Type I Error ($\alpha$):** We conclude that more protein causes a different weight gain when, in reality (in the population), there is no difference ($\mu_{\text{more}} = \mu_{\text{less}}$). This is a **false positive**.
- **Type II Error ($\beta$):** We fail to reject the null hypothesis, concluding that more protein does not cause a different weight gain when, in reality, there is a difference ($\mu_{\text{more}} \neq \mu_{\text{less}}$). This is a **false negative**.
- **Power ($1-\beta$):** The probability that we correctly reject $H_0$ and conclude that there is a difference in weight gain, given that a real difference in the population actually exists.
\paragraph{Example 2: Two-Sided vs. One-Sided p-Values and Decisions} A researcher conducts a test with a chosen significance level of . The statistical software reports a two-sided p-value of . 1. Make the test decision for the two-sided test. 2. Calculate the corresponding one-sided p-value, assuming the sample difference is in the hypothesized direction. 3. If the researcher decided to perform a one-sided test, what local significance level should they choose to protect against inflation, and what would the test decision be? - Two-Sided Test Decision: Since , the null hypothesis is rejected. The difference is statistically significant at the level. - One-Sided p-Value Calculation: For a symmetric distribution, the one-sided p-value is exactly half of the two-sided p-value: - Significance Level for One-Sided Test: To prevent ``fishing for significance,'' the significance level of the one-sided test should be halved: Since , the null hypothesis is still rejected. \paragraph{Example 3: Multiple Testing & Bonferroni Correction} A genomic study tests independent gene sequences simultaneously. Each test is conducted at a local significance level of . 1. Calculate the global significance level (probability of at least one false positive). 2. Apply the Bonferroni correction to find the adjusted local significance level required to guarantee a global significance level of at most . - Calculate Global Significance Level: Testing 4 hypotheses simultaneously increases the false positive risk from to . - Apply Bonferroni Correction: To ensure the global error probability is at most , we adjust the local significance level for each test: We reject a local null hypothesis only if its local p-value is less than . \paragraph{Example 4: Significance vs. Relevance (Systolic Blood Pressure)} A study defines a systolic blood pressure difference of more than mmHg as clinically relevant. For each of the following confidence intervals for the mean difference (), classify the result (using the significance limit and relevance limit ): 1. Interval mmHg 2. Interval mmHg 3. Interval mmHg 4. Interval mmHg - Analyze : The interval covers , so the difference is not statistically significant (). However, the upper limit exceeds mmHg, indicating a trend to relevance. - Analyze : The interval does not cover , so the difference is statistically significant (). However, the entire interval is below mmHg, so the difference is not clinically relevant. - Analyze : The interval does not cover , so it is statistically significant (). The entire interval lies strictly above mmHg, so the difference is clinically relevant as well. - Analyze : The interval covers and lies entirely below mmHg. The result is neither statistically significant nor relevant. \subsection{Subtopic 4.2: One-Sample Tests for the Mean () [EXAM HEAVY]}
Key concepts
Unit 6
11 min read
p{5.7cm}X} Measure / Concept (Subtopic) & Formula & Variable Legend
Class Mark (Subtopic 1.2) & & : class boundaries
\addlinespace Histogram Height (Subtopic 1.2) & & : abs. freq., : rel. freq., : class width
\addlinespace Arithmetic Mean (Subtopic 1.3) & & : sample size
\addlinespace Weighted Mean (Subtopic 1.3) & & : weights (quantities, sample sizes)
\addlinespace Geometric Mean (Subtopic 1.3) & & Used for average growth rates ()
\addlinespace Harmonic Mean (Subtopic 1.3) & & Used for average speeds over equal distances
\addlinespace Sample Trans. (Subtopic 1.3) & & Transforming sample mean and variance ()
\addlinespace Sample Variance (Subtopic 1.3) & & : sample variance (unbiased, JKU exam default)
\addlinespace Population Var. (Subtopic 1.3) & & : population variance (empirical)
\addlinespace Tukey's Hinges (Subtopic 1.3) & & : median index, : hinge index
\addlinespace Grouped Mean (Subtopic 1.4) & & : abs. freq., : rel. freq., : class mark
\addlinespace Grouped Sample Var. (Subtopic 1.4) & & : total sample size
\addlinespace Grouped Pop. Var. (Subtopic 1.4) & & : population grouped variance
\addlinespace Grouped Quantile (Subtopic 1.4) & & : lower bound of class , : width, : abs. freq., : target quantile, : cumulative abs. freq. before class
\addlinespace Quantile Phrasing (Subtopic 1.4) & Below : \quad Above : & Translating text into target
\addlinespace Sample Covariance (Subtopic 1.5) & & : sample covariance (unbiased, exam default)
\addlinespace Pearson Correlation (Subtopic 1.5) & & : sample standard deviations
\addlinespace Spearman Correlation (Subtopic 1.5) & & Bravais-Pearson correlation on ranks (Ties: assign average rank)
\addlinespace Regression Line (Subtopic 1.5) & & (slope), (intercept)
\addlinespace Coeff. of Det. () (Subtopic 1.5) & & (Total), (Explained)
\subsection*{Block 2, Part A: Combinatorics & Basic Probability/Bayes' (Subtopics 2.1--2.2)}
Key concepts