Descriptive Statistics and Analysis: A Complete Practical Guide

Published On: July 20, 2026

The Complete Guide to Descriptive Statistics: From Basic Summaries to Advanced Normality Testing

Before running advanced machine learning models or inferential hypothesis tests, every data analysis journey begins with a fundamental step: descriptive analytics. Descriptive statistics transforms raw, unorganized data into clear, human-digestible insights.

As John Tukey famously noted in Exploratory Data Analysis (1977), looking at data with a critical, visual eye is essential—summary metrics alone can easily obscure key structural patterns.

1. The Core Pillars of Numerical & Graphical Analysis

Descriptive statistics is generally structured around three fundamental dimensions:

Measures of Central Tendency

Central tendency identifies the typical or center value of a dataset:

  • Mean (): The arithmetic average:​

    arithmetic average

    Best suited for symmetric data without severe outliers.

  • Median (): The exact middle value when data is ordered. Preferred for skewed distributions or datasets with extreme values.

  • Mode: The most frequently occurring observation. Essential for nominal categorical variable reporting.

Measures of Dispersion (Variability)

Dispersion quantifies how widely observations spread out around the center:

    • Standard Deviation () & Variance (): Measure average squared distance from the mean.

Standard Deviation

    • Interquartile Range (): The spread of the middle 50% of values:

In a standard normal distribution, 68.3% of data points fall within , 95.4% within , and 99.7% within (Altman & Bland, 1995).

Graphical Diagnostics: The Box Plot

Visualization is crucial for spotting anomalies that summary statistics hide (Anscombe, 1973). For identifying skewness and outliers, Tukey’s box plot remains a gold standard.

  • Box Boundaries: (25th percentile) to (75th percentile), enclosing the .

  • Center Line: Median ().

  • Whiskers: Extend to beyond the quartiles.

  • Outliers: Individual points plotted beyond the whiskers.

2. Best Practices by Field of Study

Field of Study Recommended Numerical Metrics Preferred Visual Formats Common Pitfalls
Biomedical & Clinical Sciences

Parametric: Mean ()

Non-parametric: Median ()

• Box plots & Violin plots

• Survival curves (Kaplan-Meier)

• Using Mean () for heavily skewed lab values

• Relying on 3D pie charts

Psychology & Social Sciences

• Mean, ,

• Effect sizes (Cohen’s )

• Raincloud plots

• Forest plots

• Truncating Y-axes to inflate trivial differences

• Omitting sample sizes ()

Business & Finance

• Geometric mean

• Percentiles ()

• Value-at-Risk ()

• Time-series line charts

• Waterfalls & Heatmaps

• Relying on mean income/wealth metrics distorted by extreme high earners

3. Higher-Order Moments: Skewness & Kurtosis

When analyzing distributions beyond central tendency and spread, skewness and kurtosis—the 3rd and 4th standardized moments—describe asymmetry and tail behavior.

Skewness: Measuring Asymmetry

Sample skewness () quantifies the degree of asymmetry around the mean:

  • : Symmetric distribution.

  • (Positive / Right-Skewed): Right tail is elongated ().

  • (Negative / Left-Skewed): Left tail is elongated ().

Kurtosis: Measuring Tail Weight & Extremes

Kurtosis measures the propensity for extreme values in the tails relative to a normal distribution (Westfall, 2014). Excess Kurtosis () subtracts 3 so a standard normal curve equals 0:

Kurtosis
  • Mesokurtic (): Normal tail weight.

  • Leptokurtic (): Heavy-tailed (“fat tails”), prone to frequent extreme outliers.

  • Platykurtic (): Light-tailed, fewer extreme outliers than a normal curve.

Real-World Domain Examples

  • Finance: Daily stock market returns often display negative skewness (gradual gains punctuated by severe market crashes) and leptokurtosis (fat tails causing black swan events). Assuming normality leads risk models like Value-at-Risk () to understate catastrophic losses.

  • Biology: Clinical biomarkers (e.g., viral loads or serum ferritin) exhibit positive skewness due to a lower physical bound at zero and hyper-reactive outliers. Conversely, birth weights often show platykurtosis due to stabilizing natural selection.

4. Diagnostic Testing for Normality

To verify whether sample data satisfies normality assumptions, analysts combine visual tools with formal hypothesis testing.

Quantile-Quantile (Q-Q) Plots

A Normal Q-Q plot graphs empirical sample quantiles against theoretical quantiles from a standard normal distribution :

  1. Order sample observations: x(1)​ ≤ x(2)​ ≤ ⋯ ≤ x(n).

  2. Calculate empirical cumulative ranks:

    empirical cumulative ranks

  3. Obtain theoretical standard normal quantiles:

theoretical standard normal quantiles

  1. Plot zi​,x(i) alongside a linear reference line.

Q-Q Plot Diagnostic

Q-Q Plot Diagnostics for Skewness, Kurtosis, and Tail Weight. Source: GitHub Pages

  • Heavy Tails (Leptokurtic): Points form an “S” shape—falling below the line on the left, exceeding it on the right.

  • Right Skew: Points curve upward in an inverted “U” shape.

  • Left Skew: Points curve downward in a “U” shape.

Formal Normality Tests

When visual inspection requires statistical confirmation, two primary goodness-of-fit tests are widely used:

1. The Shapiro-Wilk Test ()

The Shapiro-Wilk test (1965) evaluates the correlation between ordered sample values and normal order statistics:

  • Mechanism: lies between 0 and 1 (). Deviations in skewness or kurtosis reduce the numerator, driving below 1.

  • Best Use: Highly powerful for small to moderate sample sizes ( up to ).

2. The Jarque-Bera Test ()

The Jarque-Bera test (1980) directly evaluates sample skewness () and excess kurtosis ():

Jarque-Bera Test

  • Mechanism: Under (normality), asymptotically follows a Chi-Square distribution with 2 degrees of freedom (). Any non-zero skewness or excess kurtosis inflates .

  • Best Use: Large datasets (), particularly in econometrics and financial modeling.

References

  • Altman, D. G., & Bland, J. M. (1995). Statistics notes: The normal distribution. BMJ, 310(6975), 298.

  • Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27(1), 17–21.

  • Jarque, C. M., & Bera, A. K. (1980). Efficient tests for normality, homoscedasticity and serial independence of regression residuals. Economics Letters, 6(3), 255–259.

  • Shapiro, S. S., & Wilk, M. B. (1965). An analysis of variance test for normality (complete samples). Biometrika, 52(3/4), 591–611.

  • Tukey, J. W. (1977). Exploratory Data Analysis. Addison-Wesley.

  • Westfall, P. H. (2014). Kurtosis as peakedness, 1905–2014. R.I.P. The American Statistician, 68(3), 191–195.

Datawise Firm: Precision in Practice

We empower the future leaders of industry and academia with the analytical tools they need for better solutions. Discover how our 20+ years of expertise in statistical analysis can elevate your next project.

The Data Alchemist

📊 Accomplished Founder & Senior Statistician | Data Science Diplomat | IBM Certified SPSS Profissional | Statistical Training Expert

Leave A Comment