Graphing • Multivariate Statistics & Distribution Analysis

Scatterplot with Marginal Histograms Calculator

Plot bivariate coordinate observations alongside aligned top and right univariate marginal frequency histograms. Analyze joint correlation, Gaussian Kernel Density Estimation (KDE) curves, Ordinary Least Squares trendlines, and distribution modality across both dimensional projections simultaneously.

|
Last Updated: September 2026
|
Verified: OLS • Pearson r • Gaussian KDE
Bivariate Distribution Archetypes Click to load sample dataset
Observation Series Data (X & Y)
35 points
35 points
Histogram & Overlay Settings
Marginal Histogram Bins: 8 Bins
4 (Coarse) Auto (Sturges) 25 (Fine)
Bivariate Correlation & Moments
Pearson r +0.992
R² Variance 98.4%
Covariance 495.2
Sample (n) 35
OLS Regression Model: Strong Positive
ŷ = 2.148x + 2.083
Series Mean (μ) Std Dev (σ) Min → Max
X (Top) 37.14 15.42 12.0 → 65.0
Y (Right) 81.86 33.40 28.0 → 142.0
Joint Scatter & Marginal Projections Retina 2D Canvas
x: 0.00, y: 0.00
Scatter Observations (X, Y)
OLS Linear Trendline
Marginal Histograms
Drag to pan • Scroll wheel to zoom
•
Direct Answer & Overview
Verified Educational Guide

Multivariate Probability

Foundations of Joint Distributions & Marginal Histograms

In exploratory data analysis and mathematical statistics, analyzing a single variable in isolation (univariate analysis) frequently conceals critical real-world phenomena, while inspecting a two-dimensional scatterplot without marginal context can mislead researchers regarding the underlying distributions of individual dimensions. The scatterplot with marginal histograms resolves this dilemma by unifying bivariate coordinate projections and univariate density distributions into a cohesive, synchronized visual framework.

Mathematically, let (X, Y) be a continuous bivariate random vector governed by the joint probability density function f_{XY}(x, y). The central scatterplot visualizes empirical draws sampled from this joint density. The probability that an observation falls within an infinitesimal Cartesian region dx × dy is:

P(x ≤ X ≤ x + dx, y ≤ Y ≤ y + dy) = f_{XY}(x, y) dx dy

The marginal probability density function of X, denoted f_X(x), represents the total probability distribution of X integrated over all plausible values of Y. Conversely, the marginal density of Y, denoted f_Y(y), integrates over all possible values of X:

f_X(x) = ∫_-∞^{+∞} f_{XY}(x, y) dy
f_Y(y) = ∫_-∞^{+∞} f_{XY}(x, y) dx

By aligning the horizontal bins of the top marginal histogram precisely with the X-axis coordinate domain of the central scatter plane, and rotating the vertical bins of the right marginal histogram by 90° to match the Y-axis range, this visualization preserves spatial geometry. Every point in the scatterplot vertically projects its mass into the top histogram bin and horizontally projects its mass into the right histogram bin.

Analytical Mechanics

Mathematical Architecture: Pearson r, Covariance & Regression

The interactive joint plot computes essential bivariate moments and ordinary least squares linear regression parameters across the n coordinate pairs {(x_1, y_1), (x_2, y_2), …, (x_n, y_n)}.

1. Univariate Centroids & Standard Deviations

The sample means x̄ and ȳ locate the balance centroid of the joint distribution, displayed as the dashed crosshair intersection:

x̄ = (1/n) ∑_{i=1}^n x_i,   s_x = √[ (1/(n-1)) ∑_{i=1}^n (x_i - x̄)² ]

2. Sample Covariance & Pearson Product-Moment Correlation (r)

The degree of linear co-variation between X and Y is quantified by sample covariance Cov(X, Y) and normalized Pearson correlation r:

Cov(X, Y) = s_{xy} = (1/(n-1)) ∑_{i=1}^n (x_i - x̄)(y_i - ȳ)
r = s_{xy} / (s_x · s_y) = ∑ (x_i - x̄)(y_i - ȳ) / √[ ∑ (x_i - x̄)² · ∑ (y_i - ȳ)² ]

Where -1 ≤ r ≤ +1. The coefficient of determination R² = r² represents the proportion of variance in Y predictable from X under a linear specification.

3. Ordinary Least Squares (OLS) Trendline

The fitted trendline ŷ = b_1 x + b_0 minimizes the sum of squared vertical residuals ∑ (y_i - ŷ_i)²:

b_1 = s_{xy} / s_x² = r · (s_y / s_x),   b_0 = ȳ - b_1 · x̄
Nonparametric Smoothing

Kernel Density Estimation (KDE) & Nonparametric Smoothing

While histograms provide an intuitive summary of counts partitioned into discrete intervals, their visual shape depends heavily on arbitrary bin widths and bin boundary choices. A minor shift in origin can dramatically alter whether a histogram appears unimodal or bimodal. Kernel Density Estimation (KDE) circumvents this discretization bias by placing a continuous, smooth weighting function (the kernel) over each observed data point and aggregating them into a smooth density curve.

f̂_h(x) = (1 / (n · h)) ∑_{i=1}^n K((x - x_i) / h)

In our calculator engine, we implement the standard Gaussian Kernel, satisfying the regularity conditions ∫ K(u) du = 1 and ∫ u K(u) du = 0:

K(u) = (1 / √(2π)) exp(- ½ u²)

Bandwidth Selection: Silverman's Rule of Thumb

The smoothing parameter h (bandwidth) controls the bias-variance tradeoff of the estimated density. An excessively small h leads to undersmoothing (spurious peaks from individual points), whereas an excessively large h oversmooths, flattening true multimodal features. Our engine calculates the optimal bandwidth using Silverman's robust formula:

h_{opt} = 1.06 · min(s_x, IQR_x / 1.34) · n-1/5

Using the minimum between the sample standard deviation s_x and the scaled Interquartile Range (IQR / 1.34) ensures stability against extreme outliers and skewed tails.

Data Discretization

Optimal Binning Rules: Freedman-Diaconis, Scott & Sturges

The number of histogram bins k and bin width w dictate the resolution of the marginal histograms. Statistical literature defines three primary formulations for determining binning:

Freedman-Diaconis Rule

Optimized to minimize the L² integrated mean squared error without assuming Gaussian normality. Highly robust against outliers:

w = 2 · IQR · n-1/3
k = ⌈ (max - min) / w ⌉

Scott's Normal Rule

Asymptotically optimal for unimodal Gaussian random variables by minimizing mean integrated squared error (MISE):

w = 3.49 · s · n-1/3
k = ⌈ (max - min) / w ⌉

Sturges' Rule

Based on the binomial expansion (1 + 1)^n. Tends to over-smooth for n > 200 or non-normal data:

k = ⌈ log_2(n) + 1 ⌉
w = (max - min) / k

Our calculator defaults to k = 10 bins while providing an interactive slider (4 ≤ k ≤ 25) so you can dynamically inspect how changing bin resolution impacts the perception of modal peaks and edge tails.

Diagnostic Heuristics

Diagnosing Bivariate Phenomena: Modality, Skewness & Variance

When examining joint distributions, marginal histograms expose essential statistical anomalies that a standard 2D scatterplot might obscure:

1. Multimodality and Mixture Distributions

A scatterplot displaying points scattered across a wide region might mask distinct clusters if their bounding ellipses overlap. If the top marginal histogram shows two separated peaks (bimodal distribution), it confirms that X consists of a mixture of two distinct sub-populations with different latent means.

2. Asymmetric Skewness & Long-Tailed Residuals

Standard linear regression assumes that vertical residuals are normally distributed with zero mean. If the right marginal histogram of Y exhibits severe positive skewness (a long tail stretching toward higher values), linear regression coefficients will be disproportionately pulled by high-leverage outliers, violating ordinary least squares assumptions.

3. Heteroscedasticity (Non-Constant Variance)

In a funnel-shaped scatterplot where point dispersion expands as X increases, conditional variance Var(Y|X=x) is non-constant. Inspecting the top marginal histogram clarifies whether observations are uniformly dense across X or concentrated primarily in the narrow-variance region.

Computational Workflow

Step-by-Step Numerical Algorithm for Constructing Marginal Histograms

The algorithmic pipeline executed by our 2D canvas rendering engine proceeds through seven deterministic phases:

1

Coordinate Domain Parsing & Padding

Extract n coordinate pairs. Compute bounds [x_{min}, x_{max}] and [y_{min}, y_{max}]. Apply a 5% safety margin padding on each boundary to ensure peripheral points and histogram edges do not touch canvas view boundaries.

2

Linear Affine Transform Mapping

Map world coordinates (x, y) to pixel screen coordinates (u, v) within the central bounding box:
u = u_{min} + ((x - x_{min}) / Δx) · W_{scatter},   v = v_{max} - ((y - y_{min}) / Δy) · H_{scatter}

3

Univariate Partitioning into k Uniform Bins

Partition domain intervals into k equal intervals of width w_x = Δx / k and w_y = Δy / k. Increment frequency counts c_j for each coordinate falling in [B_j, B_{j+1}).

4

Top Marginal Histogram Bar Rendering

Calculate bar heights proportional to c_j / c_{max,x}. Render vertical rectangular bars anchored to the top baseline, perfectly spanning horizontal intervals [u_j, u_{j+1}].

5

Right Marginal Histogram Bar Rendering (90° Rotation)

Render horizontal rectangular bars anchored to the right boundary of the central scatter plane. Bar lengths extend to the right proportional to c_m / c_{max,y}, matching vertical screen coordinates [v_{m+1}, v_m].

6

Continuous Kernel Density Curve Evaluation

Sample 100 evaluation steps along each margin. Evaluate Gaussian density f̂(t) = (1/(n·h)) ∑ K((t - t_i)/h). Scale maximum curve peak to 92% of margin height and stroke with emerald and amber path geometries.

7

Dynamic Canvas Resizing & High-DPI Rendering

Scale canvas buffer by device pixel ratio window.devicePixelRatio to ensure razor-sharp graphics on Retina and mobile screens.

Typology & Diagnostics

Comprehensive Bivariate Distribution Archetypes Comparison Table

The table below contrasts standard joint distribution archetypes against their diagnostic signatures across the scatter plane, marginal histograms, and correlation metrics:

Archetype Scatter Morphology Top Marginal (X) Right Marginal (Y) Pearson r Primary Risk / Fallacy
Positive Linear Correlation Narrow diagonal elliptical swath sloping upward Unimodal, approximately normal Unimodal, approximately normal 0.80 to 0.95 Conflating correlation with direct causation
Bivariate Gaussian (Spherical) Circular cloud with radial density decay Symmetric Gaussian bell curve Symmetric Gaussian bell curve ≈ 0.00 Assuming r = 0 implies total statistical independence
Bimodal Mixture (Two Clusters) Two distinct, disconnected point clouds Twin peaks (bimodal distribution) Twin peaks (bimodal distribution) Spurious positive or negative Simpson's Paradox; ignoring latent sub-groups
Right-Skewed Exponential Tail Dense cluster at bottom-left with scattered outliers Severe right skew (long right tail) Severe right skew (long top tail) Artificially inflated by outliers OLS violation; failure to apply logarithmic transform
Non-Linear Quadratic (U-Curve) Symmetric parabolic curve y ≈ x² Uniform or symmetric normal Strongly asymmetric; mass at minimum ≈ 0.00 Falsely concluding no relationship exists
Practice & Application

Graded Worked Statistical Problems with Full Solutions

Problem 1: Manual Calculation of Joint Moments & Marginal Bins Foundational

Consider the sample of n = 5 observations: {(2, 4), (4, 7), (5, 8), (7, 11), (10, 15)}.
Tasks: (a) Calculate sample means x̄ and ȳ. (b) Compute sample covariance s_{xy} and Pearson correlation r. (c) Construct k = 2 uniform bins for the marginal distribution of X and state bin counts.

Complete Step-by-Step Solution:

Step 1: Compute Centroids.
x̄ = (2 + 4 + 5 + 7 + 10) / 5 = 28 / 5 = 5.6
ȳ = (4 + 7 + 8 + 11 + 15) / 5 = 45 / 5 = 9.0

Step 2: Deviations and Cross-Products.
i=1: (2 - 5.6)(4 - 9) = (-3.6)(-5) = +18.0
i=2: (4 - 5.6)(7 - 9) = (-1.6)(-2) = +3.2
i=3: (5 - 5.6)(8 - 9) = (-0.6)(-1) = +0.6
i=4: (7 - 5.6)(11 - 9) = (+1.4)(+2) = +2.8
i=5: (10 - 5.6)(15 - 9) = (+4.4)(+6) = +26.4
Sum of cross-products: SS_{xy} = 18.0 + 3.2 + 0.6 + 2.8 + 26.4 = 51.0.
Sample covariance: s_{xy} = 51.0 / (5 - 1) = 51.0 / 4 = 12.75.

Step 3: Sum of Squared Deviations & Correlation.
SS_x = (-3.6)² + (-1.6)² + (-0.6)² + 1.4² + 4.4² = 12.96 + 2.56 + 0.36 + 1.96 + 19.36 = 37.2
SS_y = (-5)² + (-2)² + (-1)² + 2² + 6² = 25 + 4 + 1 + 4 + 36 = 70.0
r = 51.0 / √(37.2 · 70.0) = 51.0 / √(2604) = 51.0 / 51.0294 ≈ +0.9994.

Step 4: Marginal Bins for X.
Domain of X: [x_{min}, x_{max}] = [2, 10]. Range Δx = 8. For k = 2 bins, bin width is w = 8 / 2 = 4.
Bin 1: [2, 6) → values {2, 4, 5} → Count = 3 (60%).
Bin 2: [6, 10] → values {7, 10} → Count = 2 (40%).

Problem 2: Dissecting Simpson's Paradox in Marginal Projections Advanced Diagnostic

An observational study measures hours studied (X) versus exam error score (Y). The aggregate dataset shows r = +0.72 (suggesting more studying yields more errors). However, the top marginal histogram of X exhibits two isolated peaks at x = 4 and x = 14.
Question: How does the marginal histogram reveal the underlying confounder, and what occurs when the clusters are conditioned separately?

Diagnostic Explanation:

The prominent bimodal signature in the top marginal histogram alerts the analyst that the sample is not a homogeneous univariate draw, but an aggregated mixture of two distinct cohorts: Group A (Introductory level students studying 2–6 hours, baseline errors 2–5) and Group B (Advanced Quantum Mechanics students studying 12–16 hours, baseline errors 15–20).

Within Group A, the conditional correlation is r_{A} = -0.85. Within Group B, r_{B} = -0.82. In both separate groups, studying reduces errors. However, because the advanced group has both higher study demands and inherently higher test complexity, aggregating the groups shifts the centroid from (4, 3.5) to (14, 17.5), producing a spurious aggregate positive slope r_{agg} = +0.72. The marginal histogram's twin peaks provide the visual proof required to identify Simpson's Paradox.

Empirical Science

Cross-Disciplinary Real-World Applications

Genomics & Bioinformatics

In single-cell RNA sequencing (scRNA-seq), joint plots graph expression levels of two candidate genes across tens of thousands of individual cells. Marginal histograms reveal zero-inflation and bimodal expression states (active transcription vs. transcriptional silencing).

Quantitative Finance & Risk

Risk managers plot asset returns against benchmark market index returns. Right and top marginal histograms expose heavy negative tails (leptokurtic excess kurtosis), proving that asset returns deviate from Gaussian assumptions and possess severe tail-risk downside.

Meteorology & Climate

Climate scientists plot atmospheric temperature against relative humidity. Joint plots identify dew point saturation boundaries where the joint distribution truncates abruptly, while marginal histograms reflect seasonal temperature oscillations.

Analytical Caveats

Common Pitfalls, Artifacts & Simpson's Paradox to Avoid

1. Bin Edge Placement Artifacts

Selecting too few bins (k < 5) averages away subtle multimodal peaks, while selecting too many bins (k > 30) produces jagged, noisy combs where individual bins contain single observations. Always toggle the Gaussian KDE curve to confirm whether a histogram peak is an authentic distribution mode or an artifact of bin boundary placement.

2. Confusing Marginal Independence with Joint Independence

Two random variables are independent if and only if f_{XY}(x, y) = f_X(x) · f_Y(y) for all coordinates. Observing that marginal histograms are symmetric and bell-shaped does not guarantee that the joint distribution is bivariate normal. Highly non-linear dependencies (such as circular manifolds or parabolic relationships) can possess perfectly normal marginals while being entirely dependent.

3. Over-Reliance on Pearson r in Non-Linear Regimes

Pearson's correlation coefficient r measures strictly linear association. For data following a symmetric U-shaped parabola y = x² over [-10, 10], Pearson r ≈ 0.00 despite a deterministic mathematical relationship. Always visually inspect the central scatterplot rather than relying solely on summary statistics.

Interactive Ecosystem

Connected Graphing & Statistics Ecosystem Hub

Deepen your exploratory statistical and regression analysis with complementary mathematical graphing instruments across our suite:

Fact-Checked & Verified • Computational Accuracy Standards
Updated September 2026 • Editorial Policy
Authored By
Sanjay Samanta

Lead Developer & Founder of Basic Math Tools. Specializes in browser-native computational algorithms and applied mathematics.

Reviewed & Verified By
Academic Review Board

Mathematics & curriculum specialists. Audited against standard algebraic and arithmetic principles.

Found an error or have an improvement suggestion? Report a calculation issue

Frequently Asked Questions

What is a scatterplot with marginal histograms (joint distribution plot)?
A scatterplot with marginal histograms is a composite statistical visualization combining a central two-dimensional scatterplot with aligned univariate frequency histograms positioned along the top (horizontal X-axis distribution) and right (vertical Y-axis distribution) margins. This layout enables researchers to simultaneously observe the bivariate joint relationship (correlation, clustering, covariance, and trend direction) alongside individual marginal properties (univariate skewness, modality, kurtosis, and dispersion).
How do marginal distributions differ from conditional and joint distributions?
A joint distribution f(X, Y) describes the simultaneous probability or occurrence of variable X and variable Y across coordinate space. A marginal distribution isolates one variable regardless of the values taken by the other by integrating or summing across the omitted dimension: f_X(x) = ∫ f(X=x, Y=y) dy. A conditional distribution f(Y|X=x) isolates the probability distribution of Y strictly constrained to an exact given slice where X equals a specific value x.
What is Kernel Density Estimation (KDE) and why is it overlaid on marginal histograms?
Kernel Density Estimation (KDE) is a nonparametric technique for estimating the continuous probability density function (PDF) of a random variable without assuming an underlying parametric family. When overlaid on discrete histogram bins, a Gaussian KDE curve f̂(x) = (1 / (n·h·√(2π))) ∑ exp(-½((x - x_i)/h)²) smooths out binning discretization artifacts, revealing underlying modalities and distribution peaks that might otherwise be masked by arbitrary bin widths or bin edge placements.
How does the Freedman-Diaconis rule determine the optimal histogram bin width?
The Freedman-Diaconis rule calculates bin width based on sample size n and the Interquartile Range (IQR): h = 2 · IQR(X) · n^(-1/3). Unlike Sturges' rule—which assumes Gaussian normality and under-bins skewed or multimodal data—the Freedman-Diaconis rule is robust to outliers and heavy-tailed non-Gaussian empirical distributions because IQR measures the spread of the middle 50% of the sample.
How can a scatterplot with marginal histograms reveal Simpson's Paradox?
Simpson's Paradox occurs when a statistical trend observed within distinct subgroups reverses or vanishes when the subgroups are aggregated. A standard 2D scatterplot of aggregated data may show an overall negative correlation, but the marginal histograms may expose distinct bimodal peaks along the X and Y axes. These separate marginal modes alert the analyst to hidden latent sub-populations (such as demographic cohorts or experimental treatments) that must be conditioned separately.
Can variables have identical marginal distributions yet completely different joint relationships?
Yes. Two bivariate datasets can have perfectly identical univariate marginal distributions along both axes (for example, standard normal distributions for both X and Y), yet possess radically different joint dependencies. In one dataset, X and Y might have a Pearson correlation of r = +0.95; in a second dataset, r = 0; in a third, a circular or parabolic non-linear dependency. Marginal histograms alone cannot establish dependency without the central scatterplot coordinate plane.