Scatterplot with Marginal Histograms Calculator
Plot bivariate coordinate observations alongside aligned top and right univariate marginal frequency histograms. Analyze joint correlation, Gaussian Kernel Density Estimation (KDE) curves, Ordinary Least Squares trendlines, and distribution modality across both dimensional projections simultaneously.
| Series | Mean (μ) | Std Dev (σ) | Min → Max |
|---|---|---|---|
| X (Top) | 37.14 | 15.42 | 12.0 → 65.0 |
| Y (Right) | 81.86 | 33.40 | 28.0 → 142.0 |
Foundations of Joint Distributions & Marginal Histograms
In exploratory data analysis and mathematical statistics, analyzing a single variable in isolation (univariate analysis) frequently conceals critical real-world phenomena, while inspecting a two-dimensional scatterplot without marginal context can mislead researchers regarding the underlying distributions of individual dimensions. The scatterplot with marginal histograms resolves this dilemma by unifying bivariate coordinate projections and univariate density distributions into a cohesive, synchronized visual framework.
Mathematically, let (X, Y) be a continuous bivariate random vector governed by the joint probability density function f_{XY}(x, y). The central scatterplot visualizes empirical draws sampled from this joint density. The probability that an observation falls within an infinitesimal Cartesian region dx × dy is:
The marginal probability density function of X, denoted f_X(x), represents the total probability distribution of X integrated over all plausible values of Y. Conversely, the marginal density of Y, denoted f_Y(y), integrates over all possible values of X:
By aligning the horizontal bins of the top marginal histogram precisely with the X-axis coordinate domain of the central scatter plane, and rotating the vertical bins of the right marginal histogram by 90° to match the Y-axis range, this visualization preserves spatial geometry. Every point in the scatterplot vertically projects its mass into the top histogram bin and horizontally projects its mass into the right histogram bin.
Mathematical Architecture: Pearson r, Covariance & Regression
The interactive joint plot computes essential bivariate moments and ordinary least squares linear regression parameters across the n coordinate pairs {(x_1, y_1), (x_2, y_2), …, (x_n, y_n)}.
1. Univariate Centroids & Standard Deviations
The sample means x̄ and ȳ locate the balance centroid of the joint distribution, displayed as the dashed crosshair intersection:
2. Sample Covariance & Pearson Product-Moment Correlation (r)
The degree of linear co-variation between X and Y is quantified by sample covariance Cov(X, Y) and normalized Pearson correlation r:
r = s_{xy} / (s_x · s_y) = ∑ (x_i - x̄)(y_i - ȳ) / √[ ∑ (x_i - x̄)² · ∑ (y_i - ȳ)² ]
Where -1 ≤ r ≤ +1. The coefficient of determination R² = r² represents the proportion of variance in Y predictable from X under a linear specification.
3. Ordinary Least Squares (OLS) Trendline
The fitted trendline ŷ = b_1 x + b_0 minimizes the sum of squared vertical residuals ∑ (y_i - ŷ_i)²:
Kernel Density Estimation (KDE) & Nonparametric Smoothing
While histograms provide an intuitive summary of counts partitioned into discrete intervals, their visual shape depends heavily on arbitrary bin widths and bin boundary choices. A minor shift in origin can dramatically alter whether a histogram appears unimodal or bimodal. Kernel Density Estimation (KDE) circumvents this discretization bias by placing a continuous, smooth weighting function (the kernel) over each observed data point and aggregating them into a smooth density curve.
In our calculator engine, we implement the standard Gaussian Kernel, satisfying the regularity conditions ∫ K(u) du = 1 and ∫ u K(u) du = 0:
Bandwidth Selection: Silverman's Rule of Thumb
The smoothing parameter h (bandwidth) controls the bias-variance tradeoff of the estimated density. An excessively small h leads to undersmoothing (spurious peaks from individual points), whereas an excessively large h oversmooths, flattening true multimodal features. Our engine calculates the optimal bandwidth using Silverman's robust formula:
Using the minimum between the sample standard deviation s_x and the scaled Interquartile Range (IQR / 1.34) ensures stability against extreme outliers and skewed tails.
Optimal Binning Rules: Freedman-Diaconis, Scott & Sturges
The number of histogram bins k and bin width w dictate the resolution of the marginal histograms. Statistical literature defines three primary formulations for determining binning:
Freedman-Diaconis Rule
Optimized to minimize the L² integrated mean squared error without assuming Gaussian normality. Highly robust against outliers:
k = ⌈ (max - min) / w ⌉
Scott's Normal Rule
Asymptotically optimal for unimodal Gaussian random variables by minimizing mean integrated squared error (MISE):
k = ⌈ (max - min) / w ⌉
Sturges' Rule
Based on the binomial expansion (1 + 1)^n. Tends to over-smooth for n > 200 or non-normal data:
w = (max - min) / k
Our calculator defaults to k = 10 bins while providing an interactive slider (4 ≤ k ≤ 25) so you can dynamically inspect how changing bin resolution impacts the perception of modal peaks and edge tails.
Diagnosing Bivariate Phenomena: Modality, Skewness & Variance
When examining joint distributions, marginal histograms expose essential statistical anomalies that a standard 2D scatterplot might obscure:
1. Multimodality and Mixture Distributions
A scatterplot displaying points scattered across a wide region might mask distinct clusters if their bounding ellipses overlap. If the top marginal histogram shows two separated peaks (bimodal distribution), it confirms that X consists of a mixture of two distinct sub-populations with different latent means.
2. Asymmetric Skewness & Long-Tailed Residuals
Standard linear regression assumes that vertical residuals are normally distributed with zero mean. If the right marginal histogram of Y exhibits severe positive skewness (a long tail stretching toward higher values), linear regression coefficients will be disproportionately pulled by high-leverage outliers, violating ordinary least squares assumptions.
3. Heteroscedasticity (Non-Constant Variance)
In a funnel-shaped scatterplot where point dispersion expands as X increases, conditional variance Var(Y|X=x) is non-constant. Inspecting the top marginal histogram clarifies whether observations are uniformly dense across X or concentrated primarily in the narrow-variance region.
Step-by-Step Numerical Algorithm for Constructing Marginal Histograms
The algorithmic pipeline executed by our 2D canvas rendering engine proceeds through seven deterministic phases:
Coordinate Domain Parsing & Padding
Extract n coordinate pairs. Compute bounds [x_{min}, x_{max}] and [y_{min}, y_{max}]. Apply a 5% safety margin padding on each boundary to ensure peripheral points and histogram edges do not touch canvas view boundaries.
Linear Affine Transform Mapping
Map world coordinates (x, y) to pixel screen coordinates (u, v) within the central bounding box:
u = u_{min} + ((x - x_{min}) / Δx) · W_{scatter}, v = v_{max} - ((y - y_{min}) / Δy) · H_{scatter}
Univariate Partitioning into k Uniform Bins
Partition domain intervals into k equal intervals of width w_x = Δx / k and w_y = Δy / k. Increment frequency counts c_j for each coordinate falling in [B_j, B_{j+1}).
Top Marginal Histogram Bar Rendering
Calculate bar heights proportional to c_j / c_{max,x}. Render vertical rectangular bars anchored to the top baseline, perfectly spanning horizontal intervals [u_j, u_{j+1}].
Right Marginal Histogram Bar Rendering (90° Rotation)
Render horizontal rectangular bars anchored to the right boundary of the central scatter plane. Bar lengths extend to the right proportional to c_m / c_{max,y}, matching vertical screen coordinates [v_{m+1}, v_m].
Continuous Kernel Density Curve Evaluation
Sample 100 evaluation steps along each margin. Evaluate Gaussian density f̂(t) = (1/(n·h)) ∑ K((t - t_i)/h). Scale maximum curve peak to 92% of margin height and stroke with emerald and amber path geometries.
Dynamic Canvas Resizing & High-DPI Rendering
Scale canvas buffer by device pixel ratio window.devicePixelRatio to ensure razor-sharp graphics on Retina and mobile screens.
Comprehensive Bivariate Distribution Archetypes Comparison Table
The table below contrasts standard joint distribution archetypes against their diagnostic signatures across the scatter plane, marginal histograms, and correlation metrics:
| Archetype | Scatter Morphology | Top Marginal (X) | Right Marginal (Y) | Pearson r | Primary Risk / Fallacy |
|---|---|---|---|---|---|
| Positive Linear Correlation | Narrow diagonal elliptical swath sloping upward | Unimodal, approximately normal | Unimodal, approximately normal | 0.80 to 0.95 | Conflating correlation with direct causation |
| Bivariate Gaussian (Spherical) | Circular cloud with radial density decay | Symmetric Gaussian bell curve | Symmetric Gaussian bell curve | ≈ 0.00 | Assuming r = 0 implies total statistical independence |
| Bimodal Mixture (Two Clusters) | Two distinct, disconnected point clouds | Twin peaks (bimodal distribution) | Twin peaks (bimodal distribution) | Spurious positive or negative | Simpson's Paradox; ignoring latent sub-groups |
| Right-Skewed Exponential Tail | Dense cluster at bottom-left with scattered outliers | Severe right skew (long right tail) | Severe right skew (long top tail) | Artificially inflated by outliers | OLS violation; failure to apply logarithmic transform |
| Non-Linear Quadratic (U-Curve) | Symmetric parabolic curve y ≈ x² | Uniform or symmetric normal | Strongly asymmetric; mass at minimum | ≈ 0.00 | Falsely concluding no relationship exists |
Graded Worked Statistical Problems with Full Solutions
Consider the sample of n = 5 observations: {(2, 4), (4, 7), (5, 8), (7, 11), (10, 15)}.
Tasks: (a) Calculate sample means x̄ and ȳ. (b) Compute sample covariance s_{xy} and Pearson correlation r. (c) Construct k = 2 uniform bins for the marginal distribution of X and state bin counts.
Complete Step-by-Step Solution:
Step 1: Compute Centroids.
x̄ = (2 + 4 + 5 + 7 + 10) / 5 = 28 / 5 = 5.6
ȳ = (4 + 7 + 8 + 11 + 15) / 5 = 45 / 5 = 9.0
Step 2: Deviations and Cross-Products.
i=1: (2 - 5.6)(4 - 9) = (-3.6)(-5) = +18.0
i=2: (4 - 5.6)(7 - 9) = (-1.6)(-2) = +3.2
i=3: (5 - 5.6)(8 - 9) = (-0.6)(-1) = +0.6
i=4: (7 - 5.6)(11 - 9) = (+1.4)(+2) = +2.8
i=5: (10 - 5.6)(15 - 9) = (+4.4)(+6) = +26.4
Sum of cross-products: SS_{xy} = 18.0 + 3.2 + 0.6 + 2.8 + 26.4 = 51.0.
Sample covariance: s_{xy} = 51.0 / (5 - 1) = 51.0 / 4 = 12.75.
Step 3: Sum of Squared Deviations & Correlation.
SS_x = (-3.6)² + (-1.6)² + (-0.6)² + 1.4² + 4.4² = 12.96 + 2.56 + 0.36 + 1.96 + 19.36 = 37.2
SS_y = (-5)² + (-2)² + (-1)² + 2² + 6² = 25 + 4 + 1 + 4 + 36 = 70.0
r = 51.0 / √(37.2 · 70.0) = 51.0 / √(2604) = 51.0 / 51.0294 ≈ +0.9994.
Step 4: Marginal Bins for X.
Domain of X: [x_{min}, x_{max}] = [2, 10]. Range Δx = 8. For k = 2 bins, bin width is w = 8 / 2 = 4.
Bin 1: [2, 6) → values {2, 4, 5} → Count = 3 (60%).
Bin 2: [6, 10] → values {7, 10} → Count = 2 (40%).
An observational study measures hours studied (X) versus exam error score (Y). The aggregate dataset shows r = +0.72 (suggesting more studying yields more errors). However, the top marginal histogram of X exhibits two isolated peaks at x = 4 and x = 14.
Question: How does the marginal histogram reveal the underlying confounder, and what occurs when the clusters are conditioned separately?
Diagnostic Explanation:
The prominent bimodal signature in the top marginal histogram alerts the analyst that the sample is not a homogeneous univariate draw, but an aggregated mixture of two distinct cohorts: Group A (Introductory level students studying 2–6 hours, baseline errors 2–5) and Group B (Advanced Quantum Mechanics students studying 12–16 hours, baseline errors 15–20).
Within Group A, the conditional correlation is r_{A} = -0.85. Within Group B, r_{B} = -0.82. In both separate groups, studying reduces errors. However, because the advanced group has both higher study demands and inherently higher test complexity, aggregating the groups shifts the centroid from (4, 3.5) to (14, 17.5), producing a spurious aggregate positive slope r_{agg} = +0.72. The marginal histogram's twin peaks provide the visual proof required to identify Simpson's Paradox.
Cross-Disciplinary Real-World Applications
Genomics & Bioinformatics
In single-cell RNA sequencing (scRNA-seq), joint plots graph expression levels of two candidate genes across tens of thousands of individual cells. Marginal histograms reveal zero-inflation and bimodal expression states (active transcription vs. transcriptional silencing).
Quantitative Finance & Risk
Risk managers plot asset returns against benchmark market index returns. Right and top marginal histograms expose heavy negative tails (leptokurtic excess kurtosis), proving that asset returns deviate from Gaussian assumptions and possess severe tail-risk downside.
Meteorology & Climate
Climate scientists plot atmospheric temperature against relative humidity. Joint plots identify dew point saturation boundaries where the joint distribution truncates abruptly, while marginal histograms reflect seasonal temperature oscillations.
Common Pitfalls, Artifacts & Simpson's Paradox to Avoid
1. Bin Edge Placement Artifacts
Selecting too few bins (k < 5) averages away subtle multimodal peaks, while selecting too many bins (k > 30) produces jagged, noisy combs where individual bins contain single observations. Always toggle the Gaussian KDE curve to confirm whether a histogram peak is an authentic distribution mode or an artifact of bin boundary placement.
2. Confusing Marginal Independence with Joint Independence
Two random variables are independent if and only if f_{XY}(x, y) = f_X(x) · f_Y(y) for all coordinates. Observing that marginal histograms are symmetric and bell-shaped does not guarantee that the joint distribution is bivariate normal. Highly non-linear dependencies (such as circular manifolds or parabolic relationships) can possess perfectly normal marginals while being entirely dependent.
3. Over-Reliance on Pearson r in Non-Linear Regimes
Pearson's correlation coefficient r measures strictly linear association. For data following a symmetric U-shaped parabola y = x² over [-10, 10], Pearson r ≈ 0.00 despite a deterministic mathematical relationship. Always visually inspect the central scatterplot rather than relying solely on summary statistics.
Connected Graphing & Statistics Ecosystem Hub
Deepen your exploratory statistical and regression analysis with complementary mathematical graphing instruments across our suite:
Lead Developer & Founder of Basic Math Tools. Specializes in browser-native computational algorithms and applied mathematics.
Mathematics & curriculum specialists. Audited against standard algebraic and arithmetic principles.