Linear Regression Calculator
Fit an Ordinary Least Squares (OLS) linear model ŷ = β₁x + β₀ to bivariate data. Calculate slope, intercept, Pearson correlation coefficient (r), coefficient of determination (R²), standard error (Sₑ), and inspect step-by-step residual tables and interactive scatterplot visualizations.
Lead Developer & Founder of Basic Math Tools. Specializes in browser-native computational algorithms and applied mathematics.
Mathematics & curriculum specialists. Audited against standard algebraic and arithmetic principles.
How to Calculate Simple Linear Regression
To find the linear regression equation ŷ = mx + b (or ŷ = β₁x + β₀), first calculate the means x̄ and ȳ. Compute the slope m = SS_xy / SS_xx = ∑(x - x̄)(y - ȳ) / ∑(x - x̄)². Then compute the y-intercept b = ȳ - m · x̄. The strength of the fit is given by Pearson r = SS_xy / √(SS_xx · SS_yy), where R² = r² indicates the percentage of variance in Y explained by X.
Ordinary Least Squares (OLS) and Line of Best Fit
In statistics and data analysis, Simple Linear Regression models the expected value of a continuous dependent variable $Y$ as a linear function of an independent explanatory variable $X$:
The Ordinary Least Squares (OLS) criterion selects the line ŷ = β₀ + β₁x that minimizes the Sum of Squared Residuals (SS_res):
By squaring each residual distance $e_i$, OLS penalizes large deviations quadratically, eliminates sign cancellations between positive and negative errors, and guarantees an analytically unique closed-form solution.
Slope, Intercept and Standard Error Formulas
Setting the partial derivatives of SS_res with respect to β₀ and β₁ to zero yields the fundamental normal equations:
Measures the expected change in Y for a one-unit increase in X.
The baseline predicted value of Y when the predictor X = 0.
The Standard Error of the Estimate ($S_e$) measures the standard deviation of data points around the fitted regression line:
Pearson Correlation ($r$) vs Coefficient of Determination ($R^2$)
While related, $r$ and $R^2$ describe distinct statistical dimensions of the bivariate relationship:
Quantifies both the direction (positive or negative) and strength of linear alignment. $r = +1$ indicates a perfect positive slope; $r = -1$ indicates a perfect negative slope; $r = 0$ indicates no linear correlation.
Quantifies the goodness-of-fit. An $R^2 = 0.85$ means that $85\%$ of the observed variation in $Y$ is linearly explained by $X$, while $15\%$ remains unexplained residual error.
Core Assumptions of Linear Regression (The LINE Acronym)
For hypothesis tests ($t$-tests, $F$-tests) and confidence intervals to remain statistically valid, the model must satisfy the four classical LINE assumptions:
The relationship between the mean of $Y$ and $X$ is strictly linear. Diagnosed via scatter plots and residual plots.
Residuals are uncorrelated over time or space (no autocorrelation). Verified via Durbin-Watson tests.
The error terms εᵢ are normally distributed around the line. Checked using Q-Q plots and Shapiro-Wilk tests.
The variance of residuals Var(εᵢ) = σ² is constant across all predictor values X.
Residual Analysis and ANOVA Sum of Squares
The total variability in $Y$ is partitioned into two orthogonal components: the variation explained by the model (SS_reg) and the unexplained error (SS_res):
The overall statistical significance of the regression model is tested with the ANOVA F-statistic:
Step-by-Step Worked Statistical Examples
- Means: x̄ = 30 / 5 = 6.0, ȳ = 370 / 5 = 74.0
- Deviations: SS_xx = (-4)² + (-2)² + 0² + 2² + 4² = 16 + 4 + 0 + 4 + 16 = 40
- Covariance: SS_xy = (-4)(-19) + (-2)(-9) + (0)(1) + (2)(6) + (4)(21) = 76 + 18 + 0 + 12 + 84 = 190
- Slope m = 190 / 40 = 4.75
- Intercept b = 74.0 - (4.75 · 6.0) = 74.0 - 28.5 = 45.5
Comparison: Simple Linear vs Multiple vs Polynomial Regression
| Model Type | Equation Form | Number of Predictors | Primary Use Case |
|---|---|---|---|
| Simple Linear | ŷ = β₀ + β₁X | 1 continuous predictor | Bivariate trend analysis, straightforward interpretation |
| Multiple Linear | ŷ = β₀ + β₁X₁ + β₂X₂ + ... | 2+ predictors | Controlling for confounding variables and covariates |
| Polynomial | ŷ = β₀ + β₁X + β₂X² + ... | 1 predictor (higher orders) | Capturing curvilinear or non-linear parabolic trends |