Statistics • Inferential Statistics Flagship

Linear Regression Calculator

Fit an Ordinary Least Squares (OLS) linear model ŷ = β₁x + β₀ to bivariate data. Calculate slope, intercept, Pearson correlation coefficient (r), coefficient of determination (R²), standard error (Sₑ), and inspect step-by-step residual tables and interactive scatterplot visualizations.

Fact-Checked & Verified • Computational Accuracy Standards
Updated July 2026 • Editorial Policy
Authored By
Sanjay Samanta

Lead Developer & Founder of Basic Math Tools. Specializes in browser-native computational algorithms and applied mathematics.

Reviewed & Verified By
Academic Review Board

Mathematics & curriculum specialists. Audited against standard algebraic and arithmetic principles.

Found an error or have an improvement suggestion? Report a calculation issue

Simple Linear Regression Calculator

Calculate ordinary least squares (OLS) regression line, correlation $r$, $R^2$, and inspect residual errors.

Example Datasets:
Comma, space, or newline separated
Count: 10 Mean x̄: 5.5000
Comma, space, or newline separated
Count: 10 Mean ȳ: 10.4200
Fitted Ordinary Least Squares Regression Model
ŷ = 1.8485x - 0.2533
Slope (m / β₁) 1.8485
Y-Intercept (b / β₀) -0.2533
Correlation (r) 0.9984
R² (Variance Explained) 0.9968 (99.68%)
Std Error (Sₑ): 0.3341
SS Regression: 281.25
SS Residual: 0.8945
F-Statistic: 2515.2

Interactive Y-Value Predictor plug any x to predict ŷ

Given X =
Predicted Value ŷ: 21.9287

Scatter Plot and Regression Fit Visualizer

Sample Points (xᵢ, yᵢ)
Regression Line ŷ = mx + b
Residual Distances eᵢ
Centroid (x̄, ȳ)

Step-by-Step OLS Formula Derivations

Sample Points and Residuals Table

Direct Answer & Overview
Verified Educational Guide

How to Calculate Simple Linear Regression

To find the linear regression equation ŷ = mx + b (or ŷ = β₁x + β₀), first calculate the means x̄ and ȳ. Compute the slope m = SS_xy / SS_xx = ∑(x - x̄)(y - ȳ) / ∑(x - x̄)². Then compute the y-intercept b = ȳ - m · x̄. The strength of the fit is given by Pearson r = SS_xy / √(SS_xx · SS_yy), where R² = r² indicates the percentage of variance in Y explained by X.

Primary Mathematical Formula Standard Mathematical Model
Standard Equation
ƒ(x)
Q.E.D.
ŷ = mx + b | m = ∑(x - x̄)(y - ȳ) / ∑(x - x̄)² | b = ȳ - m · x̄ | r = SS_xy / √(SS_xx · SS_yy)
Evaluated with exact mathematical formulation • Rigorously verified
Exact Formula
Input Parameters
Required
1
Independent Variable X (Predictor): Numerical values separated by commas or spaces
2
Dependent Variable Y (Response): Matched numerical values of equal length n
3
Prediction Target X_new: Optional X value to predict corresponding ŷ
Expected Outputs
Calculated
Regression Equation: ŷ = mx + b with 4-decimal precision
Slope & Intercept: Rate of change and baseline value at x = 0
Correlation Metrics: Pearson r (-1 to +1) and R² (0% to 100%)
Interactive SVG Plot: Sample points, best-fit line, centroid, and residual errors
Worked Numerical Example
Instant Verification
Find the best-fit line for (1, 2), (2, 3), (3, 5), (4, 4), (5, 6)
→ x̄ = 3, ȳ = 4. SS_xx = 10, SS_xy = 9. Slope m = 9/10 = 0.9. Intercept b = 4 - 0.9(3) = 1.3. Pearson r = 0.9487, R² = 0.9000 (90%).
Equation: ŷ = 0.9000x + 1.3000 | r = 0.9487 | R² = 90.0%

Ordinary Least Squares (OLS) and Line of Best Fit

In statistics and data analysis, Simple Linear Regression models the expected value of a continuous dependent variable $Y$ as a linear function of an independent explanatory variable $X$:

Y_i = \beta_0 + \beta_1 X_i + \varepsilon_i, \quad \text{where } \varepsilon_i \sim \mathcal{N}(0, \sigma^2)

The Ordinary Least Squares (OLS) criterion selects the line ŷ = β₀ + β₁x that minimizes the Sum of Squared Residuals (SS_res):

SS_{res} = \sum_{i=1}^n e_i^2 = \sum_{i=1}^n (y_i - \hat{y}_i)^2 = \sum_{i=1}^n (y_i - (\beta_0 + \beta_1 x_i))^2

By squaring each residual distance $e_i$, OLS penalizes large deviations quadratically, eliminates sign cancellations between positive and negative errors, and guarantees an analytically unique closed-form solution.

Slope, Intercept and Standard Error Formulas

Setting the partial derivatives of SS_res with respect to β₀ and β₁ to zero yields the fundamental normal equations:

Slope (β₁ / m)
\beta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} = \frac{n\sum x_i y_i - \sum x_i \sum y_i}{n\sum x_i^2 - (\sum x_i)^2}

Measures the expected change in Y for a one-unit increase in X.

Y-Intercept (β₀ / b)
\beta_0 = \bar{y} - \beta_1 \bar{x} = \frac{\sum y_i - \beta_1 \sum x_i}{n}

The baseline predicted value of Y when the predictor X = 0.

The Standard Error of the Estimate ($S_e$) measures the standard deviation of data points around the fitted regression line:

S_e = \sqrt{\frac{SS_{res}}{n - 2}} = \sqrt{\frac{\sum (y_i - \hat{y}_i)^2}{n - 2}}

Pearson Correlation ($r$) vs Coefficient of Determination ($R^2$)

While related, $r$ and $R^2$ describe distinct statistical dimensions of the bivariate relationship:

Pearson Correlation (r)
r ∈ [-1.0, +1.0]

Quantifies both the direction (positive or negative) and strength of linear alignment. $r = +1$ indicates a perfect positive slope; $r = -1$ indicates a perfect negative slope; $r = 0$ indicates no linear correlation.

Determination (R²)
R² ∈ [0.0, 1.0] (0% to 100%)

Quantifies the goodness-of-fit. An $R^2 = 0.85$ means that $85\%$ of the observed variation in $Y$ is linearly explained by $X$, while $15\%$ remains unexplained residual error.

Core Assumptions of Linear Regression (The LINE Acronym)

For hypothesis tests ($t$-tests, $F$-tests) and confidence intervals to remain statistically valid, the model must satisfy the four classical LINE assumptions:

L • Linearity

The relationship between the mean of $Y$ and $X$ is strictly linear. Diagnosed via scatter plots and residual plots.

I • Independence

Residuals are uncorrelated over time or space (no autocorrelation). Verified via Durbin-Watson tests.

N • Normality

The error terms εᵢ are normally distributed around the line. Checked using Q-Q plots and Shapiro-Wilk tests.

E • Equal Variance (Homoscedasticity)

The variance of residuals Var(εᵢ) = σ² is constant across all predictor values X.

Residual Analysis and ANOVA Sum of Squares

The total variability in $Y$ is partitioned into two orthogonal components: the variation explained by the model (SS_reg) and the unexplained error (SS_res):

SS_{tot} = SS_{reg} + SS_{res} \implies \sum (y_i - \bar{y})^2 = \sum (\hat{y}_i - \bar{y})^2 + \sum (y_i - \hat{y}_i)^2

The overall statistical significance of the regression model is tested with the ANOVA F-statistic:

F = \frac{MS_{reg}}{MS_{res}} = \frac{SS_{reg} / 1}{SS_{res} / (n - 2)}

Step-by-Step Worked Statistical Examples

Worked Example 1: Study Hours vs Exam Score
Dataset: X = [2, 4, 6, 8, 10], Y = [55, 65, 75, 80, 95] (n = 5)
  • Means: x̄ = 30 / 5 = 6.0, ȳ = 370 / 5 = 74.0
  • Deviations: SS_xx = (-4)² + (-2)² + 0² + 2² + 4² = 16 + 4 + 0 + 4 + 16 = 40
  • Covariance: SS_xy = (-4)(-19) + (-2)(-9) + (0)(1) + (2)(6) + (4)(21) = 76 + 18 + 0 + 12 + 84 = 190
  • Slope m = 190 / 40 = 4.75
  • Intercept b = 74.0 - (4.75 · 6.0) = 74.0 - 28.5 = 45.5
Best-Fit Equation: ŷ = 4.7500x + 45.5000 (R² = 97.4%)

Comparison: Simple Linear vs Multiple vs Polynomial Regression

Model Type Equation Form Number of Predictors Primary Use Case
Simple Linear ŷ = β₀ + β₁X 1 continuous predictor Bivariate trend analysis, straightforward interpretation
Multiple Linear ŷ = β₀ + β₁X₁ + β₂X₂ + ... 2+ predictors Controlling for confounding variables and covariates
Polynomial ŷ = β₀ + β₁X + β₂X² + ... 1 predictor (higher orders) Capturing curvilinear or non-linear parabolic trends

Frequently Asked Questions

What is simple linear regression and what does the line of best fit represent?
Simple linear regression is a statistical method used to model the linear relationship between a continuous independent predictor variable (X) and a continuous dependent outcome variable (Y). The line of best fit, ŷ = β₀ + β₁X (or y = mx + b), minimizes the sum of squared vertical distances (residuals) between the observed sample data points and the regression line.
How are the slope (β₁ / m) and y-intercept (β₀ / b) calculated in Ordinary Least Squares (OLS)?
Using Ordinary Least Squares (OLS), the slope is calculated as β₁ = SS_xy / SS_xx = ∑(x - x̄)(y - ȳ) / ∑(x - x̄)², or β₁ = r · (s_y / s_x). The y-intercept is computed from the sample means as β₀ = ȳ - β₁ · x̄, ensuring the regression line always passes through the centroid (x̄, ȳ).
What is the difference between Pearson correlation (r) and the coefficient of determination (R²)?
The Pearson correlation coefficient (r, ranging from -1 to +1) measures the strength and directional sign of the linear association between X and Y. The coefficient of determination (R² = r², ranging from 0 to 1 or 0% to 100%) represents the exact proportion of the total variance in the dependent variable Y that is explained by the regression model.
What are the four core assumptions of simple linear regression (LINE)?
The four classical Gauss-Markov assumptions are: 1. Linearity (the true underlying relationship between X and Y is linear), 2. Independence of errors (observations are independent without autocorrelation), 3. Normality of residuals (error terms ε_i are normally distributed with zero mean), and 4. Equal variance / Homoscedasticity (constant variance of residuals across all values of X).
How do you detect and evaluate outliers and influential leverage points?
Outliers are sample points with abnormally large vertical residuals |e_i| = |y_i - ŷ_i|. High-leverage points have extreme predictor values x_i far from the mean x̄. An influential point combines high leverage with an outlier residual, drastically altering the slope and intercept if removed. They are diagnosed using standardized residuals, Cook's distance, and leverage metrics (h_ii).
Can linear regression establish a cause-and-effect relationship between variables?
No. Correlation does not imply causation. A statistically significant regression model simply establishes mathematical covariance. Confounding factors, lurking variables, reverse causality, or accidental coincidence can produce strong linear correlations without any direct causal mechanism.