Data Regression Tool
Paste x and y values and find the curve that fits them best. Linear, quadratic, cubic, exponential, logarithmic and power models are fitted by least squares and ranked by adjusted R², with the fit, the residuals and a prediction table drawn for the winner.
How to Use
- Paste your data as two lines,
x: 1, 2, 3, …andy: 2, 4.1, …, or as one x y pair per line (a spreadsheet’s two columns work; lines without two numbers, such as a header, are skipped). Or click one of the six examples under the predict box. - Leave it on Auto to fit every model the data allows and pick the highest adjusted R². The starting data gives y ≈ 2.06x − 0.04.
- To force one model, click Linear, Quadratic, Cubic, Exponential, Logarithmic or Power. The comparison table still lists R² and adjusted R² for every model.
- Check the residual chart under the fit: a good model leaves the residuals (y − ŷ) scattered around zero with no curve or funnel shape.
- Type x values into the predict box, separated by commas, or leave it blank for predictions at the first, middle and last x and two steps beyond.
Data & best-fit curve
Residuals
Worked Example
The starting data, by hand. For x = 1–5 and y = 2, 4.1, 6.2, 8.1, 10.3: Σx = 15, Σy = 30.7, Σxy = 112.7 and Σx² = 55. The slope is (5 × 112.7 − 15 × 30.7) ÷ (5 × 55 − 15²) = 103 ÷ 50 = 2.06 and the intercept (30.7 − 2.06 × 15) ÷ 5 = −0.04, so y ≈ 2.06x − 0.04 with R² = 0.99962. Auto predicts 12.32 at x = 6 and 14.38 at x = 7.
Exponential growth. For the example 3, 4.5, 6.7, 10, 15, 22.4 at x = 0–5, Auto picks y ≈ 3.003·e^(0.4018x), whose R² rounds to 1. A straight line through the same points has R² = 0.92434 and predicts 27.21 at x = 7, against 50.02 from the exponential.
The common mistake: trusting the best fit outside the data. On the cooling example (100, 82, 67, 55, 45, 37, 30 at x = 0–6), Auto picks a cubic: its adjusted R² of 0.99999 just beats the exponential y ≈ 100.1·e^(−0.2001x) at 0.99997. Inside the data the two agree to within 0.12 at x = 6 (30.02 against 30.14). Beyond it they part: at x = 8 the cubic gives 17.93 and the exponential 20.2, and at x = 12 the cubic gives −13.4, below zero, while the exponential gives 9.076. Cooling follows an exponential (Newton’s law of cooling), so click Exponential before predicting.
Show Work
Formulas
Where Least Squares Came From
Adrien-Marie Legendre published the method of least squares in 1805, in a work on the orbits of comets. Carl Friedrich Gauss published his own account in 1809 and said he had used it since 1795; with it he had predicted where the dwarf planet Ceres would reappear, and astronomers found it there at the end of 1801.
The word “regression” comes from Francis Galton, who found in the 1880s that tall parents tend to have children shorter than themselves and short parents taller ones, a “regression towards mediocrity”. The adjusted R² used to rank models here was proposed by the agricultural economist Mordecai Ezekiel in 1930.
About This Tool
This tool answers “what shape is my data?”: it fits six models at once, ranks them by adjusted R², draws the winner with its residuals and predicts new values. The Regression mode of the Probability & Statistics Lab fits a straight line only, but adds the standard error, t and p-value of the slope; use it when the question is whether a linear trend is real. The Slope Calculator finds the line through exactly two points.
Everything runs in your browser; nothing is uploaded.
Related tools: Probability & Statistics Lab, Slope Calculator, Graphing Calculator, and Linear Interpolation Calculator.
Frequently Asked Questions
How does Auto choose the best model?
It fits every model the data allows and ranks them by adjusted R², which docks a little for each extra parameter. On the starting data the cubic has the highest plain R² (0.99984 against the line’s 0.99962), but its adjusted R² is lower (0.99934 against 0.9995), so the straight line wins. You can override the choice by clicking a model.
What does R² mean?
R² = 1 − SS_res ÷ SS_tot: the share of the variation in y that the model explains. For the starting data the residuals square to 0.016 and y varies by 42.452 about its mean, so R² = 1 − 0.016/42.452 = 0.99962. R² can even go negative when a model is worse than a flat line at the mean: on the parabola example the exponential scores −0.16589.
How are the exponential, logarithmic and power models fitted?
By straightening them first: ln y = ln a + bx for exponential, y = a + b ln x for logarithmic and ln y = ln a + b ln x for power, then fitting a line and transforming back. That needs all y > 0 (exponential), all x > 0 (logarithmic) or both (power); on the parabola example, with x from −3 to 3, the logarithmic and power models are left out. R² is always measured on the original data, so every model is compared on the same footing.
How many points do I need?
At least 2, and more than a model has parameters: a quadratic needs 4 points and a cubic 5. With exactly 4 points a cubic passes through every one of them (R² = 1) and says nothing about the trend, so it is left out.
What do the residuals tell me?
A residual is y − ŷ, the gap between a point and the curve. For the starting data they are −0.02, 0.02, 0.06, −0.1 and 0.04, scattered with no pattern. Force a line onto the exponential-growth example and they run 2.148, −0.1181, −1.684, −2.15, −0.9152, 2.719: a U shape, which says the curve is bending and a straight line is the wrong model, even though its R² is 0.92434.
How do I use the Data Regression Tool?
Just type your numbers. The answer shows up right away — there is no button to press. Change anything and it updates by itself.
Is it free? Does it work without internet?
Yes to both. It is free with no sign-up, and once the page has loaded it keeps working even with no internet.
Where does my data go?
Nowhere — every calculation runs on your own device. Nothing you enter is uploaded, logged, or stored.
Common Use Cases
Calibration curves
Fit sensor readings against known standards and read off the slope and offset.
Growth and doubling time
The exponential example gives b = 0.4018, so y doubles every ln 2 ÷ b = 1.725 units of x.
Diminishing returns
The logarithmic example adds 2.016 × ln 2 = 1.397 to y each time x doubles.
Scaling laws
The power-law example recovers an exponent of 1.5, the one in Kepler’s third law.
Trendlines
Paste two spreadsheet columns and get an equation to put in a report.
Last updated: