This website uses cookies to ensure you get the best experience on our website.
- Table of Contents
Most ELISA analysis software can generate a 4PL or 5PL standard curve in seconds. The curve may look smooth, the standards may appear close to the fitted line, and R² may be very close to 1.
That still does not tell you whether the fit is reliable enough for quantitation.
A reliable ELISA curve fit needs to do more than produce a smooth sigmoid and a high R². A curve can fit the data well overall while performing poorly in one part of the concentration range—and that matters if your samples fall in that region.
A more useful evaluation looks at residuals, back-calculated standards, and weighting. Residuals show where the model is missing the observed responses. Back-calculated standards show whether those misses translate into concentration error. Weighting becomes important when the variability of the response is not constant across the curve. [1–3]
For a broader workflow covering raw OD values, standard curves, and concentration calculations, see our guide to ELISA data analysis.
Quick Answer
R² alone is not enough to judge an ELISA curve fit because it summarizes overall agreement without showing where fitting errors occur. Residuals can reveal systematic error patterns, while back-calculated standards show whether those errors affect concentration estimates. If error variance changes across the curve, the weighting method should also be evaluated. [2]
Across its full working range, a quantitative ligand-binding assay rarely produces a straight concentration-response relationship.
At the lower end, responses approach the lower asymptote. Through the middle of the curve, signal changes more rapidly with concentration. At the upper end, responses approach the upper asymptote. On a logarithmic concentration axis, this commonly produces the familiar sigmoidal relationship. [1,2]
The four-parameter logistic (4PL) model describes the two asymptotes, the inflection region, and the slope. One characteristic of the 4PL is symmetry around its inflection point. [2]
The five-parameter logistic (5PL) model adds an asymmetry parameter. This allows it to describe curves in which the two sides of the sigmoid are not symmetric, and can improve concentration estimates when genuine asymmetry is present. [2,5]
ICH M10 likewise notes that ligand-binding assay calibration curves are most often fitted with 4PL or 5PL models when data are available near both asymptotes. [4]
If you are still working on how to construct and fit the standards themselves, see how to generate an ELISA standard curve.
The presence of a fifth parameter, however, does not make 5PL automatically superior. If 4PL already describes the assay adequately, more flexibility may add little practical value. Published LBA recommendations favor a sufficiently simple model that provides adequate quantitative performance. [2,3]
R² is appealing because it reduces a complicated fit to one familiar number.
The limitation is that it does not tell you where the errors occur.
Two curves can have similar R² values while producing different errors in the concentrations calculated from them. Azadeh and colleagues specifically note that statistics such as R² evaluate fit in response space and do not necessarily reflect the quality of the reportable inverse prediction—the concentration you ultimately want to know. [2]
Suppose a curve has an R² of 0.999, but the lowest three standards are consistently displaced from the fitted curve.
If your unknowns sit near the middle of the range, the practical impact may be small. If most of your samples are close to the lower end, the same fitting behavior becomes much more important.
Instead of stopping at “How high is my R²?”, ask:
Does the model recover concentrations reliably across the range I actually need?
Residuals are one way to see where that question needs closer attention.
A residual is the difference between the response observed for a standard and the response predicted by the fitted model:
If a standard produces an OD of 1.20 and the fitted model predicts 1.16, the residual is +0.04.
One residual is rarely informative by itself. The pattern across the curve is much more useful.
A well-behaved residual plot should not show persistent concentration-dependent bias. Published LBA recommendations describe random residual distribution without systematic bias at the low or high concentration levels as a desirable feature of model performance. [2]
What deserves attention is repeated structure.
If residuals tend to fall on one side of zero through one part of the range and shift to the other side elsewhere, the model may be missing something about the curve shape. Reproducible asymmetry is one possible explanation, although experimental problems should be considered before switching models.
A different pattern occurs when residual spread changes markedly across the response range. This is consistent with heteroscedasticity, or nonconstant response variance, which is well recognized in nonlinear ligand-binding calibration curves. [2,3] In that situation, weighting deserves closer examination.
An isolated large residual is different again. If one standard behaves poorly while adjacent concentrations do not, check the experimental point first. Replicate disagreement, a dilution error, pipetting, or another technical problem can make one standard look like a regression problem when it is not.
If standard preparation itself is questionable, the ELISA standard preparation guide is a better next step than immediately changing the regression model.
Residuals describe error in response units. ELISA results, however, are usually reported as concentrations.
Back-calculation connects the two.
Take a standard whose concentration is already known, treat its measured response as though it came from an unknown sample, and calculate its concentration from the fitted curve. You can then compare the result with the concentration that was actually prepared.
Example — Nominal concentration: 100 pg/mL | Back-calculated concentration: 93 pg/mL
%Recovery = (Back-calculated concentration / Nominal concentration) × 100
So the recovery is 93%.
%RE = [(Back-calculated concentration − Nominal concentration) / Nominal concentration] × 100
which gives −7%.
Findlay and Dillard recommend evaluating agreement between nominal and model-predicted calibrator values when assessing an LBA calibration model. [1] ICH M10 also requires reporting and evaluating back-calculated calibration-standard concentrations in regulated bioanalytical LBAs. [4]
| Back-calculation pattern | What deserves a closer look |
|---|---|
| Low standards are consistently underestimated | Low-end fit, signal separation, background, weighting |
| Low standards are consistently overestimated | Low-end response behavior, background, weighting |
| High standards become progressively less accurate | Upper-end compression, saturation, curve fit |
| Middle standards perform well but the extremes do not | The usable quantitative range may be narrower than the full fitted curve |
| One standard performs poorly while adjacent standards do not | Standard preparation, replicate variation, technical artifact |
| Error changes direction systematically across the range | Curve shape or regression model |
These are diagnostic clues, not one-to-one diagnoses.
There is no single acceptance threshold that should automatically be applied to every research-use ELISA.
For regulated ligand-binding bioanalytical assays, ICH M10 specifies that the accuracy and precision of back-calculated calibration-standard concentrations should be within ±20% of nominal at most levels and within ±25% at the LLOQ and ULOQ. At least 75% of calibration standards, excluding anchor points and including at least six concentration levels, should meet the specified criteria. [4]
Those thresholds belong to the regulated LBA context. They should not be presented as universal pass/fail rules for every exploratory or research-use ELISA.
For a research assay, the acceptance criteria should reflect the purpose of the experiment, assay performance, and predefined laboratory procedures.
Nonlinear ligand-binding calibration curves frequently show unequal response variability across concentration levels. This is one reason an unweighted fit can perform well in one part of the curve while giving poorer inverse predictions elsewhere. [2,3]
Common options available in curve-fitting software include unweighted regression and weighting functions such as 1/Y and 1/Y².
The important point is not that one of these is universally preferable.
Weighting changes how different observations influence estimation of the curve parameters. The appropriate approach should be based on the assay’s response-error relationship and then assessed using quantitative performance, including the accuracy and precision of back-calculated concentrations. [2,3]
Xiang and colleagues presented three quantitative LBA case studies in which weighting functions were selected after evaluating heteroscedasticity, and the resulting weighted models were assessed using back-calculated %RE and %CV. The appropriate weighting differed among their assays. [3]
That is why 1/Y² should not be treated as a default answer to every poor ELISA curve.
If changing the weighting improves the low end but worsens concentration recovery elsewhere, the error has not necessarily been solved—it may simply have been redistributed.
After changing the weighting, go back to the same questions: Does the residual pattern improve? Do the standards back-calculate more appropriately? Does the fit perform better across the range that will actually be used?
Most curve-fitting software makes it easy to switch between 4PL and 5PL. Selecting the model solely because one produces a slightly better R² is less useful.
A 4PL model is often sufficient when the calibration response is reasonably symmetric and quantitative performance across the intended range is acceptable.
A 5PL model becomes more relevant when the curve shows reproducible asymmetry. Gottschalk and Dunn showed that the additional asymmetry parameter can improve fitting and concentration estimation for asymmetric assay data. [5] Azadeh et al. similarly describe 5PL as an option when a symmetric 4PL does not adequately represent an asymmetric calibration curve. [2]
The important word is reproducible.
One unusual plate or one questionable standard is not strong evidence that the assay requires a five-parameter model.
If the 5PL improves the apparent fit but produces little meaningful improvement in back-calculated concentrations, the additional complexity may not be useful. Xiang et al. recommend comparing quantitative accuracy and precision between candidate models and favoring the simpler model when performance is comparable. [3]
In a noncompetitive assay such as a typical sandwich ELISA, the response is generally directly related to analyte concentration. In a competitive assay, the relationship can be inverse. [2]
The orientation of the sigmoid changes, but the diagnostic questions remain much the same.
You still need to know whether the model describes the standards appropriately, whether residuals show systematic bias, whether known concentrations can be recovered, and whether the weighting and quantitative range are suitable.
A descending calibration curve therefore does not require an entirely different philosophy of curve evaluation.
A useful review can be reduced to five passes through the data.
Before changing any fitting settings, look at replicate agreement, background, signal separation, saturation, and obvious dilution problems.
Many curve problems begin before the analysis stage. The ELISA experimental design checklist is useful when standards, controls, dilution strategy, or plate setup may be contributing.
If the assay already has an established analysis procedure, use it consistently. During assay development, begin with a justified model rather than trying multiple equations and choosing whichever produces the highest R². Both published best-practice recommendations and ICH M10 emphasize selecting an appropriate regression model during method development or validation. [2–4]
Residuals answer: Where is the model missing the observed response? Back-calculation asks: Does that error materially change the reported concentration? Neither needs to be treated as an isolated pass/fail metric.
If the problem is concentrated near one end of the range, investigate the assay signal, quantitative range, heteroscedasticity, and weighting. If residuals show reproducible shape-related bias, then a 4PL-versus-5PL comparison becomes more informative.
If samples themselves frequently fall outside the usable range, choosing an appropriate ELISA dilution may be more important than changing the regression equation.
A model change should improve the problem that prompted the change without creating an unacceptable problem elsewhere. Do not stop simply because the new curve looks smoother.
Looking at residuals and back-calculated standards together is more informative than treating either one alone.
| Residual pattern | Back-calculation pattern | What to investigate next |
|---|---|---|
| No obvious systematic bias | Standards recover appropriately across the working range | Fit is more likely to be adequate; continue routine QC review |
| Systematic curvature | Relative error changes direction across the curve | Curve shape, reproducible asymmetry, 4PL vs 5PL |
| Residual spread changes with response | Poor recovery is concentrated at one or both ends | Response-error relationship and weighting |
| One large residual | One standard fails while neighbors perform normally | Replicates, dilution, pipetting, standard preparation |
| Upper-end residuals and compressed responses | High standards back-calculate poorly | Saturation and upper quantitative range |
| Lower-end deviations | Lowest standards repeatedly recover poorly | Background, low-end signal separation, weighting, practical LLOQ |
| Residual pattern improves after weighting | Recovery also improves across the intended range | Weighting change is supported by quantitative performance |
| Fit statistic improves with 5PL but recovery changes little | Similar concentration accuracy with 4PL and 5PL | Additional model complexity may not be justified |
This matrix is intentionally diagnostic rather than prescriptive. The same visible pattern can have more than one cause, and the regression model should not be changed until the experimental data have also been reviewed.
Curve-fitting software can draw a smooth line through imperfect data. That does not mean the assay itself behaved well.
If several high standards produce nearly indistinguishable responses, the problem may be saturation rather than regression. In that case, see the guide to ELISA saturated signals.
If the abnormality follows plate position—for example, perimeter wells behave differently from central wells—review ELISA edge effects before attributing the pattern to the curve model.
Standard preparation, poor replicate precision, washing or incubation variation, background, and other experimental problems can also distort calibration data. The ELISA troubleshooting guide provides a broader assay-level troubleshooting workflow.
Very high analyte concentrations can create another problem. ICH M10 specifically recommends evaluating dilution linearity and high-dose hook effect in ligand-binding assay validation because high concentrations can suppress signal and produce erroneous results. [4]
None of these problems can be fixed simply by adding another regression parameter.
A concentration table alone is not enough to evaluate a calibration model.
Ideally, the analysis environment should let you review the fitted curve alongside the observed standards and provide enough output to evaluate residual behavior, nominal versus back-calculated concentrations, quantitative accuracy, the selected regression model, and the weighting method being used. [2]
Not every ELISA analysis platform exposes all of these diagnostics. If your software only displays the fitted curve and R², additional statistical output may be needed before making decisions about model reliability.
For routine online ELISA data processing and concentration analysis, Boster provides an online ELISA data analysis tool. Use the diagnostic outputs available in your analysis software when evaluating residuals, weighting, and model performance.
R² can remain part of your curve review. It simply should not make the decision by itself.
A stronger assessment combines three views of the same calibration curve:
If the raw standards are sound, residual bias is limited, known concentrations are recovered adequately across the intended quantitative range, and alternative weighting or model choices do not reveal a meaningful improvement, the fit has much stronger support than a high R² alone can provide.
The objective is not the mathematically prettiest sigmoid.
It is a calibration model that produces defensible concentration estimates where your samples actually fall.