Research guide

Equivalence testing with TOST and a justified SESOI

Why a large p-value does not prove no effect: interpret equivalence bounds, standard errors and confidence intervals with executable examples.

Illustrative charts and a data-analysis notebook in an academic workspace
Prepared by: Dr. Didgar Research Institute · Last revised: · 3 min read · Guide created:
Expected deliverables
  • A prospective SESOI rationale and test plan.
  • Reproducible intervals and both one-sided p-values.
  • A difference/equivalence report with uncertainty and limits.
Workbook and completed exampleCode, synthetic data and executable examples

Frame absence of an important effect

Failure to reject zero may reflect imprecise data. Equivalence asks whether an effect is sufficiently precise to fall within justified bounds of unimportant effects. Justify the SESOI using scientific relevance, costs, practical consequences or prior evidence before seeing results.

Give bounds explicit units. The teaching bounds −0.5 to +0.5 score points are neither universal nor standardized d. Record stakeholder agreement and the basis for the margin.

Two one-sided tests and intervals

TOST requires rejecting both an effect at or below the lower bound and an effect at or above the upper bound. With compatible tests at α=0.05, the corresponding 90% confidence interval entirely inside the bounds expresses the same decision.

Distribution, degrees of freedom and dependence matter. Our executable example uses a normal approximation with a known or large-sample standard error; it does not replace a design-specific t-test, clustered analysis or mixed model.

A reproducible numerical example

A hypothetical estimate of 0.20, SE=0.15 and bounds [−0.50,0.50] yield CI90 approximately [−0.047,0.447] and the larger one-sided p approximately 0.0228. Under this teaching model, both nonequivalence null hypotheses are rejected.

CI95 approximately [−0.094,0.494] includes zero, so this case supports equivalence without a statistically significant difference from zero. Replacing SE with 1 produces an inconclusive equivalence result. Do not confuse standard error with the data’s standard deviation.

Report four possible outcomes

Evidence may support a difference, equivalence, both or neither: they ask different questions. Do not label neither as proof of no effect. Report the estimate, sign, uncertainty, bounds, α and test.

Teaching wording: “Under the normal model and illustrative ±0.5 bounds, the estimate was 0.20 and its CI90 lay within the bounds; interpretation is restricted to these units, model and scope.” Replace this with actual study results.

Plan sample size

Specify bounds, variability, design, attrition, target power and testing method before collection. Do not reuse a difference-test power calculation without assessing its suitability for equivalence. Complex designs need appropriate formulas or simulation and specialist review.

Use sensitivity analysis for defensible margins and uncertainties. Changing bounds to obtain the preferred decision or selecting the sample only for a desired p-value is not a defensible plan.

Completed teaching worksheet

This is a hypothetical teaching case, not observed data, an actual review or a publication acceptance. Numbers illustrate decisions.

Completed teaching worksheet
Decision or recordTeaching exampleYour project action
QuestionAn effect smaller than ±0.5 pointsJustify actual units and SESOI.
Estimate0.20 with SE=0.15Obtain a design-appropriate SE.
CI90[−0.047,0.447]Check the whole interval against bounds.
Low precisionSame estimate with SE=1Do not infer absence from uncertainty.
Reportp_TOST≈0.0228 in normal exampleState method, bounds and limits.

Deliverables and completion checks

  • A prospective SESOI rationale and test plan.
  • Reproducible intervals and both one-sided p-values.
  • A difference/equivalence report with uncertainty and limits.

Sources and further reading

Official sources for verification and further reading

This guide supports research learning and planning; align implementation with the actual design and institutional requirements. Editorial policy
Back to top