- Write an estimate/interval/unit/importance-threshold table.
Keep the substantive question visible
Describe how large a difference or association is and what it means for the research question. Report original units, relevant populations and timeframes. Standardized effects can help comparison but may obscure the meaning of the original scale.
Understand what a p-value does not establish
The ASA statement warns against interpreting a p-value alone as the probability that the null hypothesis is true, effect magnitude or scientific importance. Do not reduce research decisions to a threshold. Sample size and design influence the evidence.
Report uncertainty and assumptions
Present the estimate, interval and calculation method. Frequentist coverage concerns the procedure’s repeated-sampling behavior, rather than a probability assigned to a fixed parameter after the interval is observed. Model assumptions, dependence, selection bias and missingness can limit interpretation.
Address multiple analyses
Separate primary confirmatory outcomes from exploratory analyses and explain multiplicity handling. A nonsignificant finding alone demonstrates neither equivalence nor absence of an effect. Equivalence requires its own question, margin and suitable design.
A transparent reporting template
Teaching template: estimated difference [value and units], interval [lower, upper], method [model and settings], key limitation [limitation]. These placeholders are not actual results. Add the analytical sample and sensitivity to important decisions; an isolated number is insufficient.
Practical research checklist
- State effect, direction and units.
- Report the interval and method.
- Include all primary analyses.
- Keep conclusions within design and evidence.
Worked case and implementation decisions
The following is a fictional teaching case. Do not use its numbers or wording as actual study findings.
A teaching mean difference of 2 points with CI95=[−1,5] leaves a small negative and a larger positive effect compatible with the model. It proves neither no effect nor practical usefulness. Report units, interval method and design. If 4 points was the prespecified practical threshold, this interval remains uncertain relative to it. Distinguish raw and standardized effects and explain the denominator and target population rather than treating all effect sizes as interchangeable.
| Stage | Teaching example | Verification question |
|---|---|---|
| Estimate | Hypothetical 2 points | Are units and direction clear? |
| CI95 | Illustrative [−1,5] | Are method and assumptions suitable? |
| Importance | 4 is only a teaching threshold | Was the actual threshold justified prospectively? |
| Interpretation | Uncertainty remains | Is p distinct from size and precision? |
| Report | Raw, standardized and limits | Is the standardizing denominator defined? |
Exercise output: Write an estimate/interval/unit/importance-threshold table.
Sources and further reading
Official sources for verification and further reading

