- A dataset package with dictionary and validation evidence.
- A generation-method and limitations description.
- Access, version and executable reuse documentation.
What a data paper claims
A data paper explains a dataset’s provenance, quality, value and reuse. A Scientific Data Data Descriptor is not the same article type as a hypothesis-testing paper. Follow the journal’s current instructions for structure and files.
A useful description lets another researcher understand how records were generated and where they are limited. Size, an open licence or a DOI alone does not establish measurement quality.
Design records around reuse
Teaching case: synthetic student records contain study hours, a baseline score and a final score. Their permitted example use is learning analysis, not inference about real students. Real data require documented population, sampling, dates and instruments.
For each variable specify name, definition, type, unit, range, missing code and provenance. Separate raw from derived values. Use a stable lawful linkage key and keep personal identifiers out of public deposits.
Evidence for technical validation
Example code checks record counts, unique identifiers, score ranges, missingness and inconsistent values, then saves a report. Such format checks are different from instrument validity or measurement error; those need additional evidence.
Report the check, number of failures, action and remaining limitation. “The data are completely clean” is not a substitute for Technical Validation.
Repository, rights and access
Deposit the permitted data, README, dictionary, validation code, licence and versioned files in a suitable repository with stable links. If consent or contracts forbid public release, explain controlled access and appropriate public metadata.
FAIR does not require opening every dataset. Removing names may not prevent identification. Assess combinations of variables, small groups and location under the project’s ethical decision before release.
Usage notes and inferential limits
Provide an example that loads data, handles units, selects records and reproduces a descriptive table. Include expected output so users can check their setup.
Record the version, file hash, dependencies and limits to population representation. Present new hypothesis-based findings in an appropriate article type; statistical significance does not establish data quality.
Completed teaching worksheet
This is a hypothetical teaching case, not observed data, an actual review or a publication acceptance. Numbers illustrate decisions.
| Decision or record | Teaching example | Your project action |
|---|---|---|
| Records | 12 entirely synthetic teaching records | Document actual provenance and sampling. |
| Dictionary | study_hours measured in hours | Define range and missing codes. |
| Checks | Unique IDs and scores between 0 and 100 | Save failures and actions. |
| Access | Shareable synthetic example | Assess actual permission and identification risk. |
| Reuse | Mean table from a fixed version | Provide expected output and limitations. |
Deliverables and completion checks
- A dataset package with dictionary and validation evidence.
- A generation-method and limitations description.
- Access, version and executable reuse documentation.
Version and scope references: Scientific Data — Current submission guidelines · Scientific Data — Guide to referees
Sources and further reading
Official sources for verification and further reading
- Scientific Data — Current submission guidelines Source verification: 2026-10-06
- Scientific Data — Guide to referees Source verification: 2026-10-06
- Wilkinson et al. — The FAIR Guiding Principles (2016)
- MIT Libraries — Write a data management plan

