Research guide

Research data management and reproducibility

Build a data dictionary, document cleaning and changes, protect confidential information and preserve code and analysis environments for review.

Illustration of a research desk, notes and analysis tools
Prepared by: Dr. Didgar Research Institute · Last revised: · 3 min read
Expected deliverables
  • Produce a dictionary, cleaning log and access/retention record.
Decision workbook and example

Data dictionary

Define variable names, meaning, units, types, valid ranges and missing-value codes. Separate participant identities from research identifiers. Record transformations and preserve an unchanged raw-data copy.

Cleaning and quality control

Check ranges, duplicate records and inconsistencies. Excluding outliers or imputing missing data requires methodological justification. Report sample counts through cleaning stages so their consequences are visible.

Protection and transfer

Transfer only necessary information and restrict access. Align encryption, backup and retention with institutional requirements. A public messaging channel is not a substitute for an agreed secure process for sensitive data.

Reproducing results

Retain analysis code, software versions and execution order. Data sharing requires consent, permission and regulatory compatibility. If sharing is restricted, document methods and dictionaries or clearly labelled synthetic examples.

A data-dictionary example

For a teaching variable named score, record the construct, instrument, numeric type, permitted range, measurement time, missing code and derivation. Do not impose an invented range on real data. Record formulas and inputs for derived variables.

Versions and change control

Use meaningful filenames and versions. Separate raw and processed data and preserve transformations in code or a decision log. Test backup restoration rather than merely counting copies. Assign folder responsibilities and permissions in collaborative work.

Close the project carefully

An authorized colleague should understand the dataset using the README and dictionary. Retention or deletion follows institutional rules, consent and project agreements rather than a universal period. Match publication availability statements to actual releases and permissions.

Worked case and implementation decisions

The following is a fictional teaching case. Do not use its numbers or wording as actual study findings.

In the fictional dictionary, score is numeric, measured in points, with a teaching range of 0–100 and blank missingness. A value −99 used as a documented missing code must not enter the mean, but replacing it without consulting the dictionary is also wrong. Preserve raw data and encode transformations. Separate linkage IDs from personal identifiers and restrict access in real projects. Retention follows consent, contracts and institutional rules rather than one period for all research.

Worked case and implementation decisions
StageTeaching exampleVerification question
Variablescore, points and measurement timeIs the actual definition documented?
MissingnessDocumented code, not analyst assumptionIs conversion correct before calculation?
TransformCode from raw to processedAre changes counted?
AccessResearch ID separate from identityAre permissions and controls adequate?
PreservationVersions, backup and restoration testDoes retention match project policy?

Exercise output: Produce a dictionary, cleaning log and access/retention record.

Sources and further reading

Official sources for verification and further reading

This guide supports research learning and planning; align implementation with the actual design and institutional requirements. Editorial policy
Back to top