- Deliver a documented execution bundle with known-answer tests.
Guide to choose languages according to application, design patterns and production-ready recommendations.
Python
Suitable for data analysis, machine learning and lightweight web development. Key packages include NumPy, Pandas, scikit-learn and TensorFlow. Advice on project structure and virtual environments (virtualenv/conda).
R
Specialized for statistical analysis and advanced visualization.
C++, Java, JavaScript, SQL
Application areas, performance considerations, memory management tips, project structure and input/output security.
Best practices
- Version control (Git) with proper branching
- Reproducible environments (Docker, conda)
- Documentation and unit testing
Scientific code should be reproducible
A script running on one computer is not sufficient. Document inputs, outputs, environment, dependencies and assumptions. Keep credentials out of source code. Avoid data leakage during preprocessing and model selection.
- README with exact execution instructions and sample inputs
- Versioned dependency and license information
- Meaningful checks and a minimal reproducible example
- Document validation, limitations and computational cost
A suggested scientific project layout
project/
README.md
requirements.txt
src/
tests/
data/README.md
outputs/Describe data provenance, permissions and preparation. Sensitive files need not belong in the repository. Keep regenerable outputs separate from source code and credentials outside version control.
Validating simulations
Start with a simple case with known or theoretically expected behavior before comparing complex scenarios. Document sensitivity to parameters, step size and random seeds. Numerical precision and model validity are different: precise calculations from inappropriate assumptions may mislead. Explain usage limits and how to change parameters in the final documentation.
Worked case and implementation decisions
The following is a fictional teaching case. Do not use its numbers or wording as actual study findings.
The downloadable mini-project contains 12 synthetic records and standard-library Python. It validates inputs, creates summaries and six illustrative regression paths, and records a file hash. This demonstrates a pipeline rather than actual research evidence. Quarto document dependencies are documented separately. Review statistical methods and assumptions for the real problem: successfully creating a file alone does not validate a model. Known-answer tests verify calculations, while a README explains their scope and limitations.
| Stage | Teaching example | Verification question |
|---|---|---|
| Input | synthetic.csv and dictionary | Is the original preserved? |
| Checks | IDs, types, ranges and quality flags | Are invalid inputs explicitly rejected? |
| Execution | research.py from the project folder | Are hidden personal paths avoided? |
| Outputs | CSV, JSON, SVG and hash | Do they match expected results? |
| Scope | Teaching example without causal inference | Are limitations in the README? |
Exercise output: Deliver a documented execution bundle with known-answer tests.
Sources and further reading
Official sources for verification and further reading

