MENA Open Data & Evidence Lab
Research Standards Handbook

The method includes the evidence trail.

The Lab treats provenance, versioning, correction and independent review as part of research design. A statistically sophisticated result with an unreconstructable source trail is not a high-standard release.

01

Source hierarchy

Prefer authoritative primary publishers and official release artifacts. Secondary aggregators can assist discovery or cross-checking but do not silently replace the canonical source when the primary record is available.

02

Data provenance

Material values preserve publisher, source URL or artifact, retrieval time, definition, reference period, release date or vintage, transformation record and source hash where practicable.

Full standard →
03

Replication standard

A replication record identifies the original claim, dataset and specification; documents deviations; reruns the analysis; records robustness checks, environment, code and results hash; and separates reproduced from not reproduced.

Validation standard →
04

Version control

Draft, release candidate and released states are explicit. Material changes in rows, source perimeter, transformations, methods or corrections require a new recorded state rather than a silent replacement.

Release standard →
05

Statistical review

Inferential work must state the estimand, assumptions, uncertainty treatment, clustering or dependence choices where relevant, robustness perimeter and material failure modes. Internal checks are not relabelled external review.

06

Citation standards

Every released object receives a stable suggested citation with author or institutional author, year, title, version and persistent identifier when one exists. Underlying official sources remain cited at claim level.

07

Correction policy

Confirmed errors are logged with affected version, severity, correction date and replacement state. Corrections remain inspectable rather than disappearing with an overwritten file.

Corrections →
08

Synthetic-data policy

Synthetic, simulated or randomized-synthetic evidence is labelled as such at the claim level. Synthetic benchmarks can test methods; they are not presented as observed client or population outcomes.

09

External review

Independent verification requires a different identifiable person to complete the defined source audit, second-coding or methods task and leave a reviewable record. Informal praise or discussion is not review.

10

Conflict of interest

Commissioning, funding, collaboration and review relationships are disclosed at the level relevant to the research object. A funder or client does not receive authority to change an evidence label.

Funding & independence →
11

Evidence classifications

Important claims are classified according to their actual evidentiary state: official source, real public data, provider test, synthetic, randomized synthetic, production client data, external review, independent reproduction, founder produced or pending validation.

Classification register →
12

Known limitations

Coverage gaps, unresolved source conflicts, sampling limits, missing second coding, unavailable vintages and untested assumptions remain part of the public record. A limitation is not removed because it weakens presentation.

Definition

INDEPENDENTLY REPRODUCED

A release may use this designation only when a person or team outside the primary production process reconstructs the relevant data or analysis from a clean environment, records required deviations and discrepancies, and leaves an attributable result that can be reviewed. Founder reruns, automated CI, internal code checks and second runs by the same producer remain internal QA.

Methods pipeline

Pipeline topics are research directions, not publications. They receive a publication identifier only once an actual manuscript, draft or citable methods-note record exists.

Pipeline

Statistical vintages

Reconstructing what was publicly knowable from official releases at a historical point in time.

Pipeline

PDF-origin provenance

Source and transformation traceability for official evidence published outside clean machine-readable tables.

Pipeline

Independent second coding

Separating primary extraction from independent evidence coding, reconciliation and agreement measurement.

View research agenda →

MENA Replication Series — MEL-REP

Required fields are: original claim, original dataset, original specification, reproduction result, required deviations, robustness checks, updated sample where applicable, limitations, code, environment and results hash. Replication records are methodical research objects, not attack pieces.

Machine-readable methods pipeline · Validation dashboard · Research governance.