Historical pages, PDFs, language pairs, discovery and original-vintage survival verified. Machine retrievability is partial because filename semantics and direct-file access are not uniformly robust.
Statistical Memory Loss
Can an analyst reconstruct what an official statistical agency actually published two or three years ago? Release 004 audits the survival of historical release pages, underlying files, bilingual pairs, archive discovery paths and machine retrievability across Jordan, Morocco, Saudi Arabia and Egypt.
8 source cases4 statistical systemssurvivability auditverified memory dimensions for Jordan and Morocco in the pilot sample; Saudi Arabia and Egypt retain substantial history but expose more machine-access friction.
Core finding
History can survive and still become harder to reproduce.
The audit found four distinct failure modes: misleading filenames, current databases that re-express history under new frameworks, opaque attachment identifiers, and migration from legacy archives to interfaces that are less machine-visible.
Forensic cases
Observed survivability and reproducibility issues from official-source archives.
| System | Historical evidence | Observed issue | Why it matters |
|---|---|---|---|
| Jordan · CPI Jun 2024 | English release page and PDF survive | Surviving English PDF URL is CPI_Jul_e.pdf although the document is the June 2024 CPI release | Automated collectors that infer reference period from filename alone can misclassify the observation. |
| Jordan · CPI 2023/2024 | Arabic and English release history plus PDFs remain public | Storage filenames encode publication-month conventions rather than uniquely encoding reference period | Archive preservation is strong, but provenance requires reading document metadata rather than trusting paths. |
| Morocco · GDP 2023/2024 | Old HCP pages survive with French and Arabic attachments | Attachments are served through opaque numeric identifiers while descriptive labels remain on the page | A release page is part of the provenance record; copying the attachment URL alone loses semantic context. |
| Saudi Arabia · CPI history | Historical monthly publications remain indexed | Current CPI methodology republishes historical series from 2013 under the 2023 reference framework while older period-specific publication objects also survive | “History as currently expressed” and “history as originally published” are different research objects. |
| Egypt · CPI archive | Legacy CAPMAS archive exposes years 2000–2025 and metadata records | New publication route is more JavaScript-dependent and less machine-visible in this audit | Website modernization can preserve human access while degrading automated reproducibility. |
Statistical Memory Score v0.1
Six dimensions: historical release page, underlying document, language pair, archive discovery, original-vintage survival, machine retrievability.
Old release pages and bilingual attachments survive alongside later revised vintages. Machine use remains partial because attachment identities are less self-describing than page context.
Historical publication objects survive, but direct file access, language parity and machine retrieval need a deeper byte-level audit. Current CPI history is also re-expressed under the new framework.
Legacy archive discovery is strong, but file verification, language parity, original-vintage survival and machine retrievability remain only partial in this first audit.
Scores summarize verified evidence coverage in this release; they are not permanent institutional ratings.
What this changes for empirical research
Never identify a vintage by value alone.
A revised current database can reproduce the same reference period with a different historical value.
Never identify a period from filename alone.
The Jordan June-2024 CPI case demonstrates why file paths are not reliable substitutes for document metadata.
Preserve the release page with the file.
Language, publication date, attachment label and contextual notes can disappear when researchers save only the binary document.
Hash now, compare later.
Release 004 establishes the archive schema for future SHA-256 checks. Byte-level hashing was not successfully completed in the current audit environment, so no false hash claims are published.
Audit rule
A crawler timeout is not coded as a dead link. A JavaScript-heavy page is not coded as missing. A file whose bytes could not be retrieved is not called deleted. The release records the narrower fact that was observed.
Next reproducibility layer
The next archive pass should run a scheduled byte-level fetch of a frozen file panel, record HTTP status, content length, ETag/Last-Modified when present, SHA-256, redirect chain and document title, and then compare those fields over time. That is the point at which silent replacement becomes directly measurable.
Suggested citation: MENA Open Data & Evidence Lab (2026). Research Release 004: Statistical Memory Loss. Version 1.0, 17 August 2026.