On 31 January 2025, the Centers for Disease Control and Prevention began taking down pages and data tools that contained terms flagged in a 29 January memo from the U.S. Office of Personnel Management. Within days, the removals had reached the Food and Drug Administration, the National Institutes of Health, and other parts of the Department of Health and Human Services. By the first week of February, researchers tracking federal health web content counted more than 8,000 pages that had gone missing. A data portal that a graduate student used on a Tuesday might be gone by Wednesday afternoon.
Nobody tells you that federal health data can vanish faster than a grant deadline sneaks up on you. I learned this the hard way. Years ago, I assumed a government survey file would stay at the same URL forever. My backup strategy was hope. Hope is not a preservation plan, and this is the year a lot of people found that out the same way I did.
What followed is a research data availability crisis. The phrase is not a synonym for a broken website. It describes a situation where a dataset exists, or existed, but the version a researcher needs is no longer reliably accessible, citable, or comparable to the version used in earlier analysis.
The removals began with a memo and ended up in a federal courtroom
The 29 January 2025 memorandum from the Office of Personnel Management directed agencies to remove materials related to what it called gender ideology from public-facing systems. CDC staff were told to pull pages containing words such as gender, transgender, equity, inclusion, and others. A federal judge stepped in on 11 February 2025, ordering CDC and FDA to restore pages removed under the directive. The ruling did not restore everything. Some pages reappeared with changed wording; some data files did not reappear at all.
What made this different from an ordinary website redesign was the scale and the absence of notice. Federal statistical products are supposed to be prepared and released under policies that protect them from political interference. Researchers plan studies, write grant applications, and calibrate models around the expectation that 2021 Youth Risk Behavior Surveillance System files will still be available when 2025 analyses begin. That expectation failed quickly and publicly.
A data availability crisis means version control, not just broken links
Think of a government dataset like a family recipe card. If someone retypes the card and changes one ingredient but keeps the same title, the casserole tastes different and nobody can say exactly why. The same thing happens when a public health survey file is removed, edited, and reposted without a clear version history. Researchers can end up comparing estimates that were not built on the same denominator or the same wording.
The problem is not that federal agencies never update datasets. They do, and good versioning is part of their job. The problem is that the removals were fast, the documentation was thin, and the original URLs often stopped resolving. For a university research team, that is not an inconvenience. It is a threat to replication, to peer review, and to the integrity of longitudinal public health research.
The datasets under the sharpest strain
Among the tools researchers flagged in the first weeks of February were the CDC Youth Risk Behavior Surveillance System, known as YRBSS, which tracks health risk behaviours among adolescents; the Social Vulnerability Index, used widely in disaster preparedness and health equity research; and the AtlasPlus portal for HIV, viral hepatitis, tuberculosis, and sexually transmitted disease surveillance data. These are not obscure files. YRBSS alone underpins a large share of published work on adolescent mental health and substance use in the United States.
- The YRBSS files and query tools that school health researchers use to track trends in depression, vaping, and sexual health behaviour.
- The Social Vulnerability Index, which combines census variables into a single measure used by emergency planners and university researchers studying disaster outcomes.
- The AtlasPlus portal, where state and local health departments build epidemic curves and compare rates across counties.
- Health disparities and social determinants pages that provided documentation for variables used in funded studies.
A page returning with new language is not the same as a dataset returning with its original variable dictionary. Even when a file came back, researchers had to check whether the variable names, coding, and questionnaire text matched the version they had cited in a manuscript or grant.
University data librarians became the backup system by default
Within days, data librarians and research data organisations began treating the removals as an archiving emergency. The Data Rescue Project, a community effort coordinated by data professionals, publishes guidance and gathers copies of at-risk federal datasets. The Data Rescue Project was one of the quickest paths into that work. Researchers also turned to DataLumos, an open archive built for public-sector data, and to Harvard Dataverse for versioned deposits. The Internet Archive's End of Term Web Archive had been crawling federal sites at administration changes since 2008, and its snapshots suddenly became working infrastructure for scientific papers.
University libraries now face a reversal of roles. Instead of pointing researchers to government portals, they are storing the copies researchers need. That shift is good, but it places libraries in a position they cannot sustain on their own. Archival copies are only as useful as the metadata around them: who captured the file, when, from which URL, and with what checksum.
For publication and grant compliance, missing source data is now a workflow risk
The funding side of this crisis is not separate from the data side. When the National Institutes of Health requires a data management and sharing plan, it assumes the underlying federal data will remain available. If a CDC dataset disappears mid-project, a principal investigator may be unable to satisfy the sharing requirement without archiving a copy that carries its own provenance problems. This connects directly to the funding instability covered in our earlier reporting on NIH grant terminations and to the long-running open access changes that the NIH and EU data sharing mandates set in motion.
Journal reviewers are now left asking a new question: where is the source data? A manuscript may cite a CDC URL that no longer resolves. An archived copy can answer the question, but only if the authors documented the version. That documentation is not yet standard practice in most university departments. Editors have not settled on a universal citation format for datasets that experienced a federal purge.
Practical steps for research teams and department heads
This is the point where my old instinct was to make a checklist and call it a policy. Checklists help, but only if they solve the actual workflow problem. The workflow problem here is simple: researchers treat source URLs as permanent, and they are not permanent. The steps below are the minimum that a research group should do before starting analysis on federal health data.
- Download the dataset and keep the original file, the variable dictionary, and a screenshot of the access page on the day you retrieved it.
- Record the URL, the date of access, and a checksum for every file you plan to use in publication.
- Deposit a versioned copy in an institutional repository or a community archive such as DataLumos, even if the source file is still online.
- If a dataset disappears mid-project, tell your program officer and journal editor at the same time, with an archived copy attached.
- Add a sentence to your data management plan naming an archival fallback for each federal dataset the project depends on.
This is not glamorous work. It is the academic equivalent of checking your smoke alarm batteries. But it is what turns a data purge from a publication-ending event into a solvable citation problem.
Redundancy is the only form of trust a dataset has
The court order in February did not settle the long-term question. Federal agencies may restore more pages, but the episode has already changed how university researchers budget their trust. A dataset now needs at least two homes before it should be considered stable: the official portal and an independent archive with clear provenance.
I used to think data preservation was the job of the agency that produced the data. I was wrong in the comfortable way people are wrong about things that have never broken before. The researchers and librarians who have spent the past year building mirrors are not hoarding. They are doing the unglamorous maintenance that makes science reproducible. Hope never was a preservation plan; redundancy, it turns out, actually is.
Photo by Nisuda Nirmantha on Unsplash
