---
okf_version: "0.2"
---

# COVID-19 Records

An OKF bundle over 5,247 pages of records released by the U.S. Senate
Committee on Homeland Security and Governmental Affairs — documents the
Committee subpoenaed and then published — segmented into 18,094 individual
communications spanning 2000-02-05 to 2026-04-10.

## What this bundle is, and is not

**This is a derived layer, not the records.** It carries the analysis over the
corpus: people, topics, organisations, outlets, sources, a chronology, and
transcripts for the larger conversations. It does **not** carry the text of
the 17,260 individual records, and it holds transcripts for only
301 of 2,021 conversations.

That distinction matters more than it sounds. **Searching these files is not
searching the corpus.** A grep across the bundle returns a fraction of the
true hits and nothing here marks it as a fraction, so an agent reading only
these files can produce a confident, well-cited, incomplete answer. For
anything that needs completeness, use the API — see
[Documents](documents/index.md).

## Collections

* [Chronology](chronology/) - the record by date; the only time-based entry point
* [People](people/) - 253 individuals, with correspondents, topics and statement timelines
* [Topics](topics/) - 22 subject categories, each with its principal voices
* [Organizations](organizations/) - 81 institutions
* [Outlets](outlets/) - 54 media outlets and press-direction evidence
* [Conversations](conversations/) - 301 chat sessions and email threads
* [Sources](sources/) - 14 released document packages
* [Documents](documents/) - how to retrieve the records themselves
* [Glossary](glossary.md) - acronyms, institutional shorthand and strain names

## How to use this bundle

Start from a [person](people/), a [topic](topics/) or a [date](chronology/)
and follow the links. Person concepts carry a chronological statement timeline
answering *what did this individual say, and when*. Topic concepts carry the
same organised by subject.

Then verify. Every quoted statement cites a page, and page text is retrievable
at `https://randpaulcovid.org/api/page/<source>/<page>`. The authoritative copy is the
Committee's own PDF, linked from every [source](sources/) concept.

## What the fields mean

`grade`, citations, chat audience, topic tags and the limits of coverage each
carry a specific meaning that a reader taking the numbers at face value will
get wrong. They are defined once, in machine-readable form, at
`https://randpaulcovid.org/api/guidance` — not restated here, because an agent reading `/llms.txt`,
this file and that endpoint would otherwise meet the same paragraph three
times before reaching a record.

## Provenance

Generated by `randpaul-corpus-pipeline/1.0` at 2026-08-01T12:42:27Z from the released PDF packages, which are
public and hosted by the Committee at <https://www.paul.senate.gov/readingroom/>.
This bundle adds no source documents. Text was recovered by OCR where the
productions were scanned images. Extraction is rule-based and deterministic:
re-running the pipeline reproduces this bundle.

Statement timelines are machine-extracted and carry `confidence` below 1.0 in
the database; they are evidence pointers, not adjudicated findings, and should
be read alongside the cited page.
