Complete Regulatory Archives, Delivered Daily to Client Infrastructure
An international law firm needed dependable, citation-grade archives of enforcement decisions and sanctions from official sources in several jurisdictions. We built the pipeline that collects, verifies and delivers them every day.
Results at a Glance
The Client
An international law firm with regulatory and disputes work spanning several jurisdictions. Its lawyers and analysts depend on complete, current archives of decisions published by regulators, courts and official registers. Client identity withheld under NDA.
The Challenge
Official publication portals are built for reading one document at a time, not for maintaining a complete archive. For legal work that gap matters: an archive that silently misses documents is worse than no archive at all, because it invites false confidence.
- Gated search interfaces, session-bound downloads and no bulk export on most official portals
- Citation-grade completeness required: every published decision, not most of them
- Some sources publish scanned documents with no text layer
- Documents are occasionally amended or withdrawn after publication, and the archive has to reflect that accurately
- The firm's research interests are themselves confidential, so collection had to run on dedicated infrastructure with private delivery
- Counsel must be able to explain exactly how every document was obtained and processed
Our Solution
We built a dedicated collection pipeline for each official source, with verification and provenance built in from the start.
- One collector per source: each official portal gets its own collector built around how that publisher actually works, and kept current as portals change
- Deterministic parsing, no AI: the same input produces the same output on every run, and every extraction rule is inspectable code. The archive is defensible in a way probabilistic tools cannot be
- Completeness reconciliation: every run compares collected totals against the counts the source itself publishes, flagging any gap rather than hiding it
- OCR fallback: scanned documents pass through a text-extraction stage so the whole archive stays searchable
- Withdrawal detection: documents removed or replaced at source are detected and flagged, so the archive reflects what actually stands
- Delivery to client-owned storage: documents and per-document provenance manifests sync daily to cloud storage the client controls, under the client's own keys
- Hardened, audited infrastructure: collection runs on a dedicated host with restricted outbound network access and credentials held outside the application and injected at runtime
How the Engagement Ran
- Historical backfill: full archive reconstruction from each source's earliest available records, reconciled against published totals before sign-off
- Continuous collection: daily scheduled runs with completeness checks on every run, so the archive never drifts out of date
- Incremental expansion: new sources added as the engagement grew, each brought up to the same reconciliation and provenance standard before joining the daily cycle
The client's security team reviewed the infrastructure directly, including an independent external vulnerability scan of the collection host.
Results
The firm now holds complete, continuously updated archives it could not obtain from any commercial database, on infrastructure it controls, with a collection method its lawyers can explain to a court or regulator.
- Complete archives from five official sources, maintained daily in production
- Historical reconstructions running to tens of thousands of documents, each reconciled against the totals the source itself publishes
- Every document traceable to its source through a provenance manifest recording where it came from, when it was collected, and what was done to it
- Scanned material fully text-searchable through the OCR stage
- Collection infrastructure reviewed by the client's security team, including an independent external vulnerability scan
Citation-Grade Data, Built for Legal Work
See how we collect regulatory data, or explore the open CMA dataset we publish as a working demonstration of the method.