← Atlas

Editorial standard

Methodology

The Atlas is built as a continuously updated scoping review of grey-literature company data. The pipeline composes three established reporting standards so every record can be traced, audited, and re-verified.

§ 01

PRITES — dataset-level discipline

We adopt PRITES (Huang & Lam, 2025, Monash Addiction Research Centre) as the governing frame for the web-scraped dataset: Provenance, Representativeness, Integrity, Transparency, Ethics, Sustainability.

  • Provenance — every record stores its retrieval URL, fetch time, HTTP status, and page title.
  • Representativeness — search queries are run multilingually across 18 geographic regions to avoid over-indexing on US/UK English sources.
  • Integrity — every cited URL is re-fetched and required to actually mention the company.
  • Transparency — taxonomy, scoring rubric, and source-tier rules are public on this page.
  • Ethics — only public pages, identified bot UA, no scraping behind logins or paywalls.
  • Sustainability — the validator re-runs on a schedule; flagged sources never silently disappear.

Huang, C. A., & Lam, T. (2025). PRITES: An integrative framework for investigating and assessing web-scraped HTTP-response datasets for research applications. arXiv:2511.13773.

§ 02

PRISMA-ScR — the discovery flow

Company discovery follows the four PRISMA-ScR stages, instrumented as pipeline counters:

  1. Identification — candidate names emerge from query banks across regulatory registers (FDA, EMA, BfArM DiGA, MHRA), trial registries (ClinicalTrials.gov, EU CTR, WHO ICTRP), funding aggregators, app-store searches, and topical news.
  2. Screening — the research agent retrieves each candidate and rejects out-of-scope names with a structured NOT_RELEVANT verdict.
  3. Eligibility — surviving records must have a defensible substance / behavioural / recovery focus; pure mental-health and adjacent wellness apps are excluded.
  4. Inclusion — the record is written with a confidence score and an immediate source-validation pass.

Tricco, A. C., et al. (2018). PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, 169(7), 467–473.

§ 03

AACODS — per-source appraisal

Grey-literature items (company sites, press releases, aggregator entries) are appraised per source against the AACODS checklist — Authority, Accuracy, Coverage, Objectivity, Date, Significance. The validator currently operationalises three axes automatically; the others remain editorial.

AACODS axisHow we operationalise it
AuthorityDomain classified into a tier: regulator · trial registry · peer-reviewed · primary · news · aggregator · unknown.
AccuracyPage must load (HTTP 2xx) and must literally mention the company name or all of its significant name tokens.
CoverageEach record needs ≥1 verified source from a regulator, trial registry, peer-reviewed venue, or the company's own primary domain.
ObjectivityEditorial review of marketing-dominated sources; flagged when only PR/aggregator sources remain.
Datelast_checked_at is stored per record; sources are re-validated continuously.
SignificanceMapped to the per-company evidence and monitoring-priority scores.

Tyndall, J. (2010). AACODS Checklist. Flinders University.

§ 04

Source authority tiers

The Atlas treats sources as a hierarchy. A green ✓ on a record is meaningful only when paired with the tier of the verifying source:

  • Regulator — FDA, EMA, MHRA, BfArM, PMDA, NICE, SAMHSA, NIH, WHO, gov.uk, ec.europa.eu.
  • Trial registry — ClinicalTrials.gov, EU CTR, ISRCTN, ANZCTR, DRKS, WHO ICTRP.
  • Peer-reviewed — PubMed, OpenAlex, NEJM, Lancet, BMJ, JMIR, Nature, PLOS, Frontiers, CORDIS, Cochrane.
  • Primary — the company's own domain (verified host match).
  • News — recognised trade and financial press.
  • Aggregator — Crunchbase, Pitchbook, Dealroom, LinkedIn, Wikipedia — useful for triangulation, never sufficient alone.

§ 05

What gets flagged, automatically

  • URL returns HTTP ≥ 400, times out, or fails TLS.
  • URL loads but does not mention the company or any alias.
  • A record's only verified sources are aggregators or unknown-tier pages.

Flagged items are not deleted — they are displayed inline on the company profile with the failure reason, so reviewers can override or replace them.

§ 06

Limits — what this method does not do

  • It does not verify factual claims inside a page; only that the page exists and is about the right company.
  • It does not score editorial bias.
  • Funding figures, regulatory status, and trial counts still depend on the AI research step; they should be cross-checked before citation.
Methodology v0.1 · Last revised 2026-08-26.