Skip to main content
Thirdwatchthirdwatch
Engineering

Analyze ClinicalTrials.gov Data With Python

Use pandas to analyze clinical trial phases, statuses, sponsors, enrollment, interventions, and locations without mixing study- and site-level counts.

Jul 21, 2026 · 1 min read · 95 words
See the scraper →

Load the JSON export from the ClinicalTrials.gov Scraper and establish the study grain first:

import pandas as pd

studies = pd.read_json("dataset_items.json").drop_duplicates("nct_id")
print(studies["overall_status"].value_counts(dropna=False))
sites = studies[["nct_id", "locations"]].explode("locations")
sites = pd.concat(
    [sites.drop(columns="locations"), sites.locations.apply(pd.Series)], axis=1
)

Analyze phases, status, sponsor class, and enrollment in the study table. Analyze countries and facilities in the site table. Label missing fields and define the snapshot date on every chart.

Registry data supports descriptive analysis; it does not prove efficacy or predict approval. Preserve NCT ID and study URL so every aggregate can be checked against its source.

Frequently asked questions

Why keep a study table and a location table?

Exploding locations creates multiple rows per study and otherwise inflates study counts.

Should enrollment be summed across snapshots?

No. Deduplicate by NCT ID within a snapshot and treat revisions as changing values.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.