Skip to main content
Thirdwatchthirdwatch
Business & local data

Build a Historical FX Rate Dataset for Backtests

Export daily exchange-rate history for any currency pair as structured JSON, then shape it into a clean time series for backtests and reports.

Sep 21, 2026 · 3 min read · 693 words
See the scraper →

TL;DR — The Currency Rates Scraper exports daily exchange-rate history for 340+ currencies as JSON: pair, date, rate, inverse rate, converted amount, provider. Set a date range and currency lists, then pivot the output into a backtest-ready time series with pandas. Built for developers and analysts.

Why build a historical FX dataset

A backtest is only as honest as its exchange rates. Models that convert foreign revenue at today's rate are silently assuming FX never moves — and FX always moves. The BIS triennial survey counts over $7.5 trillion in daily FX turnover; even a few percent of drift rewrites a multi-currency P&L.

The job-to-be-done is a clean, reproducible rate table: one row per pair per day, a named source per row, and no hand-downloaded CSVs that go stale the day you save them.

How does this compare to the alternatives?

Manual CSV downloads Terminal data feeds Thirdwatch actor
Cost Free but tedious Expensive subscription Pay per result
Reliability Stale the day you save it Excellent Free providers, named per row
Setup time Per query, per provider Weeks of onboarding Minutes
Maintenance Re-download to refresh Vendor-managed Re-run to refresh

Terminals are the right answer for trading desks. For backtests, reporting, and model features, daily fixes from free keyless sources are usually enough — the actor makes them scriptable.

How to build the dataset in 4 steps

How do I pull a date range?

Set startDate and endDate and list your pairs through the currency arrays:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/currency-rates-scraper").call(
    run_input={
        "baseCurrencies": ["USD"],
        "targetCurrencies": ["EUR", "GBP", "INR", "JPY"],
        "startDate": "2026-01-01",
        "endDate": "2026-08-31",
        "provider": "auto",
        "maxResults": 250,
    }
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())

One run covers the cross-product of your lists across the whole range — no per-day loop needed.

How do I shape rows into a series?

The records are already tidy: one row per pair per date. Pivot to wide format:

import pandas as pd

df = pd.DataFrame(rows)
df["date"] = pd.to_datetime(df["date"])
rates = df.pivot_table(index="date", columns="pair", values="rate").sort_index()
rates = rates.ffill()  # bridge weekend/holiday gaps deliberately
print(rates.describe())

ffill is explicit: you decide to carry the last published fix forward rather than discovering silent gaps mid-backtest.

How do I sanity-check the data?

Two cheap checks catch most problems. Verify date coverage per pair, and compare providers where the same day was served twice:

gaps = rates.isna().sum()
print("Missing days per pair:\n", gaps)
prov = df.groupby(["pair", "provider"]).size()
print(prov)

If one provider dominates a date range while another covered the rest, look at the seam dates — provider switches can introduce small level shifts.

How do I keep the dataset fresh?

Rerun the actor with an startDate equal to your last stored date plus one day, and append:

last = rates.index.max()
run = client.actor("thirdwatch/currency-rates-scraper").call(
    run_input={
        "baseCurrencies": ["USD"],
        "targetCurrencies": ["EUR", "GBP", "INR", "JPY"],
        "startDate": (last + pd.Timedelta(days=1)).date().isoformat(),
        "endDate": pd.Timestamp.today().date().isoformat(),
        "maxResults": 250,
    }
)

Scheduled incremental runs keep the dataset current without re-pulling history.

Sample output

{"base_currency": "USD", "quote_currency": "JPY", "pair": "USD/JPY",
 "date": "2026-09-11", "rate": 147.21, "inverse_rate": 0.0067933,
 "amount": 1, "converted_amount": 147.21,
 "provider": "erapi", "provider_name": "ExchangeRate-API open access"}

Every row carries its provider and provider_name, so when two sources disagree you can attribute the difference instead of averaging it away.

Common pitfalls

Free sources fix once per day; a dataset built on them is daily-grain by construction. Exotic and crypto pairs have thinner coverage than majors — check gaps before assuming a series is complete. Historical pulls can be long, so raise maxResults to cover pairs × days or you will silently truncate the range. The actor records the provider on every row so coverage and source shifts stay visible in the data itself.

Related use cases

Frequently asked questions

How far back does the history go?

The actor's free sources cover years of daily fixes for major fiat pairs. Coverage depth varies by provider and currency; major pairs go back furthest. Request a range and check the returned date coverage before building on it.

What grain does the data come at?

One record per base-target pair per day. That daily grain suits reporting, hedging analysis, and strategy backtests that do not need intraday bars.

Can I pull several pairs in one run?

Yes. baseCurrencies and targetCurrencies are arrays, so a single run produces the full cross-product of pairs across the date range you set.

How do I avoid gaps in the series?

Weekends and holidays return the nearest published fix. Forward-fill in pandas if your model needs a row for every calendar date, and record the provider field so gaps are attributable.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.