Build a Historical FX Rate Dataset for Backtests
Export daily exchange-rate history for any currency pair as structured JSON, then shape it into a clean time series for backtests and reports.

TL;DR — The Currency Rates Scraper exports daily exchange-rate history for 340+ currencies as JSON: pair, date, rate, inverse rate, converted amount, provider. Set a date range and currency lists, then pivot the output into a backtest-ready time series with pandas. Built for developers and analysts.
Why build a historical FX dataset
A backtest is only as honest as its exchange rates. Models that convert foreign revenue at today's rate are silently assuming FX never moves — and FX always moves. The BIS triennial survey counts over $7.5 trillion in daily FX turnover; even a few percent of drift rewrites a multi-currency P&L.
The job-to-be-done is a clean, reproducible rate table: one row per pair per day, a named source per row, and no hand-downloaded CSVs that go stale the day you save them.
How does this compare to the alternatives?
| Manual CSV downloads | Terminal data feeds | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free but tedious | Expensive subscription | Pay per result |
| Reliability | Stale the day you save it | Excellent | Free providers, named per row |
| Setup time | Per query, per provider | Weeks of onboarding | Minutes |
| Maintenance | Re-download to refresh | Vendor-managed | Re-run to refresh |
Terminals are the right answer for trading desks. For backtests, reporting, and model features, daily fixes from free keyless sources are usually enough — the actor makes them scriptable.
How to build the dataset in 4 steps
How do I pull a date range?
Set startDate and endDate and list your pairs through the currency arrays:
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/currency-rates-scraper").call(
run_input={
"baseCurrencies": ["USD"],
"targetCurrencies": ["EUR", "GBP", "INR", "JPY"],
"startDate": "2026-01-01",
"endDate": "2026-08-31",
"provider": "auto",
"maxResults": 250,
}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())One run covers the cross-product of your lists across the whole range — no per-day loop needed.
How do I shape rows into a series?
The records are already tidy: one row per pair per date. Pivot to wide format:
import pandas as pd
df = pd.DataFrame(rows)
df["date"] = pd.to_datetime(df["date"])
rates = df.pivot_table(index="date", columns="pair", values="rate").sort_index()
rates = rates.ffill() # bridge weekend/holiday gaps deliberately
print(rates.describe())ffill is explicit: you decide to carry the last published fix forward rather than discovering silent gaps mid-backtest.
How do I sanity-check the data?
Two cheap checks catch most problems. Verify date coverage per pair, and compare providers where the same day was served twice:
gaps = rates.isna().sum()
print("Missing days per pair:\n", gaps)
prov = df.groupby(["pair", "provider"]).size()
print(prov)If one provider dominates a date range while another covered the rest, look at the seam dates — provider switches can introduce small level shifts.
How do I keep the dataset fresh?
Rerun the actor with an startDate equal to your last stored date plus one day, and append:
last = rates.index.max()
run = client.actor("thirdwatch/currency-rates-scraper").call(
run_input={
"baseCurrencies": ["USD"],
"targetCurrencies": ["EUR", "GBP", "INR", "JPY"],
"startDate": (last + pd.Timedelta(days=1)).date().isoformat(),
"endDate": pd.Timestamp.today().date().isoformat(),
"maxResults": 250,
}
)Scheduled incremental runs keep the dataset current without re-pulling history.
Sample output
{"base_currency": "USD", "quote_currency": "JPY", "pair": "USD/JPY",
"date": "2026-09-11", "rate": 147.21, "inverse_rate": 0.0067933,
"amount": 1, "converted_amount": 147.21,
"provider": "erapi", "provider_name": "ExchangeRate-API open access"}Every row carries its provider and provider_name, so when two sources disagree you can attribute the difference instead of averaging it away.
Common pitfalls
Free sources fix once per day; a dataset built on them is daily-grain by construction. Exotic and crypto pairs have thinner coverage than majors — check gaps before assuming a series is complete. Historical pulls can be long, so raise maxResults to cover pairs × days or you will silently truncate the range. The actor records the provider on every row so coverage and source shifts stay visible in the data itself.
Related use cases
- Analyze crypto market data with Python — same tidy-series workflow on crypto pairs.
- Build a crypto price monitor with CoinGecko — live monitoring counterpart.
- Guide to scraping business data.
- More on the blog hub.
Frequently asked questions
How far back does the history go?
The actor's free sources cover years of daily fixes for major fiat pairs. Coverage depth varies by provider and currency; major pairs go back furthest. Request a range and check the returned date coverage before building on it.
What grain does the data come at?
One record per base-target pair per day. That daily grain suits reporting, hedging analysis, and strategy backtests that do not need intraday bars.
Can I pull several pairs in one run?
Yes. baseCurrencies and targetCurrencies are arrays, so a single run produces the full cross-product of pairs across the date range you set.
How do I avoid gaps in the series?
Weekends and holidays return the nearest published fix. Forward-fill in pandas if your model needs a row for every calendar date, and record the provider field so gaps are attributable.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.