Data catalog
The data catalog stores NautilusTrader data in Parquet files for backtesting, live trading, and research.
Overview and architecture
ParquetDataCatalog is the Python interface to the Rust catalog and DataFusion query engine.
The Rust model and persistence crates define the Arrow schemas for built-in data. Registered custom
data supplies its schema and encode/decode handlers at runtime.
Instant timestamps use Timestamp(Nanosecond, Some("UTC")); durations remain integers. Arrow readers
and Nautilus queries preserve the nanoseconds and UTC annotation. SQL readers that map these columns
to microsecond-precision TIMESTAMPTZ can truncate sub-microsecond values.
Parquet provides compressed columnar storage and cross-language access. The catalog stores these files under one root without requiring a separate database service. A local path or object-store URI selects the storage backend.
Initializing
Pass a local path or URI as the first constructor argument:
from pathlib import Path
from nautilus_trader.persistence import ParquetDataCatalog
CATALOG_PATH = Path.cwd() / "catalog"
catalog = ParquetDataCatalog(str(CATALOG_PATH))Filesystem protocols and storage options
The catalog accepts the storage protocols supported by its Rust object-store backend.
Supported filesystem protocols
| Storage | URI schemes | Common option keys |
|---|---|---|
| Local filesystem | Plain path, file | None. |
| Amazon S3 | s3 | region, access_key_id, secret_access_key, endpoint_url. |
| Google Cloud Storage | gs, gcs | service_account_path, service_account_key, application_credentials. |
| Azure Blob Storage | az, abfs | account_name, account_key, sas_token. |
| HTTP or WebDAV | http, https | timeout, connect_timeout, proxy_url. |
Option keys are object_store configuration keys for the URI scheme:
- S3, GCS, and Azure keys also take a prefixed form, such as
aws_region,google_application_credentials, orazure_storage_sas_token. - Client options such as
timeoutapply to every remote scheme. Setallow_httptotrueto connect over plainhttp://. - S3 also accepts the legacy
keyandsecretnames. - An unknown key fails with an error that names it.
Pass credentials and other backend settings through storage_options:
catalog = ParquetDataCatalog(
"s3://my-bucket/nautilus-data/",
storage_options={
"access_key_id": "your-key",
"secret_access_key": "your-secret",
"region": "us-east-1",
},
)
azure_catalog = ParquetDataCatalog(
"abfs://container@account.dfs.core.windows.net/nautilus-data/",
storage_options={"account_key": "your-account-key"},
)The base path of a remote URI, after the bucket, container, or host, must not contain spaces,
non-ASCII characters, or characters that object-store paths encode, such as ~, %, #, ?,
^, or |. The catalog rejects such a URI when it opens, with an error that names it.
Compression and row groups
DataCatalogConfig sets how a configured catalog reads and writes data files:
| Field | Controls | Default |
|---|---|---|
batch_size | Rows per batch the catalog reads and writes | 10,000 |
compression | Codec name for written data files | zstd (level 1) |
max_row_group_size | Maximum rows per written row group | 131,072 |
from nautilus_trader.config import DataCatalogConfig
catalog = DataCatalogConfig(path="./catalog", compression="snappy", max_row_group_size=65_536)compression accepts uncompressed, snappy, gzip, brotli, lz4, lz4_raw, or zstd,
ignoring case. Construction fails for any other name and for a zero batch_size or
max_row_group_size. params carries only options for an external catalog backend, and the
Parquet catalog rejects any params key.
ParquetDataCatalog takes the same settings as batch_size, max_row_group_size, and a numeric
compression code: 0 (uncompressed), 1 (Snappy), 2 (gzip), 4 (Brotli), 5 (LZ4), or 6
(zstd).
lz4 and lz4_raw name the same codec: the catalog writes Parquet LZ4_RAW. The Parquet format
deprecates the older Hadoop-framed LZ4 codec, so the catalog never writes it, and compression code
5 writes LZ4_RAW even though Parquet numbers LZ4_RAW as 7. Files written earlier with the
Hadoop-framed codec remain readable. LZO is not supported because the Parquet writer cannot produce
it, so the name lzo and code 3 fail at construction.
Writing data
Use the writer for the concrete data type. Instrument definitions and custom data have separate writers.
catalog.write_instruments([instrument])
catalog.write_quote_ticks(quote_ticks)
catalog.write_trade_ticks(
trade_ticks,
start=1704067200000000000,
end=1704153600000000000,
)
catalog.write_bars(bars, skip_disjoint_check=True)The built-in market-data writers are:
write_quote_tickswrite_trade_tickswrite_order_book_deltaswrite_order_book_depthswrite_barswrite_mark_price_updateswrite_index_price_updateswrite_option_greeks
Each writer accepts optional start and end overrides as UNIX nanoseconds. The data in one call
must have one identity, such as one instrument ID or bar type, and must be ordered by ts_init.
File naming and data organization
The catalog names files from their timestamp range with the pattern
{start_timestamp}_{end_timestamp}.parquet. It converts each ISO 8601 timestamp to a filename-safe
form by replacing : and . with -.
Built-in data is organized in directories by data type and identifier. For instrument IDs and bar
types, the catalog removes / and replaces ^ with _ when creating the URI-safe directory name:
catalog/
├── data/
│ ├── quotes/
│ │ └── EURUSD.SIM/
│ │ └── 2024-01-01T00-00-00-000000000Z_2024-01-01T23-59-59-999999999Z.parquet
│ └── trades/
│ └── BTCUSD.BINANCE/
│ └── 2024-01-01T00-00-00-000000000Z_2024-01-01T23-59-59-999999999Z.parquetCustom data uses data/custom/<type_name>/ with optional identifier path segments.
Instruments use one directory per concrete instrument class, such as data/currency_pair/ or
data/equity/, rather than a shared instruments directory.
Catalog queries, and therefore backtests, read only these canonical directories. Legacy layouts
such as data/trade_tick/, data/quote_tick/, and the Python-written
data/custom_<snake_case>/ are not read at query time. To upgrade a catalog that still uses a
legacy layout, convert it with
nautilus catalog migrate-parquet.
Overlapping writes
By default, overlapping writes raise an OSError to maintain data integrity.
Set skip_disjoint_check=True only when the overlap is intentional.
Reading data
Use a typed query when the expected return type is known. start and end are UNIX nanoseconds:
quotes = catalog.query_quote_ticks(
identifiers=["EUR/USD.SIM"],
start=1704067200000000000,
end=1704153600000000000,
)
trades = catalog.query_trade_ticks(
identifiers=["BTC/USD.BINANCE"],
start=1704067200000000000,
end=1704153600000000000,
)BacktestDataConfig: backtest data
BacktestDataConfig defines the catalog data that a BacktestNode loads for one run.
Core parameters
data_typeis aNautilusDataTypevalue:QuoteTick,TradeTick,Bar,OrderBookDelta,OrderBookDepth,MarkPriceUpdate,IndexPriceUpdate,FundingRateUpdate,InstrumentStatus,OptionGreeks,InstrumentClose, orInstrument.Instrumentloads every instrument class the catalog holds for the selected identifiers.catalog_pathidentifies the catalog root.- One of
instrument_id,instrument_ids, orbar_typesis required. start_timeandend_timeare optional UNIX nanosecond bounds.filter_expris an optional DataFusion SQL predicate.catalog_fs_protocolprefixescatalog_pathfor remote storage.catalog_fs_rust_storage_optionssupplies the Rust backend options. If it is unset,BacktestNodefalls back tocatalog_fs_storage_options.- For bars,
bar_speccombines with the instrument ID to select anEXTERNALbar type. Explicitbar_typescan select internal, external, or composite bars. optimize_file_loadingregisters whole directories when possible.
Basic usage examples
from nautilus_trader.config import BacktestDataConfig
from nautilus_trader.model import BarAggregation
from nautilus_trader.model import BarSpecification
from nautilus_trader.model import InstrumentId
from nautilus_trader.model import NautilusDataType
from nautilus_trader.model import PriceType
quote_data = BacktestDataConfig(
data_type=NautilusDataType.QuoteTick,
catalog_path="/path/to/catalog",
instrument_id=InstrumentId.from_str("EUR/USD.SIM"),
start_time=1704067200000000000,
end_time=1704153600000000000,
)
trade_data = BacktestDataConfig(
data_type=NautilusDataType.TradeTick,
catalog_path="/path/to/catalog",
instrument_ids=[
InstrumentId.from_str("BTC/USD.BINANCE"),
InstrumentId.from_str("ETH/USD.BINANCE"),
],
)
bar_data = BacktestDataConfig(
data_type=NautilusDataType.Bar,
catalog_path="/path/to/catalog",
instrument_id=InstrumentId.from_str("AAPL.NASDAQ"),
bar_spec=BarSpecification(5, BarAggregation.MINUTE, PriceType.LAST),
)This bar config selects AAPL.NASDAQ-5-MINUTE-LAST-EXTERNAL.
Cloud storage and filtering
book_data = BacktestDataConfig(
data_type=NautilusDataType.OrderBookDelta,
catalog_path="my-bucket/nautilus-data",
catalog_fs_protocol="s3",
catalog_fs_rust_storage_options={
"access_key_id": "your-access-key",
"secret_access_key": "your-secret-key",
"region": "us-east-1",
},
instrument_id=InstrumentId.from_str("BTC/USD.COINBASE"),
filter_expr="ts_init >= 1704067200000000000",
)Integration with BacktestRunConfig
Pass the data configurations to BacktestRunConfig:
from decimal import Decimal
from nautilus_trader.config import BacktestDataConfig
from nautilus_trader.config import BacktestRunConfig
from nautilus_trader.config import BacktestVenueConfig
from nautilus_trader.execution import MakerTakerFeeModel
from nautilus_trader.model import AccountType
from nautilus_trader.model import BookType
from nautilus_trader.model import InstrumentId
from nautilus_trader.model import NautilusDataType
from nautilus_trader.model import OmsType
data_configs = [
BacktestDataConfig(
data_type=NautilusDataType.QuoteTick,
catalog_path="/path/to/catalog",
instrument_id=InstrumentId.from_str("EUR/USD.SIM"),
),
]
run_config = BacktestRunConfig(
venues=[
BacktestVenueConfig(
name="SIM",
oms_type=OmsType.HEDGING,
account_type=AccountType.MARGIN,
book_type=BookType.L1_MBP,
starting_balances=["1_000_000 USD"],
fee_model=MakerTakerFeeModel(
maker_rate=Decimal("0"),
taker_rate=Decimal("0"),
),
),
],
data=data_configs,
start=1704067200000000000,
end=1704153600000000000,
)Data loading process
When a backtest runs, the BacktestNode processes each BacktestDataConfig:
- Create a
ParquetDataCatalogfrom the configuration. - Load the required instrument definitions while building the engine.
- Build and run a DataFusion query from the configuration fields.
- Sort merged data by
ts_initand add it to the backtest engine.
Direct catalog access
Use ParquetDataCatalog to query or write a catalog directly. Use BacktestDataConfig when a
BacktestNode should load catalog data for a run. LiveNodeConfig has no counterpart for loading
catalog data; request historical data through a configured data client or query the catalog
directly. Its streaming field configures feather writing only.
Querying and filtering
The generic query takes a catalog directory name such as quotes, trades, or bars. Use it when
you need the files or optimize_file_loading controls:
catalog.query(
data_type=NautilusDataType.QuoteTick,
identifiers=["EUR/USD.SIM"],
start=1704067200000000000,
end=1704153600000000000,
where_clause="ts_event <= ts_init",
files=None,
)Typed methods such as query_quote_ticks, query_trade_ticks, and query_bars return the concrete
model type. query_custom_data resolves custom decoders through the runtime registry. query, the
typed market-data query methods, and query_custom_data use UNIX nanosecond time bounds and accept
a DataFusion SQL where_clause.
Time-zone database mismatch
With the current Cargo.lock, DataFusion SQL temporal functions resolve named time zones with the
transitive chrono-tz 0.10.4 database (IANA 2025b). Rust core time-zone operations use Jiff 0.2.35
with its bundled IANA 2026c database. Zone results can differ when zone rules change or historical
data is corrected after 2025b until DataFusion migrates.
If RustSec files unmaintained advisories for chrono or chrono-tz, maintain matching documented
ignores in .cargo/audit.toml and deny.toml until DataFusion migrates.
Catalog operations
Catalog operations rename, consolidate, or delete data files. Each operation takes a type selector:
a NautilusDataType, a NautilusRecordType, or a NautilusInstrumentType.
NautilusDataType.Instrument covers every instrument class; a NautilusInstrumentType targets
one. delete_data_range(...) takes a NautilusDataType alone, because only data families support
ranged deletes.
Reset file names
Reset Parquet file names to match their content timestamps so filename-based filtering remains
accurate. reset_all_file_names() processes the entire catalog; reset_data_file_names(...)
targets a data path. Supply an instrument ID for data types partitioned by instrument. Without one,
the operation recursively reads the type directory and renames each file within its own instrument
directory, checking every directory's intervals before renaming any file.
catalog.reset_all_file_names()
catalog.reset_data_file_names(NautilusDataType.QuoteTick, "EUR/USD.SIM")
catalog.reset_data_file_names(NautilusDataType.TradeTick, "BTC/USD.BINANCE")Recover from overlapping file names
Overlapping file names block writes to the affected directory. A rejected coverage extension never renames files, so overlap points to an earlier rename, a manual move, or a concurrent writer. When the file contents are still disjoint, recover the directory with a filename reset:
- Stop all writers to the catalog.
- Back up the affected directory. The reset renames files one at a time, so an I/O failure partway through can leave some files renamed. On object stores that rename by copying then deleting, a failure can also leave a file under both names.
- Run
reset_data_file_names(...)for the affected path. For custom data types, usereset_all_file_names()instead, which covers every leaf directory including custom layouts. - If the reset reports that a new name is held by another file, rename that file to an unused interval name outside the data range, then run the reset again. Repeat until it succeeds.
- Confirm each renamed file matches its content range, then write a later disjoint interval to confirm the directory accepts writes again.
The reset processes one directory at a time. Before renaming any file in a directory, it reads
each file's ts_init range. It fails when the content ranges overlap or when a file's new name is
the current name of another file, and it then leaves every file name in that directory unchanged.
Directories reset before the failing one keep their new names.
Use filename reset only for filename-only damage. When file contents overlap, resetting names cannot reconcile the data, so the reset fails; rebuild the affected range from source through a separately validated process instead.
Consolidate catalog
Combine small Parquet files to reduce file count and query overhead.
With no bounds, consolidate_catalog() processes each leaf data directory in the catalog.
consolidate_data(...) operates on one directory; supply an instrument ID for data types
partitioned by instrument.
catalog.consolidate_catalog()
catalog.consolidate_catalog(
start=1704067200000000000,
end=1704153600000000000,
ensure_contiguous_files=True,
)
catalog.consolidate_data(
NautilusDataType.QuoteTick,
identifier="EUR/USD.SIM",
start=1704067200000000000,
end=1706745600000000000,
)Consolidate catalog by period
Split data files into fixed periods. Durations and time bounds use nanoseconds. Both methods accept optional bounds. Supply an identifier to the data-type method for data partitioned by instrument.
The catalog-wide method processes quotes, trades, order book deltas, order book depths, bars, index
prices, mark prices, instrument closes, and registered custom types. It logs a warning and skips
other types. The data-type method rejects instrument and record selectors, which have no
period-typed rewrite; use consolidate_data(...) for those.
DAY_NS = 86_400_000_000_000
HOUR_NS = 3_600_000_000_000
catalog.consolidate_catalog_by_period(period_nanos=DAY_NS)
catalog.consolidate_catalog_by_period(
period_nanos=HOUR_NS,
start=1704067200000000000,
end=1704153600000000000,
)
catalog.consolidate_data_by_period(
data_type=NautilusDataType.QuoteTick,
identifier="EUR/USD.SIM",
period_nanos=HOUR_NS,
)
catalog.consolidate_data_by_period(
data_type=NautilusDataType.TradeTick,
identifier="EUR/USD.SIM",
period_nanos=HOUR_NS,
start=1704067200000000000,
end=1706745600000000000,
)Delete data range
Delete data within a time range, optionally limited to one data type and instrument. Omitting
start extends the range to the beginning; omitting end extends it to the end. For
delete_data_range(...), omitting both bounds removes all matching data. Supply an instrument ID
for data partitioned by instrument.
delete_data_range(...) supports quotes, trades, bars, order book deltas, order book depth, and
registered custom types. Pass NautilusDataType.OrderBookDepth for order book depth and
NautilusDataType.Custom("MarketTickPython") for a custom data type.
delete_catalog_range(...) continues after unsupported directories, logs a warning, and leaves
their data unchanged. Use delete_data_range(...) when you need to confirm that the requested
type is supported. NautilusDataType.Instrument is rejected because instrument definitions do not
support ranged deletion.
catalog.delete_catalog_range(
start=1704067200000000000,
end=1704153600000000000,
)
catalog.delete_catalog_range(end=1704067200000000000)
catalog.delete_data_range(
data_type=NautilusDataType.QuoteTick,
identifier="BTC/USD.BINANCE",
)
catalog.delete_data_range(
data_type=NautilusDataType.TradeTick,
identifier="EUR/USD.SIM",
start=1704067200000000000,
end=1706745600000000000,
)Permanent data removal
Delete operations cannot be undone. The catalog splits partially overlapping files to preserve data outside the range.
Feather streaming and conversion
The runtime can stage records in local Feather files and promote them into a catalog with
StreamingConfig(writer_path=..., catalog=DataCatalogConfig(...)). Staged records become available
to catalog queries after promotion succeeds. A staging flush and a catalog commit are separate steps.
Streaming defaults to promotion on close, no interval-based promotion, and retention of committed
Feather sources. A positive promotion_interval_ms uses live wall-clock scheduling or checks
against the supplied backtest clock during writes and flushes. See
stream data into a Parquet catalog for defaults, configuration,
query visibility, and recovery.
StreamingFeatherWriter remains available for direct staging. Its completed sessions can be converted
manually with ParquetDataCatalog.convert_stream_to_data().
InstrumentClose
InstrumentClose represents a closing price event for an instrument at a venue. It covers end-of-session closes and contract-expiry close events.
Custom Data
NautilusTrader supports custom data authored in Python or Rust. Both forms use the same runtime routing, persistence, and query pipeline as built-in data.