CSV¶
csv:// reads local CSV files and writes one. It is a CSV-only spelling of the
file:// connector: same readers, same writer, same path grammar,
same rerun behaviour. Prefer file:// for new work, and reach for csv:// only
to keep an existing command working.
URI format¶
csv://path/to/file.csv
Everything after csv:// is a filesystem path, resolved exactly as
file:// resolves it: relative to the working directory, or
absolute with an extra leading slash. Globs, gzipped files, split form
(--source-uri csv:// --source-table data/users.csv), Windows drive and UNC
paths, and format hints all behave the same way.
omniload ingest \
--source-uri 'csv://data/users.csv' \
--source-table 'users' \
--dest-uri 'duckdb:///local.duckdb' \
--dest-table 'public.users'
CSV only¶
The scheme names the file format, so it accepts the CSV family of readers and
nothing else: csv, csv_headless and csv_duckdb.
URI |
Result |
|---|---|
|
Reads every matching CSV. A glob widens the path, not the format. |
|
Reads a header-less CSV. |
|
Reads a gzipped CSV. |
|
Rejected. Use |
|
Rejected. Use |
|
Rejected. Use |
The same restriction applies when writing: csv://out.jsonl and
csv://out.dat#parquet are rejected before the load starts. A path carrying no
recognised format at all (csv://report, csv://out.dat) writes CSV, because
the scheme already names the format.
Destination connector¶
omniload ingest \
--source-uri 'postgres://user:password@host:5432/db' \
--source-table 'public.users' \
--dest-uri 'csv://export/users.csv' \
--dest-table 'public.users'
--dest-table must be <schema>.<table>; it names only dlt’s intermediate
layout, while the output file is the URI path. Parent directories are created
if they don’t exist, and an existing file is overwritten. The output drops
dlt’s internal bookkeeping columns.
Note
Rows are written in a deterministic order, but not necessarily source order.
Behaviour changes¶
This release moves csv:// onto the file:// reader and writer, which changes
seven things:
Values are typed. The reader infers column types, so a numeric column arrives as a number and
true/falseas a boolean, where the old reader yielded every value as a string. ISO date strings stay strings.Empty rows are preserved. A row whose fields are all empty is loaded as a row of nulls instead of being dropped.
An empty field keeps its quoting. An unquoted empty field loads as null and a quoted one (
"") as an empty string, where the old reader dropped every empty field and so always produced null.A rerun appends. With no
--incremental-strategy, a second load of the same file adds a second copy rather than replacing the first. Pass--incremental-strategy replacefor the old behaviour. A destination that does not implement append rejects the load outright rather than duplicating rows; the MongoDB destination is one, so acsv://to MongoDB command now needs the explicit--incremental-strategy replace.merge,delete+insertandscd2are rejected. They need an incremental or merge key, which the shared reader does not expose. Use a source that does, or--full-refreshto reset the destination.--incremental-keyis rejected, with the same errorfile://gives. The row cursor it drove compared raw strings against parsed datetimes and crashed on an interval; it is gone rather than ported. File-level selection by modification time is available instead, through--filesystem-incremental.The whole load is written. A load dlt splits across several files used to write only the first of them and still exit zero. Every row now reaches the output file.