NOETRION

Browse by topic

← All articles
DATA WORKFLOW · 11 MIN READ

CSV vs Parquet: choose a data format by workload, not by file extension

Understand row-oriented text, columnar storage, schemas, compression, pushdown, and a safe conversion workflow in Python and SQL.

Reviewed September 30, 2026. Performance depends on the engine, compression, row-group layout, storage, and query. Treat the qualitative comparisons below as design guidance, not a universal benchmark.

Editorial illustration comparing a text row grid with column-oriented compressed data blocks.
CSV serializes readable rows. Parquet groups typed column values so analytical engines can skip data a query does not need.

CSV is easy to open, inspect, generate, and exchange. Parquet is designed for efficient analytical storage and retrieval. Neither format is “better” in every situation; they optimize different work.

What a CSV file actually stores

CSV is delimited text. A header may name the fields, but the file does not provide a universal type system. The bytes 00123, 2026-09-30, and an empty field require interpretation by the reader.

order_id,date,city,total
00123,2026-09-30,Surabaya,125000
00124,2026-09-30,Jakarta,99000

This simplicity is a strength for exchange and debugging. It also moves work to the parser: delimiters, quotes, encodings, missing values, decimals, dates, and leading zeroes must be handled consistently.

How Parquet organizes the same table

Apache Parquet is a column-oriented file format. A file contains one or more row groups. Each row group contains one contiguous column chunk per column, and column chunks are divided into pages. Encoding and compression operate at the page level.

CSV layout
  1. Header text
  2. Row 1: every field
  3. Row 2: every field
  4. Row 3: every field
Parquet layout
  1. Row group
  2. order_id column chunk
  3. date column chunk
  4. city and total chunks

Because values in one column usually share a type and often repeat patterns, column-oriented encoding and compression can be effective. The file also carries schema and metadata that engines use to navigate the data.

The practical differences

ConcernCSVParquet
Human inspectionExcellent. Open in a text editor.Needs a compatible reader.
Schema and typesReader infers or receives rules separately.Stored in the file.
CompressionUsually whole-file external compression.Column-aware encoding and compression.
Select a few columnsParser generally reads each row.Projection can read only required columns.
Skip filtered regionsLimited without an external index.Statistics may allow row-group or page skipping.
Streaming appendSimple to append text rows carefully.Usually written in batches; layout choices matter.
InterchangeNearly universal.Strong across analytical tools, less universal elsewhere.

Projection and filter pushdown

Suppose a dataset contains 80 columns, but a query only needs date, city, and total. DuckDB's Parquet reader can push the projection into the scan and read only those columns. A filter such as date >= DATE '2026-09-01' can also be pushed down; metadata may let the engine skip row groups whose ranges cannot match.

SELECT city, SUM(total) AS revenue
FROM 'orders.parquet'
WHERE date >= DATE '2026-09-01'
GROUP BY city
ORDER BY revenue DESC;

Skipping is not guaranteed to be equally effective for every file. Row-group size, sort order, statistics, compression, and data distribution all affect the result. A badly organized Parquet dataset can underperform a well-designed one.

When CSV is the right choice

  • A person must inspect or edit a small extract without specialized software.
  • A public download needs the broadest compatibility.
  • A system emits simple line-oriented records and another step immediately ingests them.
  • The file is small enough that parse cost and storage are not meaningful constraints.
  • You need a transparent handoff artifact beside a formal data dictionary.

CSV remains valuable because failures are visible. If a delimiter or quote is broken, you can often open the file and see why.

When Parquet earns its complexity

  • Analysts repeatedly scan large tables.
  • Queries use subsets of columns.
  • Storage and transfer costs matter.
  • Preserving numeric, boolean, date, timestamp, nested, or null types matters.
  • Files feed engines such as DuckDB, Spark, Polars, pandas, or a data lake.

Parquet is a file format, not a database. It does not provide transactions, concurrent updates, constraints, access control, or record-level mutation by itself. If many users edit operational state, use a database and export Parquet for analytics.

Convert with an explicit schema check

import pandas as pd

df = pd.read_csv(
    "orders.csv",
    dtype={"order_id": "string", "city": "string"},
    parse_dates=["date"]
)

assert df["order_id"].notna().all()
assert (df["total"] >= 0).all()

df.to_parquet(
    "orders.parquet",
    index=False,
    compression="zstd"
)

check = pd.read_parquet("orders.parquet")
assert len(check) == len(df)
assert list(check.columns) == list(df.columns)

Explicitly protect identifiers with leading zeroes, parse dates intentionally, validate ranges and nulls, and compare row counts. A conversion can succeed technically while changing meaning.

Partitioning is not a substitute for thinking

A Parquet dataset is often split into directories such as year=2026/month=09/. Partitioning can help an engine skip files, but too many tiny files create metadata and scheduling overhead. Choose partition columns that appear in common filters and have manageable cardinality.

Layout symptomLikely problemDirection to test
Thousands of tiny filesToo much partitioning or overly frequent writes.Compact into fewer, larger files.
Selective query scans most row groupsLoose min/max ranges.Sort by commonly filtered columns before writing.
One file cannot use available parallelismRow groups are too large or too few.Test row-group sizing for the target engine.
Schema differs across filesUncontrolled evolution.Validate schemas and define compatibility rules.

A defensible workflow

Receive
Keep the raw source
→Validate
Types, nulls, keys
→Convert
Write controlled Parquet
→Verify
Counts and sample queries
Keep the raw file and conversion code so the analytical artifact can be reproduced.
  1. Store the original CSV immutably with its source date and checksum.
  2. Define expected columns and types instead of relying entirely on inference.
  3. Reject or quarantine malformed rows; do not silently drop them.
  4. Write Parquet with a compression and row-group strategy suited to the workload.
  5. Verify schema, row count, null count, numeric totals, and representative queries.
  6. Version the conversion code and record the library versions used.

Choose CSV for transparency and interchange. Choose Parquet for typed, repeated analytical scans. In many pipelines the correct answer is both: preserve the raw CSV, then produce a validated Parquet layer for analysis.