NOETRION

Browse by topic

← All articles
LESSON 3 OF 4 · Python & data foundations
DATA WORKFLOW · 8 MIN READ

Read a CSV in Python without losing track of your data

Load a small file, inspect its columns, and handle the two common surprises: delimiters and missing values.

Before computing anything, inspect what the file actually contains. A CSV is plain text with rows and separators, but the separator and missing-value conventions can vary.

Start with a tiny example

Save this as scores.csv next to your Python script:

student,score
Ada,82
Ben,91
Cora,

Python's built-in csv module can read it without installing a package.

import csv

with open('scores.csv', newline='', encoding='utf-8') as file:
    reader = csv.DictReader(file)
    rows = list(reader)

print(reader.fieldnames)  # ['student', 'score']
print(rows[0])           # {'student': 'Ada', 'score': '82'}

CSV values arrive as strings. Convert the score only after checking for blanks:

for row in rows:
    score = int(row['score']) if row['score'] else None
    print(row['student'], score)

Check the separator

If the headers appear as one combined field, the file may use semicolons. Pass delimiter=';' to csv.DictReader. Inspect the first few lines in a text editor before assuming a format.

When the file grows

For larger analysis tasks, pandas provides useful inspection tools. Its read_csv function infers common missing values and types, but you should still check the result.

import pandas as pd

df = pd.read_csv('scores.csv')
print(df.head())
print(df.info())
print(df.isna().sum())

If you do not have pandas installed, use the standard-library example above. For your own file, verify the header names, row count, types, and missing values before plotting or calculating an average.

Check your understanding

What type does csv.DictReader return for a score such as 82?

Choose one answer