Read a CSV in Python without losing track of your data
Load a small file, inspect its columns, and handle the two common surprises: delimiters and missing values.
Before computing anything, inspect what the file actually contains. A CSV is plain text with rows and separators, but the separator and missing-value conventions can vary.
Start with a tiny example
Save this as scores.csv next to your Python script:
student,score
Ada,82
Ben,91
Cora,
Python's built-in csv module can read it without installing a package.
import csv
with open('scores.csv', newline='', encoding='utf-8') as file:
reader = csv.DictReader(file)
rows = list(reader)
print(reader.fieldnames) # ['student', 'score']
print(rows[0]) # {'student': 'Ada', 'score': '82'}
CSV values arrive as strings. Convert the score only after checking for blanks:
for row in rows:
score = int(row['score']) if row['score'] else None
print(row['student'], score)
Check the separator
If the headers appear as one combined field, the file may use semicolons. Pass delimiter=';' to csv.DictReader. Inspect the first few lines in a text editor before assuming a format.
When the file grows
For larger analysis tasks, pandas provides useful inspection tools. Its read_csv function infers common missing values and types, but you should still check the result.
import pandas as pd
df = pd.read_csv('scores.csv')
print(df.head())
print(df.info())
print(df.isna().sum())
If you do not have pandas installed, use the standard-library example above. For your own file, verify the header names, row count, types, and missing values before plotting or calculating an average.