CSV (comma-separated values) is the most common format for sharing tabular data — and one of the most error-prone.
Delimiters
Not every "CSV" uses commas. Semicolons are common in regions that use commas as decimal separators (the UCI wine quality data uses semicolons), and tab-separated files are common too. Check the first few lines before loading.
import pandas as pd
df = pd.read_csv("data.csv", sep=";")
Headers
Some files have no header row; others have title lines above the header. Use header=None with names=[...], or skiprows to skip preamble lines.
Quoting
Values containing the delimiter or line breaks should be enclosed in quotes. Badly quoted files cause misaligned columns — count columns per row to spot problems.
Encoding
Text can be UTF-8, Latin-1 or Windows-1252, among others. Garbled accented characters signal the wrong encoding; specify encoding= when reading and write UTF-8 when producing files.
Types
CSV stores everything as text. Check that numbers, dates and IDs are interpreted correctly — leading zeros in postcodes and account numbers disappear if read as numbers, so read such columns as strings.
Missing Values
Specify the markers used (na_values=["?", "NA", ""]).
Large Files
Read in chunks, select only needed columns, or convert to a columnar format such as Parquet for faster repeated use.