使用pandas.read_csv读取泰坦尼克号CSV时整行存首单元格问题求助
Hey there, let's work through why your pandas read_csv call is shoving all data into the first column! This is a super common gotcha with CSVs, so let's break down the most likely fixes:
1. Check for File Encoding or Hidden BOM Characters
Sometimes CSV files (especially those saved on Windows) use non-UTF-8 encodings like cp1252, or have a hidden UTF-8 BOM that messes up pandas parsing. Try these encoding tweaks first:
# Try Windows-friendly encoding import pandas as pd titanic = pd.read_csv('titanic.csv', encoding='cp1252') print(titanic.head(10)) # Or if it's a UTF-8 file with a hidden BOM titanic = pd.read_csv('titanic.csv', encoding='utf-8-sig')
2. Verify the Actual Delimiter (It Might Not Be What It Looks Like)
Even if your sample shows commas, sometimes invisible characters (like tabs, or non-breaking spaces masquerading as commas) can throw pandas off. Let's use Python's built-in csv module to peek at how the file is actually structured:
import csv # Use the same encoding you tested above if needed with open('titanic.csv', 'r', encoding='cp1252') as f: reader = csv.reader(f) # Print first 4 rows to see how they're split for i, row in enumerate(reader): print(f"Row {i}: {row}") if i == 3: break
If the csv module splits the rows correctly but pandas doesn't, try forcing pandas to use the Python engine (instead of the default C engine, which can be stricter):
titanic = pd.read_csv('titanic.csv', engine='python', encoding='cp1252')
3. Fix Quoting Handling
Your sample has quoted fields (like the passenger names with commas inside), so sometimes pandas' default quoting rules don't catch this properly. Explicitly set the quote character:
titanic = pd.read_csv('titanic.csv', quotechar='"', engine='python', encoding='cp1252')
4. Check for Weird Line Endings
If the file was created on a different OS (e.g., Windows vs. Linux), mismatched line endings can cause parsing issues. Try specifying the line terminator:
# For Windows-style line endings titanic = pd.read_csv('titanic.csv', lineterminator='\r\n', encoding='cp1252') # For Unix-style line endings titanic = pd.read_csv('titanic.csv', lineterminator='\n', encoding='cp1252')
Start with the encoding checks first—those are usually the culprit for this exact "all data in first column" issue!
内容的提问来源于stack exchange,提问作者Karollo

