You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas.read_csv读取泰坦尼克号CSV时整行存首单元格问题求助

Troubleshooting Your Titanic CSV Loading Issue

Hey there, let's work through why your pandas read_csv call is shoving all data into the first column! This is a super common gotcha with CSVs, so let's break down the most likely fixes:

1. Check for File Encoding or Hidden BOM Characters

Sometimes CSV files (especially those saved on Windows) use non-UTF-8 encodings like cp1252, or have a hidden UTF-8 BOM that messes up pandas parsing. Try these encoding tweaks first:

# Try Windows-friendly encoding
import pandas as pd
titanic = pd.read_csv('titanic.csv', encoding='cp1252')
print(titanic.head(10))

# Or if it's a UTF-8 file with a hidden BOM
titanic = pd.read_csv('titanic.csv', encoding='utf-8-sig')

2. Verify the Actual Delimiter (It Might Not Be What It Looks Like)

Even if your sample shows commas, sometimes invisible characters (like tabs, or non-breaking spaces masquerading as commas) can throw pandas off. Let's use Python's built-in csv module to peek at how the file is actually structured:

import csv
# Use the same encoding you tested above if needed
with open('titanic.csv', 'r', encoding='cp1252') as f:
    reader = csv.reader(f)
    # Print first 4 rows to see how they're split
    for i, row in enumerate(reader):
        print(f"Row {i}: {row}")
        if i == 3:
            break

If the csv module splits the rows correctly but pandas doesn't, try forcing pandas to use the Python engine (instead of the default C engine, which can be stricter):

titanic = pd.read_csv('titanic.csv', engine='python', encoding='cp1252')

3. Fix Quoting Handling

Your sample has quoted fields (like the passenger names with commas inside), so sometimes pandas' default quoting rules don't catch this properly. Explicitly set the quote character:

titanic = pd.read_csv('titanic.csv', quotechar='"', engine='python', encoding='cp1252')

4. Check for Weird Line Endings

If the file was created on a different OS (e.g., Windows vs. Linux), mismatched line endings can cause parsing issues. Try specifying the line terminator:

# For Windows-style line endings
titanic = pd.read_csv('titanic.csv', lineterminator='\r\n', encoding='cp1252')

# For Unix-style line endings
titanic = pd.read_csv('titanic.csv', lineterminator='\n', encoding='cp1252')

Start with the encoding checks first—those are usually the culprit for this exact "all data in first column" issue!

内容的提问来源于stack exchange,提问作者Karollo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:14:01