You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解码含推文的.csv.gzip文件?读取遇Unicode解码错误求解

解决读取Twitter压缩CSV文件时的UnicodeDecodeError问题

问题场景

在Google Colab中合并多个Twitter情感分析数据集(.csv.gzip格式)时,执行读取代码触发UnicodeDecodeError,错误提示为'utf-8' codec can't decode byte 0xb8 in position 8048: invalid start byte,同时伴随列类型混合的警告。推测是推文中包含表情符号、非英文字符等特殊字符导致编码解码失败。

原读取代码:

temp_list = []

for file in apr_files:
    print(f"Reading in {file}")

    # unzip and read in the csv file as a dataframe
    temp = pd.read_csv(file, compression="gzip", header=0, index_col=0)
    
    # append dataframe to temp list
    temp_list.append(temp)

解决方案

针对编码问题和混合类型警告,通过修改pd.read_csv的参数即可解决:

1. 指定兼容的编码格式

Twitter数据集常使用latin-1编码(可兼容大部分特殊字符和表情符号),或尝试utf-8-sig(处理带BOM的UTF-8文件),在pd.read_csv中添加encoding参数。

2. 关闭低内存模式解决混合类型警告

添加low_memory=False参数,避免分块读取时因列类型识别不一致触发警告。

3. 可选:忽略/替换无法解码的字符

添加errors='replace'参数,将无法解码的字符替换为占位符�,确保程序不中断。

修改后的完整代码:

temp_list = []

for file in apr_files:
    print(f"Reading in {file}")
    temp = pd.read_csv(
        file,
        compression="gzip",
        header=0,
        index_col=0,
        encoding='latin-1',  # 优先尝试latin-1,无效可替换为utf-8-sig
        low_memory=False,
        errors='replace'  # 兜底方案,避免极端字符中断读取
    )
    temp_list.append(temp)

# 合并所有DataFrame
combined_tweets_df = pd.concat(temp_list, ignore_index=True)

参数说明

  • encoding='latin-1':Latin-1编码覆盖ASCII和西欧字符,能兼容Twitter数据中常见的特殊符号、表情的字节表示,是处理这类数据集的常用选择。
  • low_memory=False:强制Pandas一次性读取整个文件,避免分块读取时对列类型的错误推断,消除混合类型警告。
  • errors='replace':当遇到无法用指定编码解码的字符时,用占位符替代,保证文件能正常加载。

内容的提问来源于stack exchange,提问作者Queenish01

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 09:54:19