You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas转换TSV到CSV时遭遇UnicodeDecodeError错误求助

解决TSV转CSV时的UnicodeDecodeError问题

错误原因

你遇到的UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0,是因为你的TSV文件采用的是UTF-16编码(开头的0xff是UTF-16的字节顺序标记BOM),而Pandas默认用UTF-8解码,导致解码失败。

解决方案

方法1:直接指定UTF-16编码读取

修改代码,在read_table中添加encoding='utf-16'参数,读取后再导出为CSV:

import pandas as pd

# 读取TSV文件
df = pd.read_table('downloads/Survey.tsv', sep='\t+', header=None, encoding='utf-16')
print(df)

# 转换为CSV(指定UTF-8编码避免乱码)
df.to_csv('downloads/Survey.csv', index=False, encoding='utf-8')

方法2:先检测文件实际编码

如果不确定文件编码,可以用chardet库检测:

import chardet

with open('downloads/Survey.tsv', 'rb') as f:
    file_content = f.read()
    encoding_info = chardet.detect(file_content)
print(f"文件编码:{encoding_info['encoding']}")

得到编码后,将其传入read_table的encoding参数即可。

内容的提问来源于stack exchange,提问作者Avenger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 07:17:04