You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas与glob导入CSV遇编码错误,求正确编码格式

解决CSV读取的UnicodeDecodeError问题

错误根源

你遇到的UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0,是因为目标CSV文件采用的是UTF-16编码(带字节顺序标记BOM)——UTF-16 LE的BOM以0xff开头,和你看到的错误特征完全匹配。

解决方案

1. 适配正确编码

将read_csv的encoding参数改为'utf-16'(pandas会自动识别BOM),或者更明确的'utf-16-le':

df = pd.read_csv(csv_file, encoding='utf-16')

2. 修复合并逻辑错误

你当前的循环代码存在逻辑漏洞:每次合并时都用初始空的combined_df和当前文件的df拼接,最终df_LV_StoreCard只会保留最后一个文件的数据。需要在循环中持续更新combined_df,避免数据丢失。

修正后的完整代码

# Import Module
import pandas as pd
import glob

csv_files = glob.glob('PATH')

# Create an empty dataframe to store the combined data
combined_df = pd.DataFrame()

# Loop through each CSV file and append its contents to the combined dataframe
for csv_file in csv_files:
    df = pd.read_csv(csv_file, encoding='utf-16')
    combined_df = pd.concat([combined_df, df], ignore_index=True)  # ignore_index避免索引重复冲突

# Assign to target variable and print
df_LV_StoreCard = combined_df
print(df_LV_StoreCard)

补充检测方案

如果utf-16仍无效,可以用chardet库自动检测文件编码:

import chardet

with open(csv_file, 'rb') as f:
    detect_result = chardet.detect(f.read())
print(detect_result['encoding'])  # 输出检测到的编码,再用该编码读取文件

内容的提问来源于stack exchange,提问作者Manav Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 04:30:04