使用pandas与glob导入CSV遇编码错误,求正确编码格式
解决CSV读取的UnicodeDecodeError问题
错误根源
你遇到的UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0,是因为目标CSV文件采用的是UTF-16编码(带字节顺序标记BOM)——UTF-16 LE的BOM以0xff开头,和你看到的错误特征完全匹配。
解决方案
1. 适配正确编码
将read_csv的encoding参数改为'utf-16'(pandas会自动识别BOM),或者更明确的'utf-16-le':
df = pd.read_csv(csv_file, encoding='utf-16')
2. 修复合并逻辑错误
你当前的循环代码存在逻辑漏洞:每次合并时都用初始空的combined_df和当前文件的df拼接,最终df_LV_StoreCard只会保留最后一个文件的数据。需要在循环中持续更新combined_df,避免数据丢失。
修正后的完整代码
# Import Module import pandas as pd import glob csv_files = glob.glob('PATH') # Create an empty dataframe to store the combined data combined_df = pd.DataFrame() # Loop through each CSV file and append its contents to the combined dataframe for csv_file in csv_files: df = pd.read_csv(csv_file, encoding='utf-16') combined_df = pd.concat([combined_df, df], ignore_index=True) # ignore_index避免索引重复冲突 # Assign to target variable and print df_LV_StoreCard = combined_df print(df_LV_StoreCard)
补充检测方案
如果utf-16仍无效,可以用chardet库自动检测文件编码:
import chardet with open(csv_file, 'rb') as f: detect_result = chardet.detect(f.read()) print(detect_result['encoding']) # 输出检测到的编码,再用该编码读取文件
内容的提问来源于stack exchange,提问作者Manav Shah
相关产品推荐
相关产品推荐

