You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将多个.txt文件合并为单个DataFrame并解决编码错误?

解决方案

问题分析

你的代码存在三个核心问题:

  1. 变量名不匹配:定义了list = []却用li.append(),会触发NameError;
  2. 缺少文件名记录:最终DataFrame没有保留文件名列;
  3. 编码不兼容:强制指定cp1252编码,无法处理其他编码的文件,导致UnicodeDecodeError。

方案一:自动检测文件编码(推荐)

借助chardet库自动识别每个文件的编码,确保读取时不报错,同时完整保留文件名和文本内容。

步骤1:安装依赖库

pip install chardet

步骤2:完整代码

import os
import glob
import pandas as pd
import chardet

# 目标目录路径
data_dir = r'/Users/alldata'
all_files = glob.glob(os.path.join(data_dir, '*.txt'))

# 存储每个文件的文件名和文本
data_records = []

for file_path in all_files:
    # 提取不带后缀的文件名
    filename = os.path.splitext(os.path.basename(file_path))[0]
    
    # 检测文件编码
    with open(file_path, 'rb') as f:
        encoding_info = chardet.detect(f.read())
        file_encoding = encoding_info['encoding']
    
    # 读取文件内容,异常时用utf-8忽略错误兜底
    try:
        with open(file_path, 'r', encoding=file_encoding) as f:
            # 读取所有文本,可根据需求替换换行符
            text_content = f.read().replace('\n', ' ')
    except (UnicodeDecodeError, TypeError):
        with open(file_path, 'r', encoding='utf-8', errors='ignore') as f:
            text_content = f.read().replace('\n', ' ')
    
    # 将当前文件信息加入列表
    data_records.append({'filename': filename, 'text': text_content})

# 生成目标DataFrame
final_df = pd.DataFrame(data_records)
print(final_df.head())

方案二:尝试常用编码兜底(无需额外库)

如果无法安装chardet,可以依次尝试常见编码,失败时用忽略错误的方式读取:

import os
import glob
import pandas as pd

data_dir = r'/Users/alldata'
all_files = glob.glob(os.path.join(data_dir, '*.txt'))
data_records = []

# 定义常用编码列表
common_encodings = ['cp1252', 'utf-8', 'latin-1', 'gbk']

def read_text_file(file_path):
    # 依次尝试编码
    for enc in common_encodings:
        try:
            with open(file_path, 'r', encoding=enc) as f:
                return f.read().replace('\n', ' ')
        except UnicodeDecodeError:
            continue
    # 所有编码失败时,忽略错误读取
    with open(file_path, 'r', encoding='utf-8', errors='ignore') as f:
        return f.read().replace('\n', ' ')

for file_path in all_files:
    filename = os.path.splitext(os.path.basename(file_path))[0]
    text = read_text_file(file_path)
    data_records.append({'filename': filename, 'text': text})

final_df = pd.DataFrame(data_records)

代码说明

  • 两种方案都会自动提取文件名(不带.txt后缀),满足你需要的DataFrame结构;
  • 针对编码问题,要么自动检测,要么兜底尝试,避免单个文件报错中断整个流程;
  • 直接用open()读取文件比pd.read_fwf更高效,适合获取完整文本内容。

内容的提问来源于stack exchange,提问作者datascientistxxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 10:08:00