如何将多个.txt文件合并为单个DataFrame并解决编码错误?
解决方案
问题分析
你的代码存在三个核心问题:
- 变量名不匹配:定义了
list = []却用li.append(),会触发NameError; - 缺少文件名记录:最终DataFrame没有保留文件名列;
- 编码不兼容:强制指定
cp1252编码,无法处理其他编码的文件,导致UnicodeDecodeError。
方案一:自动检测文件编码(推荐)
借助chardet库自动识别每个文件的编码,确保读取时不报错,同时完整保留文件名和文本内容。
步骤1:安装依赖库
pip install chardet
步骤2:完整代码
import os import glob import pandas as pd import chardet # 目标目录路径 data_dir = r'/Users/alldata' all_files = glob.glob(os.path.join(data_dir, '*.txt')) # 存储每个文件的文件名和文本 data_records = [] for file_path in all_files: # 提取不带后缀的文件名 filename = os.path.splitext(os.path.basename(file_path))[0] # 检测文件编码 with open(file_path, 'rb') as f: encoding_info = chardet.detect(f.read()) file_encoding = encoding_info['encoding'] # 读取文件内容,异常时用utf-8忽略错误兜底 try: with open(file_path, 'r', encoding=file_encoding) as f: # 读取所有文本,可根据需求替换换行符 text_content = f.read().replace('\n', ' ') except (UnicodeDecodeError, TypeError): with open(file_path, 'r', encoding='utf-8', errors='ignore') as f: text_content = f.read().replace('\n', ' ') # 将当前文件信息加入列表 data_records.append({'filename': filename, 'text': text_content}) # 生成目标DataFrame final_df = pd.DataFrame(data_records) print(final_df.head())
方案二:尝试常用编码兜底(无需额外库)
如果无法安装chardet,可以依次尝试常见编码,失败时用忽略错误的方式读取:
import os import glob import pandas as pd data_dir = r'/Users/alldata' all_files = glob.glob(os.path.join(data_dir, '*.txt')) data_records = [] # 定义常用编码列表 common_encodings = ['cp1252', 'utf-8', 'latin-1', 'gbk'] def read_text_file(file_path): # 依次尝试编码 for enc in common_encodings: try: with open(file_path, 'r', encoding=enc) as f: return f.read().replace('\n', ' ') except UnicodeDecodeError: continue # 所有编码失败时,忽略错误读取 with open(file_path, 'r', encoding='utf-8', errors='ignore') as f: return f.read().replace('\n', ' ') for file_path in all_files: filename = os.path.splitext(os.path.basename(file_path))[0] text = read_text_file(file_path) data_records.append({'filename': filename, 'text': text}) final_df = pd.DataFrame(data_records)
代码说明
- 两种方案都会自动提取文件名(不带
.txt后缀),满足你需要的DataFrame结构; - 针对编码问题,要么自动检测,要么兜底尝试,避免单个文件报错中断整个流程;
- 直接用
open()读取文件比pd.read_fwf更高效,适合获取完整文本内容。
内容的提问来源于stack exchange,提问作者datascientistxxx
相关产品推荐
相关产品推荐

