Python:如何将迭代处理多文本文件的结果存入列表或DataFrame列
问题修复方案
你遇到的覆盖问题是代码存在几处疏漏导致的,修正后即可实现DataFrame每列对应一个文本文件的处理结果:
错误点梳理
- 未导入
string库:代码调用了string.punctuation处理标点,缺少对应导入会直接运行报错 - 路径不完整:
os.listdir()仅返回目录下的纯文件名,没有和你的目标根目录拼接,无法正确读取到目标文件 - 变量名不匹配:循环中遍历得到的文件名变量是
filename,打开文件时错误使用了未定义的infile变量 - 可优化项:
stop_words不需要每次循环都重复生成,放在循环外可以提升运行效率
修正后的完整代码
import pandas as pd import numpy as np import nltk import os import glob import string # 新增标点处理依赖导入 from nltk import sent_tokenize from nltk import word_tokenize from nltk.corpus import stopwords # 展开用户目录路径,识别~符号指向的真实目录 mydirectory = os.path.expanduser('~/Documents/Textfiles') datasentence = [] # 停用词集合放在循环外,避免重复生成浪费性能 stop_words = set(stopwords.words('english')) file_names = [] # 可选:存储文件名,后续可直接给DataFrame列命名 for filename in os.listdir(mydirectory): if filename.endswith(".txt"): # 拼接得到文件的完整路径 full_path = os.path.join(mydirectory, filename) # 用正确的完整路径打开文件,指定utf-8编码避免特殊字符报错 with open(full_path, "r", encoding="utf-8") as input_file: sentences = sent_tokenize(input_file.read()) tokens = [w.lower() for w in sentences] table = str.maketrans('', '', string.punctuation) stripped = [w.translate(table) for w in tokens] finalwords = [w for w in stripped if not w in stop_words] datasentence.append(finalwords) file_names.append(filename) # 可选:记录当前处理的文件名 # 转置后每列为一个文件的处理结果 df1 = pd.DataFrame(datasentence).T df1.columns = file_names # 可选:将列名设置为对应文件名,方便识别 display(df1.head())
效果说明
修正后datasentence会按遍历顺序依次存储每个txt文件的finalwords结果,不会出现内容覆盖的情况,转置后的DataFrame每一列就对应一个txt文件的处理结果,配置列名后可以直接看到每列对应的源文件名称。
内容的提问来源于stack exchange,提问作者jeanstaeej
相关产品推荐
相关产品推荐

