You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何将迭代处理多文本文件的结果存入列表或DataFrame列

问题修复方案

你遇到的覆盖问题是代码存在几处疏漏导致的,修正后即可实现DataFrame每列对应一个文本文件的处理结果:

错误点梳理

  • 未导入string库:代码调用了string.punctuation处理标点,缺少对应导入会直接运行报错
  • 路径不完整:os.listdir()仅返回目录下的纯文件名,没有和你的目标根目录拼接,无法正确读取到目标文件
  • 变量名不匹配:循环中遍历得到的文件名变量是filename,打开文件时错误使用了未定义的infile变量
  • 可优化项:stop_words不需要每次循环都重复生成,放在循环外可以提升运行效率

修正后的完整代码

import pandas as pd
import numpy as np
import nltk
import os
import glob
import string  # 新增标点处理依赖导入
from nltk import sent_tokenize
from nltk import word_tokenize
from nltk.corpus import stopwords

# 展开用户目录路径,识别~符号指向的真实目录
mydirectory = os.path.expanduser('~/Documents/Textfiles')

datasentence = []
# 停用词集合放在循环外,避免重复生成浪费性能
stop_words = set(stopwords.words('english'))
file_names = []  # 可选:存储文件名,后续可直接给DataFrame列命名

for filename in os.listdir(mydirectory):
    if filename.endswith(".txt"):
        # 拼接得到文件的完整路径
        full_path = os.path.join(mydirectory, filename)
        # 用正确的完整路径打开文件,指定utf-8编码避免特殊字符报错
        with open(full_path, "r", encoding="utf-8") as input_file:
            sentences = sent_tokenize(input_file.read())
            tokens = [w.lower() for w in sentences]
            table = str.maketrans('', '', string.punctuation)
            stripped = [w.translate(table) for w in tokens]
            finalwords = [w for w in stripped if not w in stop_words]
            datasentence.append(finalwords)
            file_names.append(filename)  # 可选:记录当前处理的文件名

# 转置后每列为一个文件的处理结果
df1 = pd.DataFrame(datasentence).T
df1.columns = file_names  # 可选:将列名设置为对应文件名,方便识别
display(df1.head())

效果说明

修正后datasentence会按遍历顺序依次存储每个txt文件的finalwords结果,不会出现内容覆盖的情况,转置后的DataFrame每一列就对应一个txt文件的处理结果,配置列名后可以直接看到每列对应的源文件名称。

内容的提问来源于stack exchange,提问作者jeanstaeej

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 13:21:01