基于关键词批量删除多文件夹TXT文件报错求助
我从EDGAR下载了一批10-K报告,每个公司对应单独文件夹,需要保留包含“cryptocurrency”和“blockchain”关键词的TXT格式10-K报告,删除不含这些关键词的文件。现有Python代码第一步能正确生成目录列表,但第二步读取文件时触发FileNotFoundError,代码及报错信息如下:
原Step1代码(运行正常)
import os import pandas as pd path = 'C:/test/2014/QTR1/' words = ['cryptocurrency', 'blockchain'] filelist = os.listdir(path) Path2 = [] for x in filelist: Path2.append(path + x+ '/') print(Path2)
原Step2代码(触发报错)
for i in Path2: filelist2 = os.listdir(i) for j in filelist2: if j.endswith('.txt'): each_file_content = open(j, 'r', encoding="utf-8").read() if not any(word in each_file_content for word in words): os.unlink(j)
报错信息
FileNotFoundError Traceback (most recent call last) Input In [43], in <cell line: 1>()
3 for j in filelist2:
4 if j.endswith('.txt'):
----> 6 each_file_content = open(j, 'r', encoding="utf-8").read()
7 if not any(word in each_file_content for word in words):
8 os.unlink(j)FileNotFoundError: [Errno 2] No such file or directory:
'0001000180-14-000019.txt'
问题原因与修复
报错核心原因是打开文件时只传入了文件名j,未拼接文件所在的目录路径i。Python的open()和os.unlink()默认在当前工作目录查找文件,而非i对应的子文件夹,因此找不到目标文件。
修复时只需用os.path.join()拼接完整文件路径,同时推荐用with语句读取文件,避免资源泄漏:
修复后的Step2代码
for i in Path2: filelist2 = os.listdir(i) for j in filelist2: if j.endswith('.txt'): # 拼接完整文件路径 full_file_path = os.path.join(i, j) # 用with语句安全读取文件 with open(full_file_path, 'r', encoding="utf-8") as f: each_file_content = f.read() # 检查是否不含目标关键词,是则删除 if not any(word in each_file_content.lower() for word in words): os.unlink(full_file_path)
额外优化点
- 大小写不敏感匹配:将文件内容转为小写(
each_file_content.lower()),避免因关键词大小写差异漏判(比如文件中写的是Blockchain而非blockchain)。 - 路径拼接更稳妥:
os.path.join()自动适配不同操作系统的路径分隔符,比手动字符串拼接更可靠。 - 异常处理防中断:针对编码异常的文件添加捕获逻辑,避免程序中途终止:
for i in Path2: filelist2 = os.listdir(i) for j in filelist2: if j.endswith('.txt'): full_file_path = os.path.join(i, j) try: with open(full_file_path, 'r', encoding="utf-8") as f: each_file_content = f.read().lower() if not any(word in each_file_content for word in words): os.unlink(full_file_path) except UnicodeDecodeError: print(f"无法解码文件:{full_file_path},已跳过") continue
内容的提问来源于stack exchange,提问作者Ciercy

