如何用Python从多文本文件提取内容到CSV?StopIteration报错求解
问题分析
StopIteration异常是因为执行next(f)时,文件指针已走到末尾,无更多行可读取。常见触发场景:
- TXT文件中
author:行是最后一行 author:行后没有空白行
解决方案
不要用next(f)硬跳一行,改为循环跳过所有后续空白行(兼容无空白行的情况),同时优化文本提取逻辑以匹配需求格式。
修改后的完整代码:
from glob import glob import csv path = '**/*.txt' def extract_fields(fname): with open(fname, encoding='utf-8') as f: author, txt = "", "" lines = f.readlines() # 遍历定位author行 for idx, line in enumerate(lines): stripped_line = line.strip() if stripped_line.startswith("author: "): author = stripped_line[8:].strip() # 跳过author行后的所有空行 content_start_idx = idx + 1 while content_start_idx < len(lines) and lines[content_start_idx].strip() == "": content_start_idx += 1 # 合并剩余文本,将段落换行替换为空格 txt = ''.join(lines[content_start_idx:]).strip().replace('\n\n', ' ') break return author, txt rows = [] # 加上recursive=True确保遍历所有子文件夹 for fname in glob(path, recursive=True): author, txt = extract_fields(fname) rows.append([fname, author, txt]) with open("output.csv", "w", newline="", encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(["filename", "author", "text"]) # 对齐需求中的列名 writer.writerows(rows)
关键改动说明
- 用
readlines()一次性读取所有行,便于索引定位和跳过空行 - 找到author行后循环跳过后续空白行,避免硬跳行引发的异常
- 合并文本时替换段落间换行符为空格,匹配需求中的文本格式
- 给
glob添加recursive=True参数,确保递归遍历子文件夹 - 统一指定
utf-8编码,避免跨系统编码问题 - 表头改为需求中的
text,与目标CSV列名对齐
内容的提问来源于stack exchange,提问作者hy9fesh
相关产品推荐
相关产品推荐

