You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3.5在Windows下提取文本文件两行间的多行内容

解决Python提取剧本人名的问题

Hey there! Let's get this sorted out. The issues you're facing—only getting one line and hitting StopIteration with next()—are pretty common when dealing with file iteration without proper boundary checks. Here's a solid, error-resistant solution tailored to your needs:

可行的实现代码

# 替换成你的输入/输出文件实际路径
input_script_path = "your_script.txt"
output_chars_path = "extracted_characters.txt"

characters = []
is_collecting = False

# 读取剧本文件,收集角色名
with open(input_script_path, 'r', encoding='utf-8') as infile:
    for line in infile:
        clean_line = line.strip()
        
        # 触发开始收集的标记
        if clean_line == "Characters:":
            is_collecting = True
            continue  # 跳过标记行本身
        
        # 触发停止收集的标记
        if clean_line == "First scene.":
            is_collecting = False
            break  # 找到结束标记后直接终止循环,提升效率
        
        # 只收集非空的有效角色行
        if is_collecting and clean_line:
            characters.append(clean_line)

# 将收集到的角色名写入新文件
with open(output_chars_path, 'w', encoding='utf-8') as outfile:
    for char in characters:
        outfile.write(f"{char}\n")

为什么这个方案能解决你的问题

  • 告别StopIteration错误: 不用next()手动迭代(它会在无内容可读时直接报错),而是用for循环遍历文件对象——Python会自动处理文件结束的情况,不会抛出异常。
  • 完整提取所有角色: 通过is_collecting布尔标记控制收集状态,从Characters:之后开始收集,到First scene.时停止,确保捕获中间所有有效角色行。
  • 处理边缘情况: 自动去除每行的空白字符,跳过空行,避免输出文件出现无效的空白条目;指定encoding='utf-8'解决Windows下常见的中文乱码问题。

你最初尝试可能踩的坑

如果你的代码只输出一行,大概率是只调用了一次next()而没有循环收集中间所有行;StopIteration则是因为next()在文件结束或未找到First scene.标记时,尝试读取不存在的内容导致的。

这个方法简单直观,在Python 3.5及以上版本的Windows系统上可以稳定运行,赶紧试试吧!

内容的提问来源于stack exchange,提问作者FabianPeters

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:55:45