You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:统计文本中两词同段落出现次数的代码返回异常

问题排查与修复方案

原代码存在的核心问题

  • 仅处理第一段就终止循环:else: break会在遇到第一个空行时直接跳出循环,后续所有段落都未被遍历,这是结果远低于预期的主要原因。
  • 单词匹配逻辑错误:'word1' in l是判断列表中是否存在完全等于目标单词的元素(即某一行整行就是该单词),但实际需求是检查段落中任意一行包含该单词,逻辑完全不符。
  • 遗漏无空行结尾的最后一段:如果文件最后一段没有以空行结尾,这段内容会留在列表中,循环结束后不会被检查。
  • 函数参数未复用:定义的filename参数未使用,硬编码文件名降低了代码灵活性。

修正后的代码

def count_same_paragraph(filename, word1, word2):
    book = readText(filename)
    current_paragraph = []
    count = 0
    
    for line in book:
        # 过滤仅含空白字符的行,避免误判
        if line.strip() != '':
            current_paragraph.append(line)
        else:
            # 检查当前段落是否同时包含两个目标单词
            has_word1 = any(word1 in line for line in current_paragraph)
            has_word2 = any(word2 in line for line in current_paragraph)
            if has_word1 and has_word2:
                count += 1
            # 重置段落列表,准备处理下一段
            current_paragraph = []
    # 处理文件末尾无空行结尾的最后一段
    if current_paragraph:
        has_word1 = any(word1 in line for line in current_paragraph)
        has_word2 = any(word2 in line for line in current_paragraph)
        if has_word1 and has_word2:
            count += 1
    return count

# 调用示例
print(count_same_paragraph('book.txt', 'word1', 'word2'))

修正说明

  1. 移除break语句,改为遇到空行时检查当前段落并重置列表,确保遍历完整文件
  2. 使用any(word in line for line in ...)判断段落中是否存在包含目标单词的行,匹配实际需求
  3. 增加循环结束后对剩余段落的检查,避免遗漏无空行结尾的内容
  4. 复用函数参数filename,并将目标单词设为参数,提升代码通用性
  5. 用line.strip() != ''替代line != '',过滤仅含空格/制表符的无效空行

内容的提问来源于stack exchange,提问作者mari00

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:52:51