You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文件单词提取遇重复START OF行内容丢失问题排查

测试文件内容截图

问题描述

基础任务为编写函数get_words_from_file(filename),返回目标兴趣区域内的小写单词列表。题目提供正则表达式:[a-z]+[-'][a-z]+|[a-z]+[']?|[a-z]+,用于匹配符合定义的所有单词。
编写的代码在部分测试用例下运行正常,但当标识兴趣区域起始的行重复出现时运行失败。

原代码如下:

import re

def get_words_from_file(filename):
    """Returns a list of lower case words that are with the region of 
    interest, every word in the text file, but, not any of the punctuation."""
    with open(filename,'r', encoding='utf-8') as file:
        flag = False
        words = []
        count = 0
        for line in file:
            if line.startswith("*** START OF"):
                while count < 1:
                    flag=True
                    count += 1
            elif line.startswith("*** END"):
                flag=False
                break       
            elif(flag):
                new_line = line.lower()
                words_on_line = re.findall("[a-z]+[-'][a-z]+|[a-z]+[']?|[a-z]+", 
                                           new_line)
                words.extend(words_on_line)
    
        return words

#test code:
filename = "bee.txt"
words = get_words_from_file(filename)
print(filename, "loaded ok.")
print("{} valid words found.".format(len(words)))
print("Valid word list:")
for word in words:
    print(word)

具体故障:字符串*** START OF在兴趣区域内部重复出现时,该行内容不会被纳入提取范围。

输出对比

预期输出:

bee.txt loaded ok.
16 valid words found.
Valid word list:
yes
really
this
time
start
of
synthetic
test
case
end
synthetic
test
case
i'm
in
too

实际运行输出:

bee.txt loaded ok.
11 valid words found.
Valid word list:
yes
really
this
time
end
synthetic
test
case
i'm
in
too
问题原因

原代码分支判断逻辑存在漏洞:

  • 所有以*** START OF开头的行,无论当前是否已经进入兴趣区域,都会进入第一个判断分支,不会走到后续的单词提取逻辑,直接导致兴趣区域内重复出现的START行内容被跳过。
  • 代码中冗余的count变量和while循环无实际作用,可直接移除。
修复方案

调整分支判断逻辑:仅当尚未进入兴趣区域时,碰到*** START OF开头的行才判定为区域起始标记,开启提取;进入兴趣区域后,所有行(包括START开头的行)都按普通文本处理,正常提取单词,直到碰到*** END开头的结束标记为止。

修复后完整代码:

import re

def get_words_from_file(filename):
    """Returns a list of lower case words that are with the region of 
    interest, every word in the text file, but, not any of the punctuation."""
    with open(filename,'r', encoding='utf-8') as file:
        in_region = False
        words = []
        for line in file:
            # 未进入区域时,匹配起始标记
            if not in_region and line.startswith("*** START OF"):
                in_region = True
                continue
            # 匹配结束标记,直接终止提取
            if line.startswith("*** END"):
                break
            # 区域内所有行正常提取单词
            if in_region:
                new_line = line.lower()
                words_on_line = re.findall(r"[a-z]+[-'][a-z]+|[a-z]+[']?|[a-z]+", 
                                           new_line)
                words.extend(words_on_line)
    
        return words

#测试代码
filename = "bee.txt"
words = get_words_from_file(filename)
print(filename, "loaded ok.")
print("{} valid words found.".format(len(words)))
print("Valid word list:")
for word in words:
    print(word)

内容的提问来源于stack exchange,提问作者Hugo Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 12:06:16