Python文件单词提取遇重复START OF行内容丢失问题排查
问题描述
基础任务为编写函数get_words_from_file(filename),返回目标兴趣区域内的小写单词列表。题目提供正则表达式:[a-z]+[-'][a-z]+|[a-z]+[']?|[a-z]+,用于匹配符合定义的所有单词。
编写的代码在部分测试用例下运行正常,但当标识兴趣区域起始的行重复出现时运行失败。
原代码如下:
import re def get_words_from_file(filename): """Returns a list of lower case words that are with the region of interest, every word in the text file, but, not any of the punctuation.""" with open(filename,'r', encoding='utf-8') as file: flag = False words = [] count = 0 for line in file: if line.startswith("*** START OF"): while count < 1: flag=True count += 1 elif line.startswith("*** END"): flag=False break elif(flag): new_line = line.lower() words_on_line = re.findall("[a-z]+[-'][a-z]+|[a-z]+[']?|[a-z]+", new_line) words.extend(words_on_line) return words #test code: filename = "bee.txt" words = get_words_from_file(filename) print(filename, "loaded ok.") print("{} valid words found.".format(len(words))) print("Valid word list:") for word in words: print(word)
具体故障:字符串*** START OF在兴趣区域内部重复出现时,该行内容不会被纳入提取范围。
输出对比
预期输出:
bee.txt loaded ok. 16 valid words found. Valid word list: yes really this time start of synthetic test case end synthetic test case i'm in too
实际运行输出:
bee.txt loaded ok. 11 valid words found. Valid word list: yes really this time end synthetic test case i'm in too
问题原因
原代码分支判断逻辑存在漏洞:
- 所有以
*** START OF开头的行,无论当前是否已经进入兴趣区域,都会进入第一个判断分支,不会走到后续的单词提取逻辑,直接导致兴趣区域内重复出现的START行内容被跳过。 - 代码中冗余的
count变量和while循环无实际作用,可直接移除。
修复方案
调整分支判断逻辑:仅当尚未进入兴趣区域时,碰到*** START OF开头的行才判定为区域起始标记,开启提取;进入兴趣区域后,所有行(包括START开头的行)都按普通文本处理,正常提取单词,直到碰到*** END开头的结束标记为止。
修复后完整代码:
import re def get_words_from_file(filename): """Returns a list of lower case words that are with the region of interest, every word in the text file, but, not any of the punctuation.""" with open(filename,'r', encoding='utf-8') as file: in_region = False words = [] for line in file: # 未进入区域时,匹配起始标记 if not in_region and line.startswith("*** START OF"): in_region = True continue # 匹配结束标记,直接终止提取 if line.startswith("*** END"): break # 区域内所有行正常提取单词 if in_region: new_line = line.lower() words_on_line = re.findall(r"[a-z]+[-'][a-z]+|[a-z]+[']?|[a-z]+", new_line) words.extend(words_on_line) return words #测试代码 filename = "bee.txt" words = get_words_from_file(filename) print(filename, "loaded ok.") print("{} valid words found.".format(len(words))) print("Valid word list:") for word in words: print(word)
内容的提问来源于stack exchange,提问作者Hugo Smith
相关产品推荐
相关产品推荐

