Python筛选指定子串开头行写入新文件失败求助
问题
需要清理TXT文件,逐行读取后仅保留以预定义双字母组合(关键词)开头的行并写入新文件。预期保留所有以ag开头的行,但运行代码后生成的文件为空。
原始文件样本:
agigolón. (Tb. ajigolón). m. 1. El Salv., Guat., Méx. y Nic. Prisa, ajetreo. U. m. en pl. 112. El Salv., Guat., Hond., Méx. y Nic. Apuro, aprieto. U. m. en pl. 113. Guat., Méx. y Nic. Fatiga, cansancio. agigotar. tr. desus. hacer gigote. ágil. (Del lat. agilis). adj. 1. Que se mueve con soltura y rapidez. Estuvo muy ágil y esquivó el golpe. 12. Dicho de un movimiento: Hábil y rápido. Camina con paso ágil. 1 3. Que actúa o se desarrolla con rapidez o prontitud. Tiene una prosa ágil. agílibus. m. coloq. agibílibus. agilidad. (Del lat. agilítas, -atis). f. 1. Cualidad de ágil. 12. Rel. Una de las cuatro dotes de los cuerpos gloriosos, que consiste en la facultad de trasladarse de un lugar a otro instantáneamente, por grande que sea la distancia.
使用的代码:
from itertools import product path = r'C:\Users\Usuario\Desktop' spanish_alphabet = 'aábcdeéfghiíjklmnñoópqrstuúvwxyz' keywords = [''.join(i) for i in product(spanish_alphabet, repeat = 2)] Keywords = [i.capitalize() for i in keywords] keywords = keywords + Keywords A_keywords = [i for i in keywords if i.startswith(('A', 'a', 'Á', 'á'))] with open(path + '\raw_text.txt', 'r', encoding='utf-8') as input_file: with open(path + '\clean_text.txt', 'w', encoding ='utf-8') as output_file: for line in input_file: # If line begins with given keyword, then write it in clean_text file if line.strip("\n").startswith(tuple(A_keywords)): output_file.write(line + '\n')
问题排查与修复
核心问题
- 行首空白未处理:
line.strip("\n")仅移除换行符,未处理行首可能存在的空格/空白字符,导致目标行的实际开头不是纯双字母,匹配失败。 - 额外添加换行符:原代码中
output_file.write(line + '\n')会生成重复换行符,因为读取的每行本身已包含换行符。 - 冗余的关键词生成:生成所有双字母组合完全没必要,针对需求直接定义目标前缀即可提升效率。
修复后的代码(针对仅保留ag开头的行)
path = r'C:\Users\Usuario\Desktop' # 定义需要匹配的所有目标前缀(含大小写、重音变体) target_prefixes = {'ag', 'Ag', 'ág', 'Ág'} with open(path + '\raw_text.txt', 'r', encoding='utf-8') as input_file: with open(path + '\clean_text.txt', 'w', encoding='utf-8') as output_file: for line in input_file: # 先去除行首所有空白字符,再检查前缀匹配 stripped_line = line.lstrip() if stripped_line.startswith(tuple(target_prefixes)): output_file.write(line) # 写入原始行,避免重复换行
扩展方案(保留所有A/a/Á/á开头的双字母行)
如果需要保留所有以A/a/Á/á开头的双字母行,保留原关键词生成逻辑,但需处理行首空白:
from itertools import product path = r'C:\Users\Usuario\Desktop' spanish_alphabet = 'aábcdeéfghiíjklmnñoópqrstuúvwxyz' keywords = [''.join(i) for i in product(spanish_alphabet, repeat = 2)] Keywords = [i.capitalize() for i in keywords] keywords = keywords + Keywords A_keywords = [i for i in keywords if i.startswith(('A', 'a', 'Á', 'á'))] with open(path + '\raw_text.txt', 'r', encoding='utf-8') as input_file: with open(path + '\clean_text.txt', 'w', encoding='utf-8') as output_file: for line in input_file: stripped_line = line.lstrip() if stripped_line.startswith(tuple(A_keywords)): output_file.write(line)
内容的提问来源于stack exchange,提问作者Javi
相关产品推荐
相关产品推荐

