You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python筛选指定子串开头行写入新文件失败求助

问题

需要清理TXT文件,逐行读取后仅保留以预定义双字母组合(关键词)开头的行并写入新文件。预期保留所有以ag开头的行,但运行代码后生成的文件为空。

原始文件样本:

agigolón. (Tb. ajigolón). m. 1. El Salv., Guat., Méx. y Nic. Prisa, ajetreo. U.
m. en pl. 112. El Salv., Guat., Hond., Méx. y Nic. Apuro, aprieto. U. m. en
pl. 113. Guat., Méx. y Nic. Fatiga, cansancio.

agigotar. tr. desus. hacer gigote.
ágil. (Del lat. agilis). adj. 1. Que se mueve con soltura y rapidez. Estuvo
muy ágil y esquivó el golpe. 12. Dicho de un movimiento: Hábil y rápido.
Camina con paso ágil. 1 3. Que actúa o se desarrolla con rapidez o
prontitud. Tiene una prosa ágil.

agílibus. m. coloq. agibílibus.
agilidad. (Del lat. agilítas, -atis). f. 1. Cualidad de ágil. 12. Rel. Una de las
cuatro dotes de los cuerpos gloriosos, que consiste en la facultad de
trasladarse de un lugar a otro instantáneamente, por grande que sea la
distancia.

使用的代码:

from itertools import product

path = r'C:\Users\Usuario\Desktop'

spanish_alphabet = 'aábcdeéfghiíjklmnñoópqrstuúvwxyz'
keywords = [''.join(i) for i in product(spanish_alphabet, repeat = 2)]
Keywords = [i.capitalize() for i in keywords]
keywords = keywords + Keywords

A_keywords = [i for i in keywords if i.startswith(('A', 'a', 'Á', 'á'))]

with open(path + '\raw_text.txt', 'r', encoding='utf-8') as input_file:
    with open(path + '\clean_text.txt', 'w', encoding ='utf-8') as output_file:
        for line in input_file:
            # If line begins with given keyword, then write it in clean_text file
            if line.strip("\n").startswith(tuple(A_keywords)):
                output_file.write(line + '\n')
问题排查与修复

核心问题

  1. 行首空白未处理:line.strip("\n")仅移除换行符,未处理行首可能存在的空格/空白字符,导致目标行的实际开头不是纯双字母,匹配失败。
  2. 额外添加换行符:原代码中output_file.write(line + '\n')会生成重复换行符,因为读取的每行本身已包含换行符。
  3. 冗余的关键词生成:生成所有双字母组合完全没必要,针对需求直接定义目标前缀即可提升效率。

修复后的代码(针对仅保留ag开头的行)

path = r'C:\Users\Usuario\Desktop'

# 定义需要匹配的所有目标前缀(含大小写、重音变体)
target_prefixes = {'ag', 'Ag', 'ág', 'Ág'}

with open(path + '\raw_text.txt', 'r', encoding='utf-8') as input_file:
    with open(path + '\clean_text.txt', 'w', encoding='utf-8') as output_file:
        for line in input_file:
            # 先去除行首所有空白字符,再检查前缀匹配
            stripped_line = line.lstrip()
            if stripped_line.startswith(tuple(target_prefixes)):
                output_file.write(line)  # 写入原始行,避免重复换行

扩展方案(保留所有A/a/Á/á开头的双字母行)

如果需要保留所有以A/a/Á/á开头的双字母行,保留原关键词生成逻辑,但需处理行首空白:

from itertools import product

path = r'C:\Users\Usuario\Desktop'

spanish_alphabet = 'aábcdeéfghiíjklmnñoópqrstuúvwxyz'
keywords = [''.join(i) for i in product(spanish_alphabet, repeat = 2)]
Keywords = [i.capitalize() for i in keywords]
keywords = keywords + Keywords

A_keywords = [i for i in keywords if i.startswith(('A', 'a', 'Á', 'á'))]

with open(path + '\raw_text.txt', 'r', encoding='utf-8') as input_file:
    with open(path + '\clean_text.txt', 'w', encoding='utf-8') as output_file:
        for line in input_file:
            stripped_line = line.lstrip()
            if stripped_line.startswith(tuple(A_keywords)):
                output_file.write(line)

内容的提问来源于stack exchange,提问作者Javi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 12:01:28