如何修正代码实现将匹配短语的拆分结果与未匹配句子存入列表
问题描述
现有两个列表,分别包含完整句子和短语:
my_sentence=['This is string', 'This is string too', 'That is not string', 'That are not sentence'] my_phrase=['This is', 'That is']
需求:
- 找出
my_sentence中包含my_phrase元素的句子 - 用匹配到的短语拆分句子,将剩余部分(去除多余空格)存入
my_result列表 - 未匹配的句子直接存入该列表
预期结果:
my_result=['string','string too', 'not string', 'That are not sentence']
尝试的代码:
result=[] for sentence in my_sentence: for phrase in my_phrase: if phrase in sentence: res=sentence.split(phrase) result.append(res) else: res=sentence result.append(res) print(result)
得到的错误结果:
[['', ' string'], 'This is string', ['', ' string too'], 'This is string too', 'That is not string', ['', ' not string'], 'That are not sentence', 'That are not sentence']
错误原因
- 双重循环重复添加:每个句子会被
my_phrase中的两个短语依次遍历,无论匹配与否都会执行append,导致每个句子被重复添加两次 - split结果未处理:
split(phrase)返回包含空字符串和剩余部分的列表,直接append会把整个列表存入结果,而非提取目标文本 - 未处理多余空格:拆分后的剩余部分开头存在空格,未做清理
修正后的代码
my_result = [] for sentence in my_sentence: matched = False for phrase in my_phrase: if phrase in sentence: # 拆分后取第二个元素,去除首尾空格 remaining = sentence.split(phrase)[1].strip() my_result.append(remaining) matched = True break # 找到匹配短语后立即停止遍历,避免重复处理 if not matched: # 无匹配短语时,直接添加原句子 my_result.append(sentence) print(my_result)
代码逻辑说明
- 给每个句子设置
matched标记,用于判断是否找到匹配短语 - 匹配到短语后,通过
split拆分句子,取拆分后的第二个元素(短语位于句子开头),再用strip()去除多余空格 - 添加剩余部分后,设置
matched=True并break退出短语循环,避免同一个句子被多个短语重复处理 - 若遍历完所有短语都未匹配,直接将原句子加入结果列表
验证结果
运行修正后的代码,输出与预期完全一致:
['string', 'string too', 'not string', 'That are not sentence']
内容的提问来源于stack exchange,提问作者Anis
相关产品推荐
相关产品推荐

