Python字典迭代时if-else条件二次循环失效问题解决
问题分析与解决
问题场景
我需要用正则表达式判断字典value中是否存在指定关键词,若存在则提取关键词之后的所有文本,并更新对应key的value。但运行代码时,第二个键值对无法正常工作,错误复用了第一个键的处理结果。
原代码
import re x = {'text1' : 'whatever text may or may not come here KEYWORD yada yada yada whatever blah', 'text2' : 'this is also KEYWORD another text that might contain yada yada yada blah'} for k, v in x.items(): match1 = re.search(r"regex_pattern", v, flags = re.IGNORECASE) match2 = re.search(r"regex_pattern", v, flags = re.IGNORECASE) match3 = re.search(r"regex_pattern", v, flags = re.IGNORECASE) match4 = re.search(r"regex_pattern", v, flags = re.IGNORECASE) if match1: print('match1 found for ', k) key_text1 = v[match1.end():] elif match2: print('match2 found for ', k) key_text2 = v[match2.end():] elif match3: print('match3 found for ', k) key_text3 = v[match3.end():] elif match4: print('match4 found for ', k) key_text4 = v[match4.end():] else: print('None of the above matches! found for ', k) if key_text1: x[k] = key_text1 elif key_text2: x[k] = key_text2 elif key_text3: x[k] = key_text3 elif key_text4: x[k] = key_text4 else: print('nothing!')
期望输出
x = {'text1' : ' yada yada yada whatever blah', 'text2' : ' another text that might contain yada yada yada blah'}
实际输出
x = {'text1' : ' yada yada yada whatever blah', 'text2' : ' yada yada yada whatever blah'}
原因分析
你的猜测完全正确:key_text1/key_text2/key_text3/key_text4这些变量在循环中没有被重置。第一次迭代时,key_text1被赋值为第一个文本的匹配结果;第二次迭代时,即使没有匹配到match1,key_text1仍然保留第一次的值,导致if key_text1:条件触发,错误地将旧值赋给第二个key。
解决方案
方案1:每次循环重置结果变量
在循环开始时,将所有结果变量初始化为None,确保每次迭代都是独立的:
import re x = {'text1' : 'whatever text may or may not come here KEYWORD yada yada yada whatever blah', 'text2' : 'this is also KEYWORD another text that might contain yada yada yada blah'} for k, v in x.items(): # 每次循环初始化结果变量为None key_text1 = key_text2 = key_text3 = key_text4 = None match1 = re.search(r"regex_pattern1", v, flags=re.IGNORECASE) match2 = re.search(r"regex_pattern2", v, flags=re.IGNORECASE) match3 = re.search(r"regex_pattern3", v, flags=re.IGNORECASE) match4 = re.search(r"regex_pattern4", v, flags=re.IGNORECASE) if match1: print('match1 found for ', k) key_text1 = v[match1.end():] elif match2: print('match2 found for ', k) key_text2 = v[match2.end():] elif match3: print('match3 found for ', k) key_text3 = v[match3.end():] elif match4: print('match4 found for ', k) key_text4 = v[match4.end():] else: print('None of the above matches! found for ', k) # 用is not None判断,避免空字符串被误判 if key_text1 is not None: x[k] = key_text1 elif key_text2 is not None: x[k] = key_text2 elif key_text3 is not None: x[k] = key_text3 elif key_text4 is not None: x[k] = key_text4 else: print('nothing!') print(x)
方案2:简化代码,用单个变量存储结果
如果四个正则模式是不同关键词,可以合并成一个正则表达式,用单个变量存储匹配结果,避免多变量的混乱:
import re x = {'text1' : 'whatever text may or may not come here KEYWORD yada yada yada whatever blah', 'text2' : 'this is also KEYWORD another text that might contain yada yada yada blah'} # 合并四个关键词为一个正则模式,用|分隔 pattern = r"KEYWORD|pattern2|pattern3|pattern4" # 替换为实际的四个关键词/模式 for k, v in x.items(): match = re.search(pattern, v, flags=re.IGNORECASE) if match: print(f'match found for {k}') x[k] = v[match.end():] else: print(f'None of the above matches! found for {k}') print(x)
这个方案不仅解决了变量复用的问题,还大幅简化了代码逻辑,更易维护。
内容的提问来源于stack exchange,提问作者user21785694
相关产品推荐
相关产品推荐

