Python文本字符串搜索异常:所有结果均判定为未找到
问题排查:Python文件字符串搜索异常
问题说明
调试一段用于在两个文件中搜索字符串的Python代码,所有符合条件的条目都进入了「Ca和Co均未找到」的分支,但单独用小段测试代码却能正常匹配到字符串。
逻辑流程:
- 先过滤
try_ID.txt中不含Ca或Co的行 - 对剩余行,按以下规则分支处理:
- 若
try_ID.txt中的Ca和Co均未在try.txt与try_C.txt中出现,进入第一个if分支 - 若仅能在其中一个文件找到Ca或Co,进入elif分支
- 若两个文件中均能找到Ca和Co,进入else分支
- 若
相关代码与文件内容
原代码
import re with open("try_ID.txt", 'r') as fin, \ open("try_C.txt", 'r') as co_splice, \ open("try.txt", 'r') as ca_splice: for row in fin: if len(re.findall("Ca", row)) == 0 or len(re.findall("Co", row)) == 0: pass else: # problem starts from here name = str(row.split()[1]) + "_blast" if not row.split()[1] in ca_splice.read() and not row.split()[2] in co_splice.read(): print(row.split()[0:2]) elif row.split()[1] in ca_splice.read() and not row.split()[2] in col_splice.read(): print(row.split()[1] + "Ca") elif not row.split()[1] in can_splice.read() and row.split()[2] in col_splice.read(): print(row.split()[2] + "Co") else: ne_name = name + "recip" print(ne_name)
文件内容
try_ID.txt
H21911 Ca29092.1t A05340.1 H21912 Ca19588.1t Co27353.1t A05270.1 H21913 Ca19590.1t Co14899.1t A05260.1 H21914 Ca19592.1t Co14897.1t A05240.1 H21915 Co14877.1t A05091.1 S25338 Ca12595.1t Co27352.1t A53970.1 S20778 Ca29091.1t Co24326.1t A61120.1 S26552 Ca20916.1t Co14730.1t A16155.1
try_C.txt
Co14730.1t;Co14730.2t Co27352.1t;Co27352.2t;Co27352.3t;Co27352.4t;Co27352.5t Co14732.1t;Co14732.2t Co4217.1t;Co4217.2t Co27353.1t;Co27353.2t Co14733.1t;Co14733.2t
try.txt
Ca12595.1t;Ca12595.2t Ca29091.1t;Ca29091.2t Ca1440.1t;Ca1440.2t Ca29092.1t;Ca29092.2t Ca20916.1t;Ca20916.2t
测试代码(可正常运行)
row = "H20118 Ca12595.1t Co18779.1t A01010.1" text_file = "try.txt" with open(text_file, 'r') as fin: if row.split()[1] in fin.read(): print(True) else: print(False)
错误原因分析
- 文件指针耗尽问题:文件对象的
read()方法只能读取一次,第一次调用后文件指针会移动到文件末尾,后续调用read()会返回空字符串。原代码在循环中多次调用ca_splice.read()和co_splice.read(),导致除第一次判断外,后续所有判断都基于空字符串,自然会判定为「未找到」。 - 变量名拼写错误:代码中出现
col_splice、can_splice的错误变量名,正确的应该是co_splice、ca_splice,这会直接导致运行时报错。 - 过滤逻辑冗余:用
len(re.findall("Ca", row)) == 0判断是否包含Ca完全没必要,直接用"Ca" not in row更简洁高效。
修复后的代码
# 先把文件内容读取到变量中,避免重复读取 with open("try_C.txt", 'r') as f: co_content = f.read() with open("try.txt", 'r') as f: ca_content = f.read() with open("try_ID.txt", 'r') as fin: for row in fin: row = row.strip() if not row: continue # 简化过滤逻辑:同时包含Ca和Co的行才处理 if "Ca" not in row or "Co" not in row: continue parts = row.split() ca_id = parts[1] co_id = parts[2] name = f"{ca_id}_blast" ca_found = ca_id in ca_content co_found = co_id in co_content if not ca_found and not co_found: print(parts[0:2]) elif ca_found and not co_found: print(f"{ca_id}Ca") elif not ca_found and co_found: print(f"{co_id}Co") else: ne_name = f"{name}recip" print(ne_name)
内容的提问来源于stack exchange,提问作者zzz
相关产品推荐
相关产品推荐

