Python新手问题:文本取数匹配CSV行失败及正则适配难题
问题描述
我是Python初学者,有一个文本文件numberstofind.txt,内容如下:
-49 -56 -62
需要在CSV文件midvalues1.csv中查找这些数字,找到后输出对应行。但脚本始终找不到第二个数字-56(确认它存在于CSV中)。
原有脚本:
import os import csv import sys with open("numberstofind.txt", 'r') as fp: data = ",{0}".format(fp.readline().strip()) with open("midvalues1.csv", "r") as csv: found = False for line in csv: if data in line: print(line) print('Found !') f = open("ident1.txt", "w+") f.write(line) f.close() found = True if not found: print('Data not found!') with open("numberstofind.txt", 'r') as pd: phrasedeux = pd.readlines() phrasedeux=phrasedeux[1] print(phrasedeux) with open("midvalues1.csv", "r") as csvd: found = False for line in csvd: if phrasedeux in line: print(line) print('Found !') f = open("ident2.txt", "w+") f.write(line) f.close() found = True if not found: print('Data not found!') with open("numberstofind.txt", 'r') as ph: phrasetrois = ph.readlines() phrasetrois=phrasetrois[2] with open("midvalues1.csv", "r") as csv: found = False for line in csv: if phrasetrois in line: print(line) print('Found !') f = open("ident3.txt", "w+") f.write(line) f.close() found = True if not found: print('Data not found!')
输出结果:
876,6,-49,1754,John Doe Found ! -56 Data not found! 411,6,-62,48,Some other name Found !
尝试用正则表达式:
import re import csv with open("numberstofind.txt", 'r') as d: dt = d.readlines() dt=dt[1] print(dt) with open('midvalues1.csv', 'r') as csv: lines = csv.readlines() for line in lines: if re.search(r'-56', line): print(line) break
硬编码-56时可以正常工作,但换成读取的变量dt就失效了。问实现从文本文件读取数字并在CSV中匹配对应行的最优方案是什么?
问题分析与最优方案
问题根源
- 换行符残留:用
readlines()读取的字符串会保留末尾的\n换行符,比如phrasedeux = pd.readlines()[1]得到的是'-56\n',直接用这个字符串匹配CSV行自然找不到——CSV行里的-56没有后续换行符。 - 脚本冗余重复:原有脚本重复编写三次查找逻辑,既不高效也易出错。
- 字符串匹配风险:直接用
,{数字}格式匹配可能出现误判(比如-5会匹配到-56),或者适配不了数字在CSV行首尾的情况。
最优实现方案
步骤1:正确读取目标数字
读取numberstofind.txt,过滤空行并去除每行的空白字符(换行、空格等),得到干净的目标数字列表。
步骤2:用CSV模块处理文件
Python的csv模块可正确解析CSV每行,将行拆分为字段列表,精准匹配字段,避免字符串匹配的歧义。
步骤3:批量查找并保存结果
遍历目标数字,在CSV中查找包含该数字的行,找到后统一处理(打印、保存到文件)。
完整代码
import csv # 读取目标数字,过滤空行并清洗 target_numbers = [] with open("numberstofind.txt", 'r') as f: for line in f: stripped_line = line.strip() if stripped_line: # 跳过空行 target_numbers.append(stripped_line) # 遍历每个目标数字,查找CSV对应行 for idx, num in enumerate(target_numbers, 1): found = False with open("midvalues1.csv", 'r', newline='') as csvfile: reader = csv.reader(csvfile) for row in reader: if num in row: # 将行转换为CSV格式字符串 row_str = ','.join(row) print(row_str) print('Found !') # 保存到对应文件 with open(f"ident{idx}.txt", 'w') as outfile: outfile.write(row_str + '\n') # 加换行符保持格式 found = True # 若只需第一个匹配行则break;需所有匹配行则删除break break if not found: print(f'Data {num} not found!')
代码关键点说明
- 清洗目标数字:用
strip()去除每行空白字符,if stripped_line跳过空行,确保得到干净的数字字符串。 - CSV模块解析:
csv.reader将每行拆分为字段列表,直接判断数字是否在列表中,精准匹配无歧义。 - 批量处理:用
enumerate自动生成文件序号(ident1.txt、ident2.txt...),避免重复代码。 - 格式保持:保存行时添加
\n,确保输出文件格式与CSV原行一致。
正则变量失效的原因
dt = d.readlines()[1]得到的是包含换行符的字符串(比如'-56\n'),正则搜索时会匹配-56加换行符,而CSV行里的-56后无换行符,因此匹配失败。解决方法是对dt做strip()处理:
dt = d.readlines()[1].strip()
但相比正则,用CSV模块解析字段的方法更可靠。
内容的提问来源于stack exchange,提问作者Damien
相关产品推荐
相关产品推荐

