Python正则表达式匹配文本导出CSV报错求助
问题解决:正则匹配日期时间写入CSV时的TypeError错误
错误核心原因
j.readlines()返回的是字符串列表,但re.search()仅接受单个字符串/类字节对象,这是触发TypeError的直接原因。- 内层无限
while True循环会重复读取文件内容,导致死循环。 - 条件判断
if match is not None in filtered_messages逻辑混乱,无法正确过滤目标行。 - 正则表达式未定义命名分组,使用
match.groupdict()无法获取有效数据。 - 每次循环用
w模式打开output.csv,会覆盖之前写入的内容。
修复步骤
- 替换
readlines()为逐行遍历文件的方式,确保给re.search()传入单个字符串。 - 移除内层无限循环,改用
for line in j遍历文件行。 - 修正条件判断:先确认匹配结果存在,再检查行内容是否包含目标关键词。
- 用
match.group()获取完整匹配的日期时间字符串,替代无效的groupdict()。 - 调整
output.csv的打开模式为a(追加),或提前一次性写入表头避免重复。 - 处理
test1.txt读取时的空行,防止尝试打开空文件名。
完整修复代码
import re import csv def output(): filtered_messages = ['Current date & time'] fieldnames = ['currenttime', 'output'] # 初始化CSV文件并写入表头(仅执行一次) with open('output.csv', 'w', newline='') as csv_file: writer = csv.DictWriter(csv_file, delimiter=',', fieldnames=fieldnames) writer.writeheader() with open("test1.txt", "r") as f: for line in f: filename = line.rstrip("\n") # 跳过空行,避免打开无效文件 if not filename: continue try: with open(filename, "r") as j: for content_line in j: # 检查行是否包含目标关键词 if any(msg in content_line for msg in filtered_messages): # 正则匹配日期时间格式,传入单个字符串 match = re.search(r'\d{4}-\d{2}-\d{2}_\d{2}-\d{2}-\d{2}', content_line) if match: # 追加写入匹配结果到CSV with open('output.csv', 'a', newline='') as csv_file: writer = csv.DictWriter(csv_file, delimiter=',', fieldnames=fieldnames) writer.writerow({ 'currenttime': match.group(), 'output': content_line.strip() }) except FileNotFoundError: print(f"文件 {filename} 不存在,跳过") output()
关键说明
- 正则表达式中的
\-可简化为-,非特殊位置无需转义。 - 使用
writeheader()替代手动写入表头,符合csv.DictWriter规范。 - 添加
try-except捕获文件不存在的异常,提升代码健壮性。 - 用
newline=''避免CSV写入时出现多余空行。
内容的提问来源于stack exchange,提问作者Srinivas
相关产品推荐
相关产品推荐

