Python CSV Reader遇无分隔符行中断的处理方案咨询
处理手动生成CSV的异常行,避免循环中断
核心问题是你直接用csv.reader拆分后的行来处理,遇到异常行时会因索引越界(比如调用row[1]但字段不存在)导致隐性错误中断循环。解决方案是先预处理每一行原始文本,修复异常格式后再交给CSV解析逻辑,同时用正则精准匹配私有IP来拆分主机名和IP。
修复后的完整代码
import csv import re import argparse # 初始化设备列表 switches = [] routers = [] firewalls = [] termservers = [] sitecodes = [] device_error = '' # 私有IP正则(仅匹配10/172/192开头的私有地址) PRIVATE_IP_PATTERN = r'(?:10\.|172\.(?:1[6-9]|2[0-9]|3[01])\.|192\.168\.)(?:\d{1,3}\.){1}\d{1,3}' # 主机名正则(假设为字母数字组合,不含点号) HOSTNAME_PATTERN = r'^[a-zA-Z0-9]+' def fix_broken_line(line): # 剔除注释、空行和首尾空白 line = line.strip() if not line or line.startswith('#') or line.startswith('!'): return None # 情况1:无逗号但同时包含主机名和私有IP ip_match = re.search(PRIVATE_IP_PATTERN, line) if ip_match and ',' not in line: ip = ip_match.group() hostname = line.replace(ip, '').strip() if hostname: return f"{hostname},{ip}" else: return f",{ip}" # 情况2:仅主机名(无IP、无逗号) if re.match(HOSTNAME_PATTERN, line) and not re.search(PRIVATE_IP_PATTERN, line): return f"{line}," # 情况3:仅私有IP(无主机名、无逗号) if re.fullmatch(PRIVATE_IP_PATTERN, line): return f",{line}" # 情况4:处理多余逗号,保留最多两个有效字段 parts = [p.strip() for p in line.split(',') if p.strip()] if len(parts) == 1: if re.fullmatch(PRIVATE_IP_PATTERN, parts[0]): return f",{parts[0]}" else: return f"{parts[0]}," elif len(parts) >= 2: # 取前两个有效字段(主机名+IP) return f"{parts[0]},{parts[1]}" return line # 解析命令行参数(保留原有逻辑) parser = argparse.ArgumentParser() parser.add_argument('filename', nargs=1) args = parser.parse_args() try: with open(args.filename[0], 'r') as file: for line_num, raw_line in enumerate(file, start=1): fixed_line = fix_broken_line(raw_line) if not fixed_line: continue # 解析修复后的行 row = next(csv.reader([fixed_line], delimiter=',')) # 确保行始终有两个字段(空字段补全) while len(row) < 2: row.append('') hostname, ip = row[0].strip(), row[1].strip() if not hostname and not ip: continue # 验证IP是否为有效私有IP valid_ip = re.fullmatch(PRIVATE_IP_PATTERN, ip) if ip else None if valid_ip: # 检查站点码一致性 if sitecodes: if not all(sc == hostname[:5] for sc in sitecodes): device_error += f"\n=> Check hostname {hostname}, line {line_num}" else: sitecodes.append(hostname[:5]) print(f"{line_num} {hostname}") # 分类设备 if re.search(r'.*(rt101|rt102)', hostname): routers.append(ip) elif re.search(r'.*(rt1[1-2]|sw).*', hostname): switches.append(ip) else: # 错误信息收集 if re.fullmatch(PRIVATE_IP_PATTERN, hostname): device_error += f"\n=> Invalid row with IP in first field, line {line_num}" elif hostname and not ip: device_error += f"\n=> Missing IP for hostname {hostname}, line {line_num}" elif ip and not hostname: device_error += f"\n=> Missing hostname for IP {ip}, line {line_num}" else: device_error += f"\n=> Invalid IP {ip} for hostname {hostname}, line {line_num}" except FileNotFoundError: print(f"Error: File {args.filename[0]} not found.") except Exception as e: print(f"Unexpected error: {str(e)}") # 输出错误汇总 if device_error: print("\n=== Error Summary ===") print(device_error)
关键修复说明
- 预处理原始行:跳过注释、空行,直接处理文本内容,避免
csv.reader提前拆分导致的结构混乱 - 异常行修复逻辑:
- 无逗号但含主机名+IP:用私有IP正则拆分,自动插入逗号
- 仅主机名:补全逗号和空IP字段
- 仅私有IP:补全逗号和空主机名字段
- 多余逗号:过滤空字段,保留前两个有效字段
- 避免循环中断:确保处理后的行始终有两个字段(空字段补全),不会出现索引越界;用
enumerate跟踪行号,错误信息更精准 - 精准私有IP匹配:替换原有通用IP正则,只匹配10/172/192开头的私有地址,贴合需求
测试验证
用你提供的测试CSV运行,所有异常行都会被正确修复:
dntsw110110.2.3.1→ 修复为dntsw1101,10.2.3.1dntsw1101→ 修复为dntsw1101,10.2.3.1→ 修复为,10.2.3.1dntsw1101,,10.2.3.1→ 修复为dntsw1101,10.2.3.1
循环不会中断,后续正常行可继续处理,错误信息会准确记录异常情况。
内容的提问来源于stack exchange,提问作者ZeusTM
相关产品推荐
相关产品推荐

