You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中循环使用多个str.startswith()提取值时遇异常

解决CSV文件解析中多条件str.startswith()失效的问题

问题背景

编写Python函数解析目录下的CSV文件,通过str.startswith()提取特定行的值:

  • 首行以TrakPro开头、提取Serial Number:行的逻辑正常工作
  • 添加Test Name:行的提取逻辑后,无输出也无报错,该部分代码未生效
  • 需求:在循环中添加多个str.startswith()判断,提取CSV中各类冒号后的关键词值(如Test Name:对应的值)

原代码

def get_csv_file_list(root):
    for r, d, f in os.walk(root):
        for file in f:

            if file.endswith('.csv'):
                path = os.path.join(r, file)
                dir_name = path.split(os.path.sep)[-2]
                file_name = os.path.basename(file)

                try:
                    with open(path) as k:
                        firstline = k.readline()

                        if firstline.startswith('TrakPro'):
                            file_list.append(path)
                            file_list.append(dir_name)
                            file_list.append(file_name)

                            txt = 'Serial Number:'
                            if txt.startswith('Serial'):
                                for row in list(k)[3:4]:
                                    file_list.append(row[15:26])

                            txt2 = 'Test Name:'
                            if txt2.startswith('Test'):
                                for rows in list(k)[4:5]:
                                    print(rows)
                                    file_list.append(row[11:])

CSV内容示例

TrakPro Version 5.2.0.0 ASCII Data File
Instrument Name:,SidePak
Model Number:,TK0W02
Serial Number:,120k2136005
Test Name:,13270
Start Date:,04/17/2021
Start Time:,01:53:29
Duration (dd:hh:mm:ss):,00:07:13:00
Log Interval (mm:ss):,01:00
Number of points:,433
Description:,

问题原因

  1. 文件指针偏移:第一次调用list(k)[3:4]时,会将文件指针移动到文件末尾,后续list(k)[4:5]无法读取到任何内容
  2. 冗余判断:txt = 'Serial Number:'后判断txt.startswith('Serial')是恒成立的,属于无效逻辑
  3. 变量名错误:第二个循环中使用未定义的row变量(循环变量是rows),会导致隐性错误

修改后的代码

import os

def get_csv_file_list(root):
    file_list = []  # 确保函数内初始化列表,避免全局变量依赖
    for r, d, f in os.walk(root):
        for file in f:
            if file.endswith('.csv'):
                path = os.path.join(r, file)
                dir_name = os.path.basename(os.path.dirname(path))  # 更可靠的上级目录获取方式
                file_name = os.path.basename(file)

                try:
                    with open(path, 'r', encoding='utf-8') as k:
                        lines = k.readlines()  # 一次性读取所有行到列表,避免指针问题
                        first_line = lines[0].strip()
                        
                        if first_line.startswith('TrakPro'):
                            file_list.extend([path, dir_name, file_name])

                            # 提取Serial Number
                            serial_line = lines[3].strip()
                            if serial_line.startswith('Serial Number:'):
                                serial_num = serial_line.split(',')[1].strip()
                                file_list.append(serial_num)

                            # 提取Test Name
                            test_line = lines[4].strip()
                            if test_line.startswith('Test Name:'):
                                test_name = test_line.split(',')[1].strip()
                                print(test_name)
                                file_list.append(test_name)
                except Exception as e:
                    print(f"处理文件 {path} 时出错: {str(e)}")
    return file_list

关键优化点

  • 一次性读取所有行到lines列表,彻底解决文件指针偏移问题
  • 使用split(',')分割行内容提取值,比固定索引更健壮(避免因行内容长度变化导致的取值错误)
  • 移除冗余的字符串判断逻辑,精简代码
  • 修复变量名错误,确保逻辑正确性
  • 增加异常捕获与提示,便于排查文件读取异常
  • 优化目录名获取方式,适配不同操作系统路径格式

内容的提问来源于stack exchange,提问作者ksu1980

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 04:25:22