You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

排查CSV中第二组连续零对应列值的Python代码故障

问题:CSV文件中查找第二次连续零对应的值失败

问题背景

需要遍历指定文件夹内的所有CSV文件,找到索引1列中第二次连续出现的零,并返回对应行的索引0列的值。但现有代码运行后始终提示未找到目标值,手动核对数据确认存在该情况,之前使用纯数字数据时代码正常,怀疑当前数据混合数字与文字导致问题。

数据示例

| Trial #| 0|
| Unnamed: 0 | filename |
| 0 | 0.0 |
| 1 | 1.0 |
| 2 | 0.0 |
| 3 | 1.0 |
| 4 | 1.0 |
| 5 | 0.0|
| 6 | 1.0|
| 7 | 1.0|
| 8 | 1.0 |
| 9 | 1.0 |
| 10 | 0.0 |
| 11 | 0.0 |
| 12 | 1.0 |
| 13 | 0.0 |
| 14 | 1.0 |
| 15 | 1.0 |

当前代码

import os
import pandas as pd

folder_path = "/content/drive/session 1 & 2"

def find_column1_value_for_second_zero(file_path):
    try:
        df = pd.read_csv(file_path)
        consecutive_zeros = 0
        column1_value = None

        for _, row in df.iterrows():
            if row.iloc[1] == 0:
                consecutive_zeros += 1
                if consecutive_zeros == 2:
                    column1_value = row.iloc[0]
                    break
            else:
                consecutive_zeros = 0

        return column1_value
    except Exception as e:
        print(f"Error reading file '{file_path}': {str(e)}")
        return None

for filename in os.listdir(folder_path):
    if filename.endswith(".csv"):  # Assuming your files are CSV format
        file_path = os.path.join(folder_path, filename)
        
        column1_value = find_column1_value_for_second_zero(file_path)
        
        if column1_value is not None:
            print(f"In file '{filename}', the value in column 1 for the second zero in column 2 is: {column1_value}")
        else:
            print(f"In file '{filename}', no second zero in column 2 was found.")

预期与实际结果

  • 预期结果:返回索引1列第二次连续零对应的索引0列值,示例中应为11
  • 实际结果:所有文件均返回“未找到列2中的第二个连续零”

问题分析与修复方案

核心问题

  1. 表头读取错误:CSV文件存在两行表头,直接用pd.read_csv()会把第二行表头当成数据行,导致列索引混乱,后续数值读取异常。
  2. 类型匹配失效:数据中的零是0.0(浮点数),代码中用==0(整数)匹配;若因表头问题导致列类型变为object(混合文本与数字),直接比较会完全失效。

修复后的代码

import os
import pandas as pd

folder_path = "/content/drive/session 1 & 2"

def find_second_consecutive_zero(file_path):
    try:
        # 跳过第一行无效表头,用第二行作为列名,从第三行开始读数据
        df = pd.read_csv(file_path, skiprows=[0], header=0)
        df = df.reset_index(drop=True)
        
        consecutive_zeros = 0
        target_value = None
        
        for _, row in df.iterrows():
            # 处理混合类型:尝试转为浮点数,失败则视为非零
            try:
                val = float(row.iloc[1])
                is_zero = abs(val - 0) < 1e-9  # 兼容浮点数精度问题
            except (ValueError, TypeError):
                is_zero = False
            
            if is_zero:
                consecutive_zeros += 1
                if consecutive_zeros == 2:
                    target_value = row.iloc[0]
                    break
            else:
                consecutive_zeros = 0
        
        return target_value
    except Exception as e:
        print(f"处理文件 '{file_path}' 出错: {str(e)}")
        return None

for filename in os.listdir(folder_path):
    if filename.endswith(".csv"):
        file_path = os.path.join(folder_path, filename)
        result = find_second_consecutive_zero(file_path)
        
        if result is not None:
            print(f"文件 '{filename}' 中,第二次连续零对应的索引0列值为: {result}")
        else:
            print(f"文件 '{filename}' 未找到第二次连续零。")

关键修复点

  • 修正表头读取逻辑,跳过第一行无效内容,确保数据行正确加载
  • 增加混合类型处理逻辑,尝试将列值转为浮点数,避免文本干扰
  • 用浮点数精度比较替代直接相等判断,解决数值类型匹配问题

内容的提问来源于stack exchange,提问作者rpn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 15:57:40