排查CSV中第二组连续零对应列值的Python代码故障
问题:CSV文件中查找第二次连续零对应的值失败
问题背景
需要遍历指定文件夹内的所有CSV文件,找到索引1列中第二次连续出现的零,并返回对应行的索引0列的值。但现有代码运行后始终提示未找到目标值,手动核对数据确认存在该情况,之前使用纯数字数据时代码正常,怀疑当前数据混合数字与文字导致问题。
数据示例
| Trial #| 0|
| Unnamed: 0 | filename |
| 0 | 0.0 |
| 1 | 1.0 |
| 2 | 0.0 |
| 3 | 1.0 |
| 4 | 1.0 |
| 5 | 0.0|
| 6 | 1.0|
| 7 | 1.0|
| 8 | 1.0 |
| 9 | 1.0 |
| 10 | 0.0 |
| 11 | 0.0 |
| 12 | 1.0 |
| 13 | 0.0 |
| 14 | 1.0 |
| 15 | 1.0 |
当前代码
import os import pandas as pd folder_path = "/content/drive/session 1 & 2" def find_column1_value_for_second_zero(file_path): try: df = pd.read_csv(file_path) consecutive_zeros = 0 column1_value = None for _, row in df.iterrows(): if row.iloc[1] == 0: consecutive_zeros += 1 if consecutive_zeros == 2: column1_value = row.iloc[0] break else: consecutive_zeros = 0 return column1_value except Exception as e: print(f"Error reading file '{file_path}': {str(e)}") return None for filename in os.listdir(folder_path): if filename.endswith(".csv"): # Assuming your files are CSV format file_path = os.path.join(folder_path, filename) column1_value = find_column1_value_for_second_zero(file_path) if column1_value is not None: print(f"In file '{filename}', the value in column 1 for the second zero in column 2 is: {column1_value}") else: print(f"In file '{filename}', no second zero in column 2 was found.")
预期与实际结果
- 预期结果:返回索引1列第二次连续零对应的索引0列值,示例中应为
11 - 实际结果:所有文件均返回“未找到列2中的第二个连续零”
问题分析与修复方案
核心问题
- 表头读取错误:CSV文件存在两行表头,直接用
pd.read_csv()会把第二行表头当成数据行,导致列索引混乱,后续数值读取异常。 - 类型匹配失效:数据中的零是
0.0(浮点数),代码中用==0(整数)匹配;若因表头问题导致列类型变为object(混合文本与数字),直接比较会完全失效。
修复后的代码
import os import pandas as pd folder_path = "/content/drive/session 1 & 2" def find_second_consecutive_zero(file_path): try: # 跳过第一行无效表头,用第二行作为列名,从第三行开始读数据 df = pd.read_csv(file_path, skiprows=[0], header=0) df = df.reset_index(drop=True) consecutive_zeros = 0 target_value = None for _, row in df.iterrows(): # 处理混合类型:尝试转为浮点数,失败则视为非零 try: val = float(row.iloc[1]) is_zero = abs(val - 0) < 1e-9 # 兼容浮点数精度问题 except (ValueError, TypeError): is_zero = False if is_zero: consecutive_zeros += 1 if consecutive_zeros == 2: target_value = row.iloc[0] break else: consecutive_zeros = 0 return target_value except Exception as e: print(f"处理文件 '{file_path}' 出错: {str(e)}") return None for filename in os.listdir(folder_path): if filename.endswith(".csv"): file_path = os.path.join(folder_path, filename) result = find_second_consecutive_zero(file_path) if result is not None: print(f"文件 '{filename}' 中,第二次连续零对应的索引0列值为: {result}") else: print(f"文件 '{filename}' 未找到第二次连续零。")
关键修复点
- 修正表头读取逻辑,跳过第一行无效内容,确保数据行正确加载
- 增加混合类型处理逻辑,尝试将列值转为浮点数,避免文本干扰
- 用浮点数精度比较替代直接相等判断,解决数值类型匹配问题
内容的提问来源于stack exchange,提问作者rpn
相关产品推荐
相关产品推荐

