如何将DtypeWarning视为异常处理并按需读取CSV文件
解决pandas读取CSV时DtypeWarning的捕获与自动处理
Python的try-except确实只能捕获异常,没法直接处理警告。要捕获DtypeWarning,得用Python标准库的warnings模块,结合上下文管理器实现。下面是修改后的可行代码:
import os import warnings import pandas as pd def read_file(file_path): processed_files = [] for file in os.listdir(file_path): if file.endswith('.csv'): file_full_path = os.path.join(file_path, file) # 捕获警告上下文 with warnings.catch_warnings(record=True) as w: warnings.simplefilter("always") # 强制记录所有警告 csv_file = pd.read_csv(file_full_path) # 检查是否触发了目标警告 if any(issubclass(warn.category, pd.errors.DtypeWarning) for warn in w): # 重新读取并删除单位行 csv_file = pd.read_csv(file_full_path) csv_file = csv_file.drop(index=0).reset_index(drop=True) processed_files.append(csv_file) # 按需返回:返回最后一个文件或所有处理后的文件列表 return processed_files[-1] if processed_files else None
关键细节说明:
warnings.catch_warnings(record=True)会把期间触发的所有警告存入w列表warnings.simplefilter("always")确保即使是默认被忽略的警告也会被记录- 遍历警告列表,判断是否属于
DtypeWarning,触发则执行删首行操作 - 修复了原代码中路径变量不一致的问题,统一用函数参数传入路径
如果明确知道单位行的特征(比如首行是非数值内容),也可以提前判断避免触发警告,效率更高:
def read_file(file_path): processed_files = [] for file in os.listdir(file_path): if file.endswith('.csv'): file_full_path = os.path.join(file_path, file) # 读取首行判断是否为单位行 first_row = pd.read_csv(file_full_path, nrows=1) # 根据实际列名判断,比如col_A、col_B是否为非数值类型 if not pd.api.types.is_numeric_dtype(first_row['col_A']) or not pd.api.types.is_numeric_dtype(first_row['col_B']): csv_file = pd.read_csv(file_full_path, skiprows=[1]) # 跳过单位行(注意表头是第1行,单位行是第2行) else: csv_file = pd.read_csv(file_full_path) processed_files.append(csv_file) return processed_files[-1] if processed_files else None
内容的提问来源于stack exchange,提问作者serdar_bay
相关产品推荐
相关产品推荐

