You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对文件列表调用数据清洗函数?解决传入列表类型报错问题

解决批量文件数据清洗的报错问题

问题根源

你的clean_data函数设计为接收单个文件路径字符串,但你直接传入了整个文件路径列表file_list,而pd.read_csv无法识别列表类型的输入,因此抛出Invalid file path or buffer object type: <class 'list'>错误。


解决方案1:遍历文件列表逐个处理

这是最直接的修复方式,循环遍历每个文件路径,单独调用清洗函数:

import pandas as pd

# 生成文件列表(优化:split()默认分割任意空白字符,避免多余空格产生空值)
files = "/data1.txt /data2.txt"
file_list = files.split()

# 批量清洗并收集结果
cleaned_datasets = []
for file_path in file_list:
    cleaned_df = clean_data(file_path)
    cleaned_datasets.append(cleaned_df)

# 可选:合并所有清洗后的数据集
combined_df = pd.concat(cleaned_datasets, ignore_index=True)

解决方案2:修改函数支持批量输入

如果希望直接传入列表调用函数,可以修改clean_data,让它同时兼容单个路径和路径列表:

def clean_data(file_input):
    # 处理单个文件路径
    if isinstance(file_input, str):
        df = pd.read_csv(file_input, sep="\n|\\t", decimal='.', names=['nM','I'], engine='python', header=None)
        data_start = df.loc[df['nM'].str.startswith('#DATA')].index[0] + 1
        return df[data_start:].reset_index(drop=True)
    # 处理文件路径列表
    elif isinstance(file_input, list):
        cleaned_dfs = []
        for file_path in file_input:
            df = pd.read_csv(file_path, sep="\n|\\t", decimal='.', names=['nM','I'], engine='python', header=None)
            data_start = df.loc[df['nM'].str.startswith('#DATA')].index[0] + 1
            cleaned_dfs.append(df[data_start:].reset_index(drop=True))
        return cleaned_dfs
    else:
        raise TypeError("输入必须是单个文件路径字符串或文件路径列表")

此时直接调用clean_data(file_list)即可得到所有清洗后的DataFrame列表。


额外优化建议

  1. 修复文件列表生成的潜在问题:
    如果原始字符串存在多余空格,split(" ")会生成空字符串,建议用过滤逻辑:
    file_list = [f for f in files.split(" ") if f]
    
  2. 增加异常处理:
    避免文件不存在、未找到#DATA标记等情况导致程序崩溃:
    def clean_data(file_path):
        try:
            df = pd.read_csv(file_path, sep="\n|\\t", decimal='.', names=['nM','I'], engine='python', header=None)
            # 检查#DATA标记是否存在
            data_marker = df.loc[df['nM'].str.startswith('#DATA', na=False)]
            if data_marker.empty:
                raise ValueError(f"文件 {file_path} 中未找到#DATA标记")
            data_start = data_marker.index[0] + 1
            return df[data_start:].reset_index(drop=True)
        except FileNotFoundError:
            print(f"错误:文件 {file_path} 不存在")
            return None
        except Exception as e:
            print(f"处理文件 {file_path} 时出错:{str(e)}")
            return None
    

内容的提问来源于stack exchange,提问作者J.V.D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 04:50:21