You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量修改CSV文件表头并读取时遇解析错误求助

问题:批量处理CSV文件时出现“No columns to parse from file”错误

需处理30多个格式混乱的CSV文件,分布在两个子文件夹中。单个文件手动将[Data]表头改为['x','y'],并跳过指定行、忽略最后一行可正常读取,但自行编写的Python批量处理函数触发“No columns to parse from file”错误。

示例CSV文件内容

#Name,
#Comment,""
#ExtComment,""
#Source,
[Data]
1,2
3,4
5,6
#[END_OF_FILE]

编写的批量处理函数

#sets - refers to the set containing the name of each file (i.e. [file1, file2])
#df - the dataframe which you are going to store the data in 
#dataLabels - the headers you want to search for within the .csv file
#skip - the number of rows you want to skip
#newHeader - what you want to change the column headers to be
#pathName - provide path where files are located

def reader (sets, df, dataLabels, skip, newHeader, pathName):
     for i in range(len(sets)):
        
        df_temp = pd.read_csv(glob.glob(pathName+ sets[i]+".csv"), sep=r'\s*,', skiprows = skip, engine = 'python')[:-1] 
        df_temp.column.value[0] = [newHeader]
        for j in range(len(dataLabels)):
           df_temp[dataLabels[j]] = pd.to_numeric(df_temp[dataLabels[j]],errors = 'coerce')       
        df.append(df_temp)       
     return df

问题分析与修复方案

核心错误点

  1. 文件路径传入错误:glob.glob()返回列表,而pd.read_csv()需单个文件路径,直接传入列表会导致解析失败。
  2. 固定跳过行数不通用:硬编码skiprows无法适配不同文件中[Data]行的位置,可能跳过所有有效数据行。
  3. 表头修改写法错误:df_temp.column.value[0] = [newHeader]是无效语法,无法正确替换表头。
  4. 已弃用API使用:df.append()在新版pandas中已被弃用,易引发兼容性问题。

修复后的通用函数

import pandas as pd
import glob

def reader(sets, dataLabels, newHeader, pathName):
    df_list = []
    for filename in sets:
        # 获取单个文件路径,避免glob返回列表的问题
        file_path = glob.glob(f"{pathName}{filename}.csv")[0]
        
        # 动态查找[Data]行的位置,适配不同文件格式
        with open(file_path, 'r') as f:
            skip_rows = 0
            for line in f:
                if line.strip() == '[Data]':
                    break
                skip_rows += 1
        
        # 读取数据:跳过[Data]之前的行,将[Data]作为表头,忽略最后一行
        df_temp = pd.read_csv(
            file_path,
            sep=r'\s*,',
            skiprows=skip_rows,
            header=0,
            engine='python'
        )[:-1]
        
        # 替换为自定义表头
        df_temp.columns = newHeader
        
        # 将指定列转为数值类型
        for col in dataLabels:
            df_temp[col] = pd.to_numeric(df_temp[col], errors='coerce')
        
        df_list.append(df_temp)
    
    # 合并所有文件的数据
    return pd.concat(df_list, ignore_index=True)

使用说明

  1. 函数不再需要传入初始df,直接返回合并后的完整DataFrame。
  2. 动态查找[Data]行位置,无需手动指定跳过行数,适配所有同格式文件。
  3. 支持自定义任意表头名称,满足通用性需求。
  4. 用pd.concat()替代已弃用的df.append(),保证新版本pandas兼容性。

额外优化建议

  • 若文件分布在多级子文件夹,可改用glob.glob(f"{pathName}/**/*.csv", recursive=True)批量获取所有CSV路径,无需手动传入sets列表。
  • 添加异常处理逻辑,避免单个文件解析失败导致整个批量任务中断。

内容的提问来源于stack exchange,提问作者bigmac42

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 08:24:17