批量修改CSV文件表头并读取时遇解析错误求助
问题:批量处理CSV文件时出现“No columns to parse from file”错误
需处理30多个格式混乱的CSV文件,分布在两个子文件夹中。单个文件手动将[Data]表头改为['x','y'],并跳过指定行、忽略最后一行可正常读取,但自行编写的Python批量处理函数触发“No columns to parse from file”错误。
示例CSV文件内容
#Name, #Comment,"" #ExtComment,"" #Source, [Data] 1,2 3,4 5,6 #[END_OF_FILE]
编写的批量处理函数
#sets - refers to the set containing the name of each file (i.e. [file1, file2]) #df - the dataframe which you are going to store the data in #dataLabels - the headers you want to search for within the .csv file #skip - the number of rows you want to skip #newHeader - what you want to change the column headers to be #pathName - provide path where files are located def reader (sets, df, dataLabels, skip, newHeader, pathName): for i in range(len(sets)): df_temp = pd.read_csv(glob.glob(pathName+ sets[i]+".csv"), sep=r'\s*,', skiprows = skip, engine = 'python')[:-1] df_temp.column.value[0] = [newHeader] for j in range(len(dataLabels)): df_temp[dataLabels[j]] = pd.to_numeric(df_temp[dataLabels[j]],errors = 'coerce') df.append(df_temp) return df
问题分析与修复方案
核心错误点
- 文件路径传入错误:
glob.glob()返回列表,而pd.read_csv()需单个文件路径,直接传入列表会导致解析失败。 - 固定跳过行数不通用:硬编码
skiprows无法适配不同文件中[Data]行的位置,可能跳过所有有效数据行。 - 表头修改写法错误:
df_temp.column.value[0] = [newHeader]是无效语法,无法正确替换表头。 - 已弃用API使用:
df.append()在新版pandas中已被弃用,易引发兼容性问题。
修复后的通用函数
import pandas as pd import glob def reader(sets, dataLabels, newHeader, pathName): df_list = [] for filename in sets: # 获取单个文件路径,避免glob返回列表的问题 file_path = glob.glob(f"{pathName}{filename}.csv")[0] # 动态查找[Data]行的位置,适配不同文件格式 with open(file_path, 'r') as f: skip_rows = 0 for line in f: if line.strip() == '[Data]': break skip_rows += 1 # 读取数据:跳过[Data]之前的行,将[Data]作为表头,忽略最后一行 df_temp = pd.read_csv( file_path, sep=r'\s*,', skiprows=skip_rows, header=0, engine='python' )[:-1] # 替换为自定义表头 df_temp.columns = newHeader # 将指定列转为数值类型 for col in dataLabels: df_temp[col] = pd.to_numeric(df_temp[col], errors='coerce') df_list.append(df_temp) # 合并所有文件的数据 return pd.concat(df_list, ignore_index=True)
使用说明
- 函数不再需要传入初始
df,直接返回合并后的完整DataFrame。 - 动态查找
[Data]行位置,无需手动指定跳过行数,适配所有同格式文件。 - 支持自定义任意表头名称,满足通用性需求。
- 用
pd.concat()替代已弃用的df.append(),保证新版本pandas兼容性。
额外优化建议
- 若文件分布在多级子文件夹,可改用
glob.glob(f"{pathName}/**/*.csv", recursive=True)批量获取所有CSV路径,无需手动传入sets列表。 - 添加异常处理逻辑,避免单个文件解析失败导致整个批量任务中断。
内容的提问来源于stack exchange,提问作者bigmac42
相关产品推荐
相关产品推荐

