You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取多份CSV文件时出现列重复/异常列的原因排查及正常DataFrame生成方法咨询

读取多份CSV文件时出现列重复/异常列的原因排查及正常DataFrame生成方法咨询

我现在在处理30份CSV文件,这些文件的列结构不太统一——有的列是一样的,有的却各不相同。最开始我用这段代码读取并合并它们:

mycsvdir = r'C:\t\...\dict_full'

csvfiles = glob.glob(os.path.join(mycsvdir, '*.csv'))

dataframes = [] 
for csvfile in csvfiles:
    df = pd.read_csv(csvfile, encoding='UTF-16 LE', sep='\t') #, usecols=["event_rus", "event_category", "event_action"])
    dataframes.append(df)

result = pd.concat(dataframes, ignore_index=True)

result.head()

运行后得到的DataFrame信息是这样的:

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 877 entries, 0 to 876
Data columns (total 51 columns):

而且生成的DataFrame结构特别怪异!

之后我尝试清理列名,还合并了一些类似的列:

result.columns = result.columns.str.replace(r' |  |\s|\xa0|\-|depreciated', '', regex=True).str.lower()
result.columns = result.columns.str.strip().str.replace(r' ', '')

result['event_action'] = result[['eventaction', 'eventactiondeprecated']].apply(lambda row: row.dropna().iloc[0] if not row.dropna().empty else None, axis=1)
result['event_category'] = result[['eventcategory', 'eventcategorydeprecated']].apply(lambda row: row.dropna().iloc[0] if not row.dropna().empty else None, axis=1)
result['event_label'] = result[['eventlabel', 'propertyeventlabel', 'eventlabel']].apply(lambda row: row.dropna().iloc[0] if not row.dropna().empty else None, axis=1)
...

做完这些后,我删掉或者合并了一些没用的列,但还是存在列重复的问题,现在我整个人都懵了😂

我现在已经知道怎么合并特定列的内容(试过而且成功了),但我特别想搞明白:为什么会出现这些怪异的情况?有没有办法能直接生成结构正常的DataFrame?如果需要的话,我可以提供DataFrame的类型信息。

备注:内容来源于stack exchange,提问作者march_1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 15:14:28