读取多份CSV文件时出现列重复/异常列的原因排查及正常DataFrame生成方法咨询
读取多份CSV文件时出现列重复/异常列的原因排查及正常DataFrame生成方法咨询
我现在在处理30份CSV文件,这些文件的列结构不太统一——有的列是一样的,有的却各不相同。最开始我用这段代码读取并合并它们:
mycsvdir = r'C:\t\...\dict_full' csvfiles = glob.glob(os.path.join(mycsvdir, '*.csv')) dataframes = [] for csvfile in csvfiles: df = pd.read_csv(csvfile, encoding='UTF-16 LE', sep='\t') #, usecols=["event_rus", "event_category", "event_action"]) dataframes.append(df) result = pd.concat(dataframes, ignore_index=True) result.head()
运行后得到的DataFrame信息是这样的:
<class 'pandas.core.frame.DataFrame'> RangeIndex: 877 entries, 0 to 876 Data columns (total 51 columns):
而且生成的DataFrame结构特别怪异!
之后我尝试清理列名,还合并了一些类似的列:
result.columns = result.columns.str.replace(r' | |\s|\xa0|\-|depreciated', '', regex=True).str.lower() result.columns = result.columns.str.strip().str.replace(r' ', '') result['event_action'] = result[['eventaction', 'eventactiondeprecated']].apply(lambda row: row.dropna().iloc[0] if not row.dropna().empty else None, axis=1) result['event_category'] = result[['eventcategory', 'eventcategorydeprecated']].apply(lambda row: row.dropna().iloc[0] if not row.dropna().empty else None, axis=1) result['event_label'] = result[['eventlabel', 'propertyeventlabel', 'eventlabel']].apply(lambda row: row.dropna().iloc[0] if not row.dropna().empty else None, axis=1) ...
做完这些后,我删掉或者合并了一些没用的列,但还是存在列重复的问题,现在我整个人都懵了😂
我现在已经知道怎么合并特定列的内容(试过而且成功了),但我特别想搞明白:为什么会出现这些怪异的情况?有没有办法能直接生成结构正常的DataFrame?如果需要的话,我可以提供DataFrame的类型信息。
备注:内容来源于stack exchange,提问作者march_1
相关产品推荐
相关产品推荐

