如何在循环中合并Pandas DataFrame且不添加列后缀?
解决合并DataFrame时列名带后缀的问题并实现多级索引
你的问题根源在于逐步合并时,初始的full_df包含空的Magnetometer值,导致每次合并都新增列而非追加行。更好的做法是先收集所有传感器的DataFrame,再一次性拼接,最后构建多级索引。
修改后的代码
import os import pandas as pd from ai import cdas import datetime def magdf( start = datetime.datetime(2016, 1, 24, 0, 0, 0), end = datetime.datetime(2016, 1, 25, 0, 0, 0), maglist_a = ['upn', 'umq', 'gdh', 'atu', 'skt', 'ghb'], # Arctic magnetometers maglist_b = ['pg0', 'pg1', 'pg2', 'pg3', 'pg4', 'pg5'], # Antarctic magnetometers is_saved = False, ): # Magnetometer parameter dict so that we don't have to type the full string: d = {'Bx':'MAGNETIC_NORTH_-_H', 'By':'MAGNETIC_EAST_-_E','Bz':'VERTICAL_DOWN_-_Z'} d_i = dict((v, k) for k, v in d.items()) # inverted mapping for col renaming later if is_saved: fname = 'output/' + str(start) + '_' + '.csv' if os.path.exists(fname): print(f'Looks like {fname} has already been generated.') return # 改为收集所有传感器DataFrame的列表,而非逐步合并 sensor_dfs = [] for mags in [maglist_a, maglist_b]: for magname in mags: # For each magnetometer, pull data and add to list: print(f'Pulling data for magnetometer: {magname.upper()}') try: data = cdas.get_data( 'sp_phys', 'THG_L2_MAG_'+ magname.upper(), start, end, ['thg_mag_'+ magname] ) data['UT'] = pd.to_datetime(data['UT']) df = pd.DataFrame(data) df.rename(columns=d_i, inplace=True) # 重命名为简洁列名 # 添加传感器标识列 df['Magnetometer'] = magname.upper() # 移除多余的UT_1列(如果不需要的话) df.drop(columns=['UT_1'], inplace=True) sensor_dfs.append(df) except Exception as e: print(e) continue # 拼接所有传感器数据 full_df = pd.concat(sensor_dfs, ignore_index=True) # 统一时间精度为秒 full_df['UT'] = full_df['UT'].astype('datetime64[s]') # 构建多级索引:先按UT,再按Magnetometer full_df.set_index(['UT', 'Magnetometer'], inplace=True) # 可选:按索引排序 full_df.sort_index(inplace=True) if is_saved: print('Saving as a CSV.') full_df.to_csv(fname) return full_df foo = magdf(maglist_a=['upn'], maglist_b=['pg1']) print(foo)
关键修改点
- 替换逐步合并为列表收集+拼接:避免重复列名后缀问题,直接将所有传感器数据行追加在一起
- 移除多余列:删除不需要的
UT_1列(如果需要保留可以去掉这行) - 直接构建多级索引:通过
set_index(['UT', 'Magnetometer'])一步到位,满足你的后续需求 - 代码简化:去掉不必要的初始
full_df设置,逻辑更清晰
修改后的输出效果
最终的DataFrame会以UT和Magnetometer为多级索引,Bx/By/Bz为单列,示例如下:
Bx By Bz UT Magnetometer 2016-01-24 00:00:00 UPN 7841.68 320.623 55419.2 2016-01-24 00:00:00 PG1 10.25 -570.860 61.98 2016-01-24 00:00:01 PG1 10.50 -571.110 61.98 2016-01-24 00:01:00 UPN 7844.55 314.371 55418.7 ... ... ... ...
这样既解决了列名后缀问题,又直接得到了你需要的多级索引结构。
内容的提问来源于stack exchange,提问作者Iotatron
相关产品推荐
相关产品推荐

