You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在循环中合并Pandas DataFrame且不添加列后缀?

解决合并DataFrame时列名带后缀的问题并实现多级索引

你的问题根源在于逐步合并时,初始的full_df包含空的Magnetometer值,导致每次合并都新增列而非追加行。更好的做法是先收集所有传感器的DataFrame,再一次性拼接,最后构建多级索引。

修改后的代码

import os
import pandas as pd
from ai import cdas
import datetime

def magdf(
    start = datetime.datetime(2016, 1, 24, 0, 0, 0), 
    end = datetime.datetime(2016, 1, 25, 0, 0, 0), 
    maglist_a = ['upn', 'umq', 'gdh', 'atu', 'skt', 'ghb'],  # Arctic magnetometers
    maglist_b = ['pg0', 'pg1', 'pg2', 'pg3', 'pg4', 'pg5'],  # Antarctic magnetometers
    is_saved = False, 
):
    # Magnetometer parameter dict so that we don't have to type the full string: 
    d = {'Bx':'MAGNETIC_NORTH_-_H', 'By':'MAGNETIC_EAST_-_E','Bz':'VERTICAL_DOWN_-_Z'}
    d_i = dict((v, k) for k, v in d.items()) # inverted mapping for col renaming later
    
    if is_saved:
        fname = 'output/' + str(start) + '_' + '.csv'
        if os.path.exists(fname):
            print(f'Looks like {fname} has already been generated.')
            return 
    
    # 改为收集所有传感器DataFrame的列表,而非逐步合并
    sensor_dfs = []
    
    for mags in [maglist_a, maglist_b]:
        for magname in mags:   # For each magnetometer, pull data and add to list:
            print(f'Pulling data for magnetometer: {magname.upper()}')
            try:                
                data = cdas.get_data(
                    'sp_phys',
                    'THG_L2_MAG_'+ magname.upper(),
                    start,
                    end,
                    ['thg_mag_'+ magname]
                )
                data['UT'] = pd.to_datetime(data['UT'])
                df = pd.DataFrame(data)
                df.rename(columns=d_i, inplace=True)    # 重命名为简洁列名
                
                # 添加传感器标识列
                df['Magnetometer'] = magname.upper()
                
                # 移除多余的UT_1列(如果不需要的话)
                df.drop(columns=['UT_1'], inplace=True)
                
                sensor_dfs.append(df)
                
            except Exception as e:
                print(e)
                continue
    
    # 拼接所有传感器数据
    full_df = pd.concat(sensor_dfs, ignore_index=True)
    
    # 统一时间精度为秒
    full_df['UT'] = full_df['UT'].astype('datetime64[s]')
    
    # 构建多级索引:先按UT,再按Magnetometer
    full_df.set_index(['UT', 'Magnetometer'], inplace=True)
    
    # 可选:按索引排序
    full_df.sort_index(inplace=True)
    
    if is_saved:
        print('Saving as a CSV.')
        full_df.to_csv(fname)
    
    return full_df

foo = magdf(maglist_a=['upn'], maglist_b=['pg1'])
print(foo)

关键修改点

  • 替换逐步合并为列表收集+拼接:避免重复列名后缀问题,直接将所有传感器数据行追加在一起
  • 移除多余列:删除不需要的UT_1列(如果需要保留可以去掉这行)
  • 直接构建多级索引:通过set_index(['UT', 'Magnetometer'])一步到位,满足你的后续需求
  • 代码简化:去掉不必要的初始full_df设置,逻辑更清晰

修改后的输出效果

最终的DataFrame会以UT和Magnetometer为多级索引,Bx/By/Bz为单列,示例如下:

Bx       By       Bz
UT                  Magnetometer                        
2016-01-24 00:00:00 UPN          7841.68  320.623  55419.2
2016-01-24 00:00:00 PG1            10.25 -570.860    61.98
2016-01-24 00:00:01 PG1            10.50 -571.110    61.98
2016-01-24 00:01:00 UPN          7844.55  314.371  55418.7
...                               ...      ...      ...

这样既解决了列名后缀问题,又直接得到了你需要的多级索引结构。

内容的提问来源于stack exchange,提问作者Iotatron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 05:35:12