如何在Pandas中合并两个DataFrame并实现指定分组展示形式
实现目标格式的方法
首先,你需要先合并两个DataFrame,再通过Pandas的样式工具(Styler)自定义显示效果,具体步骤如下:
1. 合并并整理数据结构
先使用concat合并两个DataFrame,调整列顺序并排序索引:
import pandas as pd # 模拟你的两个输入DataFrame df1 = pd.DataFrame( {'mean': [0, 2, 3], 'std': [4, 1, 2], 'epoch': [300, 300, 300]}, index=['file1', 'file2', 'file3'] ) df2 = pd.DataFrame( {'mean': [1, 3, 3], 'std': [4, 5, 6], 'epoch': [400, 400, 400]}, index=['file1', 'file2', 'file3'] ) # 合并DataFrame combined_df = pd.concat([df1, df2]) # 调整列顺序为你想要的epoch、mean、std combined_df = combined_df[['epoch', 'mean', 'std']] # 按索引排序,确保同一文件的记录连续 combined_df = combined_df.sort_index()
此时combined_df的数据结构已经符合逻辑,只是显示时会重复索引值,接下来处理显示样式。
2. 自定义样式隐藏重复索引
通过Pandas的Styler工具,编写函数隐藏重复的索引单元格内容:
def hide_duplicate_index(styler): # 标记重复的索引行 duplicate_mask = styler.index.duplicated() # 设置表格对齐样式 styler.set_properties(**{'text-align': 'center'}).set_table_styles([ {'selector': 'th.row_heading', 'props': [('text-align', 'left')]} ]) # 遍历行,隐藏重复索引的显示 for idx, is_duplicate in enumerate(duplicate_mask): if is_duplicate: styler.set_table_styles([ {'selector': f'tbody tr:nth-child({idx+1}) th', 'props': [('visibility', 'hidden')]} ], overwrite=False) return styler # 应用样式并显示结果 display(combined_df.style.pipe(hide_duplicate_index))
运行后就能得到你想要的显示效果:同一文件的重复索引会被隐藏,只在第一行显示文件名。
内容的提问来源于stack exchange,提问作者Diego Alejandro Gómez Pardo
相关产品推荐
相关产品推荐

