如何在Python中合并Pandas DataFrame中相同id的行?
如何合并Pandas DataFrame中相同id的行?
输入DataFrame
| id | test1 | test2 |
|---|---|---|
| one | 10 | |
| one | 30 | |
| two | 3 | |
| three | 10 | 5 |
期望结果
| id | test1 | test2 |
|---|---|---|
| one | 10 | 30 |
| two | 3 | |
| three | 10 | 5 |
解决方案
你可以使用groupby结合first()方法实现,该方法会为每个分组的每一列取第一个非空值,完全匹配你的需求:
- 先构造示例DataFrame(假设空值对应Pandas中的
NaN):
import pandas as pd import numpy as np df = pd.DataFrame({ 'id': ['one', 'one', 'two', 'three'], 'test1': [10, np.nan, np.nan, 10], 'test2': [np.nan, 30, 3, 5] })
- 执行合并操作:
merged_df = df.groupby('id', as_index=False).first()
说明
groupby('id'):按照id列对DataFrame分组as_index=False:保留id作为普通列,而非索引first():对每个分组的每一列,跳过空值取第一个有效数据
如果同一分组的同一列存在多个非空值,可根据业务需求替换first()为其他聚合函数,比如sum()求和、max()取最大值等。
内容的提问来源于stack exchange,提问作者user12217822
相关产品推荐
相关产品推荐

