基于其他DataFrame为行值添加前缀(Python实现)
问题描述
我有两个DataFrame:
第一个(记为df_scores):
CallId ScoreId ScoreName Weight 19 4198451 180 First 1.0 110 4198451 348 Second 3.0 124 4198451 15 Third 1.0
第二个(记为df_components):
CallId ScoreComponentName ScoreId Weight 58 4198451 Dissat 180 1.0 59 4198451 Escalation 180 0.0 60 4198451 Repeat 180 0.0 61 4198451 Transfer 180 0.0 363 4198451 Account 348 0.0 364 4198451 Activation 348 0.0 375 4198451 Categories 15 0.0
需要根据每个ScoreId对应的ScoreName,为ScoreComponentName列的每个值添加前缀,得到如下结果:
CallId ScoreComponentName ScoreId Weight 58 4198451 First.Dissat 180 1.0 59 4198451 First.Escalation 180 0.0 60 4198451 First.Repeat 180 0.0 61 4198451 First.Transfer 180 0.0 363 4198451 Second.Account 348 0.0 364 4198451 Second.Activation 348 0.0 375 4198451 Third.Categories 15 0.0
解决方案
用Pandas的关联操作和字符串拼接就能实现,步骤如下:
- 先从第一个DataFrame里提取
ScoreId与ScoreName的映射关系,避免关联时引入冗余列:
score_map = df_scores[['ScoreId', 'ScoreName']]
- 把第二个DataFrame和这个映射表按
ScoreId做左关联,确保每个组件都能匹配到对应的评分名称:
merged_df = df_components.merge(score_map, on='ScoreId', how='left')
- 用
.连接ScoreName和原ScoreComponentName,替换掉原来的列:
merged_df['ScoreComponentName'] = merged_df['ScoreName'] + '.' + merged_df['ScoreComponentName']
- 删掉多余的
ScoreName列,就是你要的结果:
final_df = merged_df.drop('ScoreName', axis=1)
嫌步骤太散的话,也可以写成一行简洁代码:
final_df = df_components.merge(df_scores[['ScoreId', 'ScoreName']], on='ScoreId')\ .assign(ScoreComponentName=lambda x: x['ScoreName'] + '.' + x['ScoreComponentName'])\ .drop('ScoreName', axis=1)
内容的提问来源于stack exchange,提问作者Hey_jude
相关产品推荐
相关产品推荐

