Python pandas:按步骤合并subtitle与dimension列生成新列
如何根据DataFrame的步骤列生成合并后的新列?
问题描述
我有一个记录用户行为步骤的DataFrame,每个步骤对应subtitle (step N)和dimension1 (step N)两列。示例数据如下:
import pandas as pd import numpy as np df = pd.DataFrame({'idVisit': [1, 2, 3], 'subtitle (step 0)': ['download', 'homepage', 'www.example.com'], 'dimension1 (step 0)': ['client', np.nan, 'internal'], 'subtitle (step 1)': ['pageview', 'pageview', 'map'], 'dimension1 (step 1)': ['client', 'client', np.nan], 'subtitle (step 2)': ['download', 'homepage', 'www.example.com'], 'dimension1 (step 2)': ['client', np.nan, 'internal'], 'subtitle (step 3)': ['pageview', 'pageview', 'map'], 'dimension1 (step 3)': ['client', 'client', np.nan]})
需要为每个步骤生成新列,规则是:
- 如果对应步骤的
dimension1列值不为空,就使用该值 - 如果
dimension1为空,则使用对应步骤的subtitle列值
比如step0列的预期输出是['client', 'homepage', 'internal'],step1列是['client', 'client', 'map'],以此类推。
解决方案
可以用Pandas的fillna()或combine_first()方法快速实现,这两个方法都能帮你用subtitle的值填充dimension1的空值,非常适合你的需求。
方法1:使用fillna()(直观易懂,适合新手)
fillna()的作用是用指定值填充列中的空值,这里我们直接用对应步骤的subtitle列来填充dimension1的空值:
# 遍历每个步骤(示例中是step0到step3,共4个步骤) for step_num in range(4): # 拼接当前步骤的列名 dim_col = f'dimension1 (step {step_num})' sub_col = f'subtitle (step {step_num})' # 生成新列,用subtitle填充dimension1的空值 df[f'step{step_num}'] = df[dim_col].fillna(df[sub_col]) # 查看处理后的核心结果 print(df[['idVisit', 'step0', 'step1', 'step2', 'step3']])
方法2:使用combine_first()
combine_first()会优先使用左边列的值,当左边列有空值时,用右边列的值补充,效果和fillna()一致:
for step_num in range(4): dim_col = f'dimension1 (step {step_num})' sub_col = f'subtitle (step {step_num})' df[f'step{step_num}'] = df[dim_col].combine_first(df[sub_col])
输出结果
运行代码后,生成的新列会符合你的预期:
| idVisit | step0 | step1 | step2 | step3 |
|---|---|---|---|---|
| 1 | client | client | client | client |
| 2 | homepage | client | homepage | client |
| 3 | internal | map | internal | map |
内容的提问来源于stack exchange,提问作者Beginner in the house
相关产品推荐
相关产品推荐

