You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas DataFrame指定列值扩展行并生成对应新列?

问题描述

我有如下所示的DataFrame:

import pandas as pd
df = pd.DataFrame({'Column 1': ['a', 'a', 'b', 'c'],
                  'Column 2': [2, 2, 3, 4],
                  'Column 3': [100, 110, 120, 130]}
                  )

打印效果如下:

Column 1 Column 2 Column 3
0 a 2 100
1 a 2 110
2 b 3 120
3 c 4 130

我需要得到如下格式的新DataFrame:

df = pd.DataFrame({'Column 1': ['a', 'a', 'a', 'a', 'b', 'b', 'b', 'c', 'c', 'c', 'c'],
                  'New Column': ['a1', 'a2', 'a3', 'a4', 'b1', 'b2', 'b3', 'c1', 'c2', 'c3', 'c4'],
                  'Column 3': [100, 100, 110, 110, 120, 120, 120, 130, 130, 130, 130]}
                  )

打印效果如下:

Column 1 New Column Column 3
0 a a1 100
1 a a2 100
2 a a3 110
3 a a4 110
4 b b1 120
5 b b2 120
6 b b3 120
7 c c1 130
8 c c2 130
9 c c3 130
10 c c4 130

我目前通过将Column 1和Column 3作为键分组,再结合两层iterrows循环实现了该需求,但运行耗时很长,请问有没有更高效的实现方式?

解决方案

你可以直接用pandas内置的向量化操作实现,全程不需要写任何Python层循环,性能远高于iterrows写法:

# 1. 按照Column 2的值将每行重复对应次数
res = df.repeat(df['Column 2']).reset_index(drop=True)
# 2. 按Column 1分组生成组内递增序号,拼接得到New Column
res['New Column'] = res['Column 1'] + res.groupby('Column 1').cumcount().add(1).astype(str)
# 3. 删除不需要的Column 2字段
res = res.drop('Column 2', axis=1)

实现逻辑说明

  • df.repeat(df['Column 2'])为pandas底层C实现的重复逻辑,会把每一行按照对应位置Column 2的值重复对应次数,生成的行数完全匹配需求
  • groupby('Column 1').cumcount()会按Column 1分组后生成从0开始的组内序号,加1后和Column 1的值拼接正好得到要求的New Column格式
  • 整个过程所有运算都在pandas底层完成,处理百万级数据也不会有性能瓶颈

内容的提问来源于stack exchange,提问作者macichocki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 06:39:00