如何以Pythonic方式实现Pandas按次数重复DataFrame的UserID?
解决方案
可以直接利用 pandas 内置的 repeat() 方法实现,这是完全向量化的操作,无需手动写循环,非常符合Pythonic风格:
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({ 'UserID': ['abc123', 'def234'], 'num_attempts': [4, 3] }) # 生成目标DataFrame result_df = pd.DataFrame(df['UserID'].repeat(df['num_attempts']), columns=['result_col']) print(result_df)
代码说明:
df['UserID'].repeat(df['num_attempts']):对UserID列的每个元素,按照对应行num_attempts的值进行重复,返回一个Series- 把这个Series传入
pd.DataFrame(),并指定列名为result_col,就得到了目标格式的DataFrame
运行后输出结果:
result_col 0 abc123 0 abc123 0 abc123 0 abc123 1 def234 1 def234 1 def234
如果不需要保留原始索引,可以加上.reset_index(drop=True)来重置索引:
result_df = pd.DataFrame( df['UserID'].repeat(df['num_attempts']), columns=['result_col'] ).reset_index(drop=True)
此时输出的索引会变成连续序列:
result_col 0 abc123 1 abc123 2 abc123 3 abc123 4 def234 5 def234 6 def234
内容的提问来源于stack exchange,提问作者FlyingPickle
相关产品推荐
相关产品推荐

