You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas重复填充DataFrame至duration总和达标 代码无限循环问题求助

问题原因排查

你的代码出现无限循环的核心原因有两个:

  • 原始数据行数过少时,sample(frac=0.05) 会返回空DataFrame:比如你示例里只有3行数据,3*0.05=0.15,pandas采样时会向下取整得到0行,每次循环拼接空数据,duration总和永远不会增长,就会陷入无限循环
  • 极端情况如果采样到的所有行的duration值都是0,总和也不会增长,同样会触发无限循环
修复后代码
import pandas as pd

def fillPlaylist(df, target_duration):
    print("inside fill playlist fn.")
    if len(df) == 0:
        print("df len is 0, cannot fill.")
        return df
    
    # 提前过滤掉duration<=0的无效行,避免总和不增长
    receivedDf = df[df['duration'] > 0].reset_index(drop=True)
    if len(receivedDf) == 0:
        print("no valid data with positive duration, cannot fill.")
        return df
    
    print("receivedDf", receivedDf, flush=True)
    print("Received df len = ", len(receivedDf), flush=True)
    print("target duration to fill ", target_duration, flush=True)
    
    current_df = df.copy()
    while current_df['duration'].sum() < target_duration:
        print("filling")
        # 采样逻辑修改:优先按0.05比例采样,不足1行时强制采样1行
        sample_n = max(int(len(receivedDf) * 0.05), 1)
        randomSampleDuplicates = receivedDf.sample(n=sample_n, replace=True).reset_index(drop=True)
        # 如果需要把新数据加在表头,调换两个参数顺序即可
        current_df = pd.concat([current_df, randomSampleDuplicates], ignore_index=True)
        print("current df duration sum: ", current_df['duration'].sum())
    
    print("after filling df len = ", len(current_df))
    return current_df
关键修改说明
  • 修复了原代码里的变量拼写错误(ramdom改为random)
  • 提前过滤原始数据中duration小于等于0的无效行,避免采样到无效数据导致总和不增长
  • 把按比例采样改为max(比例计算行数, 1),保证每次循环至少采样1条数据,杜绝空采样的情况
  • 新增replace=True参数允许重复采样,避免原始数据行数不足时采样报错
  • 拼接时添加ignore_index=True重置行索引,避免索引重复的问题
  • 先拷贝原始df再操作,避免直接修改传入的原始数据

你给出的示例数据调用该方法后,输出完全符合你的预期效果。

内容的提问来源于stack exchange,提问作者R.singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 18:09:03