Pandas重复填充DataFrame至duration总和达标 代码无限循环问题求助
问题原因排查
你的代码出现无限循环的核心原因有两个:
- 原始数据行数过少时,
sample(frac=0.05)会返回空DataFrame:比如你示例里只有3行数据,3*0.05=0.15,pandas采样时会向下取整得到0行,每次循环拼接空数据,duration总和永远不会增长,就会陷入无限循环 - 极端情况如果采样到的所有行的duration值都是0,总和也不会增长,同样会触发无限循环
修复后代码
import pandas as pd def fillPlaylist(df, target_duration): print("inside fill playlist fn.") if len(df) == 0: print("df len is 0, cannot fill.") return df # 提前过滤掉duration<=0的无效行,避免总和不增长 receivedDf = df[df['duration'] > 0].reset_index(drop=True) if len(receivedDf) == 0: print("no valid data with positive duration, cannot fill.") return df print("receivedDf", receivedDf, flush=True) print("Received df len = ", len(receivedDf), flush=True) print("target duration to fill ", target_duration, flush=True) current_df = df.copy() while current_df['duration'].sum() < target_duration: print("filling") # 采样逻辑修改:优先按0.05比例采样,不足1行时强制采样1行 sample_n = max(int(len(receivedDf) * 0.05), 1) randomSampleDuplicates = receivedDf.sample(n=sample_n, replace=True).reset_index(drop=True) # 如果需要把新数据加在表头,调换两个参数顺序即可 current_df = pd.concat([current_df, randomSampleDuplicates], ignore_index=True) print("current df duration sum: ", current_df['duration'].sum()) print("after filling df len = ", len(current_df)) return current_df
关键修改说明
- 修复了原代码里的变量拼写错误(
ramdom改为random) - 提前过滤原始数据中duration小于等于0的无效行,避免采样到无效数据导致总和不增长
- 把按比例采样改为
max(比例计算行数, 1),保证每次循环至少采样1条数据,杜绝空采样的情况 - 新增
replace=True参数允许重复采样,避免原始数据行数不足时采样报错 - 拼接时添加
ignore_index=True重置行索引,避免索引重复的问题 - 先拷贝原始df再操作,避免直接修改传入的原始数据
你给出的示例数据调用该方法后,输出完全符合你的预期效果。
内容的提问来源于stack exchange,提问作者R.singh
相关产品推荐
相关产品推荐

