Python pandas DataFrame嵌套循环实现逐行计算的问题求助
解决Pandas DataFrame嵌套循环生成true_move和price_new的问题
首先,你代码里的内层循环写法有几个小问题需要修正:
- 列名对应错误:你的数据列是
max_move_%,但代码里写的是max_move range的参数格式不对:range(start, stop)是左闭右开的,如果你要包含max_move_%的正数,得把stop设为max_move + 1- 获取当前行值的方式错误:应该用
row['max_move_%'],而不是df_merged['max_move'][row]
基于你现有iterrows的修正版本
如果你想继续用iterrows的方式,下面是可以正常运行的代码:
import pandas as pd # 你的示例数据 df_merged = pd.DataFrame({ 'product': [1], 'price': [100], 'max_move_%': [10] }) # 用来存储所有结果行的列表 result = [] # 外层循环遍历每一行 for _, row in df_merged.iterrows(): current_max = row['max_move_%'] # 内层循环:从 -current_max 到 current_max,包含两端值 for true_move in range(-current_max, current_max + 1): # 计算新价格:按照示例,100对应-10得到90,公式是 price * (1 + true_move/100) price_new = row['price'] * (1 + true_move / 100) # 把当前行信息和新生成的列打包成字典,加入结果列表 result.append({ 'product': row['product'], 'price': row['price'], 'max_move_%': row['max_move_%'], 'true_move': true_move, 'price_new': price_new }) # 把结果列表转成DataFrame final_df = pd.DataFrame(result) print(final_df)
运行后会生成包含true_move从-10到10的所有行,对应正确的price_new值。
更高效的Pandas风格实现
不过要注意,iterrows在处理大数据集时效率很低,推荐用Pandas原生的向量化操作或者explode方法,代码更简洁且速度更快:
方法1:用explode展开列表
import pandas as pd df_merged = pd.DataFrame({ 'product': [1], 'price': [100], 'max_move_%': [10] }) # 给每行生成包含所有true_move值的列表 df_merged['true_move'] = df_merged['max_move_%'].apply(lambda x: list(range(-x, x + 1))) # 展开列表列,把每个列表元素变成单独的行 final_df = df_merged.explode('true_move', ignore_index=True) # 计算price_new列 final_df['price_new'] = final_df['price'] * (1 + final_df['true_move'] / 100) # 调整列顺序(可选) final_df = final_df[['product', 'price', 'max_move_%', 'true_move', 'price_new']] print(final_df)
方法2:用apply生成子DataFrame再合并
def expand_single_row(row): max_move = row['max_move_%'] # 生成所有true_move值 true_moves = range(-max_move, max_move + 1) # 构造当前行对应的扩展DataFrame return pd.DataFrame({ 'product': row['product'], 'price': row['price'], 'max_move_%': row['max_move_%'], 'true_move': true_moves, 'price_new': row['price'] * (1 + pd.Series(true_moves)/100) }) # 对每行应用函数,合并所有子DataFrame final_df = pd.concat(df_merged.apply(expand_single_row, axis=1).tolist(), ignore_index=True) print(final_df)
这两种方法都避免了显式的嵌套循环,更符合Pandas的设计理念,数据量越大优势越明显。
内容的提问来源于stack exchange,提问作者HeadOverFeet
相关产品推荐
相关产品推荐

