You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化多变量全组合构建DataFrame的代码?

优化多变量全组合租金预测的DataFrame构建效率

问题背景

需要构建包含city、occupation、income、predicted_rent列的DataFrame,覆盖三个变量的所有组合用于预测租金。当前采用嵌套for循环实现,但处理约15万行数据耗时超24小时,且因各列表长度不等无法使用zip,需无嵌套循环的优化方案。

优化方案

1. 用itertools.product生成笛卡尔积(替代嵌套循环)

itertools.product是Python标准库中基于C实现的工具,能高效生成多列表的全组合,比纯Python嵌套循环速度快得多,代码也更简洁。

示例代码:

import itertools
import pandas as pd

# 示例变量列表
city = ["NYC", "SF", "Chicago"]
occupation = ["SWE", "Doctor", "Teacher"]
income = ["<50K", "50-100K", "100-200K"]

# 生成所有变量的笛卡尔积
all_combinations = list(itertools.product(city, occupation, income))

# 批量计算预测租金(若predict_rent可批量处理,效率还能进一步提升)
predicted_rents = [predict_rent(c, o, i) for c, o, i in all_combinations]

# 直接转换为DataFrame
df = pd.DataFrame(all_combinations, columns=["city", "occupation", "income"])
df["predicted_rent"] = predicted_rents

2. 利用Pandas的交叉连接(Cross Join)生成全组合

通过Pandas的merge方法实现交叉连接,直接生成包含所有组合的DataFrame,再批量计算预测值。

示例代码:

import pandas as pd

# 单变量DataFrame
df_city = pd.DataFrame({"city": city})
df_occupation = pd.DataFrame({"occupation": occupation})
df_income = pd.DataFrame({"income": income})

# 交叉连接生成全组合(通过临时key实现)
df = df_city.assign(key=1)\
            .merge(df_occupation.assign(key=1), on="key")\
            .merge(df_income.assign(key=1), on="key")\
            .drop("key", axis=1)

# 计算预测租金
df["predicted_rent"] = df.apply(lambda row: predict_rent(row["city"], row["occupation"], row["income"]), axis=1)

3. 并行计算核心预测逻辑(最大幅度提速)

耗时的核心是predict_rent函数,利用多核CPU并行计算可大幅缩短总时间。推荐使用concurrent.futures.ProcessPoolExecutor(适用于CPU密集型任务)或ThreadPoolExecutor(适用于IO密集型任务)。

示例代码:

import itertools
import pandas as pd
from concurrent.futures import ProcessPoolExecutor

# 生成全组合
all_combinations = list(itertools.product(city, occupation, income))

# 并行计算预测租金,max_workers设为CPU核心数(或根据资源调整)
with ProcessPoolExecutor(max_workers=4) as executor:
    predicted_rents = list(executor.map(lambda x: predict_rent(*x), all_combinations))

# 构建DataFrame
df = pd.DataFrame(all_combinations, columns=["city", "occupation", "income"])
df["predicted_rent"] = predicted_rents

额外建议

  • 如果predict_rent支持向量化输入(例如接受数组而非单个值),可以直接将整个列传入计算,这是效率最高的方式,完全避免循环。
  • 若数据量极大,可考虑分批次处理,避免内存占用过高。

内容的提问来源于stack exchange,提问作者codenoodles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 04:52:34