You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas iterrows运行过慢,如何对发送到BQ的处理代码做向量化优化

DataFrame遍历转字典列表优化方案

以下4种方案均可直接替换原有iterrows逻辑,性能从低到高排序:

方案1:Pandas apply 实现

改动最小,直接将逐行遍历逻辑封装为行处理函数即可:

def send_to_bq(self, df):
    def _process_row(row):
        return {
            "processing_timestamp": str(row["processing_timestamp"]),
            "id": row["id"],
            "embeddings_vector": [str(x) for x in row["vectors"]],
        }
    # 仅提取需要的列后按行apply
    result = df[["id", "vectors", "processing_timestamp"]].apply(_process_row, axis=1).tolist()
    # 后续原有逻辑保持不变

方案2:Pandas 列向量化实现

优先对整列做批量处理,减少行级遍历的开销,性能比apply高3~5倍:

def send_to_bq(self, df):
    # 先单独批量处理每一列,都是列级操作,无行遍历开销
    processed_df = df[["id", "vectors", "processing_timestamp"]].copy()
    processed_df["processing_timestamp"] = processed_df["processing_timestamp"].astype(str)
    processed_df["embeddings_vector"] = processed_df["vectors"].map(lambda vec: [str(x) for x in vec])
    # 直接转字典列表,Pandas底层优化过该操作
    result = processed_df[["id", "processing_timestamp", "embeddings_vector"]].to_dict("records")
    # 后续原有逻辑保持不变

方案3:Numpy 向量化实现

将Pandas列转为Numpy数组后遍历,去掉Pandas行对象的封装开销,性能比Pandas向量化再高1~2倍:

import numpy as np

def send_to_bq(self, df):
    # 转Numpy数组,减少中间对象开销
    ids = df["id"].to_numpy()
    timestamps = df["processing_timestamp"].astype(str).to_numpy()
    vectors = df["vectors"].to_numpy()
    
    result = []
    for id_val, ts, vec in zip(ids, timestamps, vectors):
        result.append({
            "processing_timestamp": ts,
            "id": id_val,
            "embeddings_vector": [str(x) for x in vec]
        })
    # 后续原有逻辑保持不变

方案4:Swifter 自动优化实现

适合超大规模数据,Swifter会自动判断最优执行方式,自动启用Dask/Ray并行处理,数据量越大优势越明显:

import swifter

def send_to_bq(self, df):
    def _process_row(row):
        return {
            "processing_timestamp": str(row["processing_timestamp"]),
            "id": row["id"],
            "embeddings_vector": [str(x) for x in row["vectors"]],
        }
    # swifter自动选择最优执行策略
    result = df[["id", "vectors", "processing_timestamp"]].swifter.apply(_process_row, axis=1).tolist()
    # 后续原有逻辑保持不变

额外优化建议

如果最终是要写入BigQuery,不需要自己构造字典列表,可以直接用google-cloud-bigquery客户端的load_table_from_dataframe方法直接导入DataFrame到BQ表,省去整个转换环节,性能提升最明显。

内容的提问来源于stack exchange,提问作者Guilherme Giuliano Nicolau

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 18:27:05