You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向含列表值的DataFrame列高效应用坐标偏移?

问题描述

现有一个Pandas DataFrame,其中bb_box列存储着list类型的边界框坐标,示例如下:

bb_box
0    [4, 565, 1088, 591]
1   [17, 820, 1092, 949]
2    [5, 746, 1084, 796]
3   [32, 240, 1104, 263]
4    [0, 187, 1111, 212]
...

需要对bb_box列中每个列表的数值应用常量偏移(比如offset_coords = [10,10,10,10]),尝试用向量化方法实现时遇到错误:

编写的代码:

def align_coordinates(align_coord, box):
    """ align_coordinates
        align_coord    A list of 4 values to add to the box coordinates
        box            The box coordinates that need to be aligned 
    """
    for idx, v in enumerate(align_coord):
        box[idx] = box[idx] + v

    return box

offset_coords = [10,10,10,10]
df['bb_box'] = np.vectorize(align_coordinates)(offset_coords, df['bb_box'])

报错信息:

ValueError: operands could not be broadcast together with shapes (4,) (5,)

目前采用非向量化的循环方式实现:

offset_coords = [10,10,10,10]
for i, v in df.iterrows():
    r_box = df.at[i,'bb_box']
    r_box = np.add(r_box, offset_coords)
    df.at[i,'bb_box'] = r_box

询问是否有更优的实现方式,如何向DataFrame列应用常量偏移列表。


优化实现方案

方法1:数组批量运算(性能最优)

将bb_box列的列表转为二维numpy数组,直接和偏移数组做元素级加法,再转回列表。这是真正的向量化操作,性能远超循环或apply:

import numpy as np

offset_coords = np.array([10,10,10,10])
# 把bb_box列转为二维数组
bbox_array = np.array(df['bb_box'].tolist())
# 批量添加偏移
bbox_array += offset_coords
# 转回列表赋值给原列
df['bb_box'] = bbox_array.tolist()

方法2:apply简化循环

如果不想转数组,用apply结合np.add可以让代码更简洁,性能略优于手动iterrows循环:

offset_coords = [10,10,10,10]
df['bb_box'] = df['bb_box'].apply(lambda x: np.add(x, offset_coords).tolist())

原代码报错原因

np.vectorize并非真正的向量化工具,只是对元素做循环的包装。你的代码中,offset_coords是长度为4的列表,df['bb_box']是包含5个长度为4的列表的Series,np.vectorize在处理时会尝试将offset_coords作为整体传递给每个box,导致内部形状匹配失败,从而报错。

内容的提问来源于stack exchange,提问作者Zexelon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 01:57:37