You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于欧氏距离与质心计算将含坐标的字典列表按每批20个分组

坐标邻近度分批次处理方案

需求说明

现有字典列表结构数据,每个字典的键为item_cd,值为二维坐标location_coordinates,需基于欧氏距离计算坐标邻近度,将数据按每20条为单位分批次,同批次数据位置尽可能邻近。

输入数据

[
 {5036885850: [92.0, 88.73]}, {5036885955: [90.0, 61.73]}, 
 {5036885984: [86.0, 73.03]}, {5036885998: [102.0, 77.54]}, 
 {5036885851: [93.0, 88.0]}, {5036885956: [91.0, 66.73]}, {5036885984: [87.0, 70.0]},
 {5036885998: [101.0, 70.54]},{5036885812: [45.0, 88.73]}, {5036885955: [76.0, 60.73]},
 {5036885911: [83.0, 74.03]}, {5036885910: [108.0, 77.54]},
 {5036885850: [89.0, 76.73]},
 {5036885800: [80.0, 69.45]},
 {50368854801: [86.0, 69.50]},
 {5036885802: [102.0, 77.54]},
 {5036885809: [92.5, 85.0]},
 {5036885803: [91.5, 65.73]},
 {5036885850: [78.0, 76.73]},
 {5036885800: [77.0, 69.45]},
 {50368854801: [85.0, 69.50]},
 {5036885802: [101.50, 89.23]},
 {5036885809: [100.5, 84.84]},
 {5036885803: [100.67, 64.23]},
]

实现规则

    1. 单批次初始计算以原点(0,0)为基准,取距离最近的项加入当前批次
    1. 后续计算以当前批次已加入所有项的质心为基准,取剩余数据中距离最近的项加入当前批次
    1. 单批次最大容量为20条,满额后自动生成新批次,新批次重新从原点开始计算邻近度
    1. 所有数据处理完成后,不足20条的剩余数据单独作为最后一个批次

输出格式要求

每个批次为独立的字典列表,参考示例如下:

[{5036885955: [90.0, 61.73]}, {5036885984: [86.0, 73.03]}, {5036885998: [102.0, 77.54]}, {5036885850: [92.0, 88.73]}]

完整可运行代码

import numpy as np

def calculate_centroid(lst):
    arr = np.array(lst)
    length = arr.shape[0]
    if length == 0:
        return np.array((0, 0))
    sum_x = np.sum(arr[:, 0])
    sum_y = np.sum(arr[:, 1])
    return np.array((sum_x/float(length), sum_y/float(length)))

# 待处理原始数据
raw_data = [
    {5036885850: [92.0, 88.73]}, {5036885955: [90.0, 61.73]},
    {5036885984: [86.0, 73.03]}, {5036885998: [102.0, 77.54]},
    {5036885851: [93.0, 88.0]}, {5036885956: [91.0, 66.73]}, {5036885984: [87.0, 70.0]},
    {5036885998: [101.0, 70.54]}, {5036885812: [45.0, 88.73]}, {5036885955: [76.0, 60.73]},
    {5036885911: [83.0, 74.03]}, {5036885910: [108.0, 77.54]},
    {5036885850: [89.0, 76.73]},
    {5036885800: [80.0, 69.45]},
    {50368854801: [86.0, 69.50]},
    {5036885802: [102.0, 77.54]},
    {5036885809: [92.5, 85.0]},
    {5036885803: [91.5, 65.73]},
    {5036885850: [78.0, 76.73]},
    {5036885800: [77.0, 69.45]},
    {50368854801: [85.0, 69.50]},
    {5036885802: [101.50, 89.23]},
    {5036885809: [100.5, 84.84]},
    {5036885803: [100.67, 64.23]},
]

BATCH_SIZE = 20
remaining_data = raw_data.copy()
all_batches = []

while remaining_data:
    current_batch = []
    current_batch_coords = []
    current_centroid = np.array((0, 0))
    
    while len(current_batch) < BATCH_SIZE and remaining_data:
        # 计算所有剩余项到当前质心的距离
        coords = [list(item.values())[0] for item in remaining_data]
        dist_list = [np.linalg.norm(current_centroid - np.array(coord)) for coord in coords]
        min_idx = dist_list.index(min(dist_list))
        
        # 取出最近项加入当前批次
        selected_item = remaining_data.pop(min_idx)
        current_batch.append(selected_item)
        current_batch_coords.append(coords[min_idx])
        
        # 更新质心
        current_centroid = calculate_centroid(current_batch_coords)
    
    all_batches.append(current_batch)

# 输出结果
for idx, batch in enumerate(all_batches, 1):
    print(f"批次{idx}:")
    print(batch)
    print("-"*50)

代码修改说明

  • 修正了原代码的拼写错误:calcualate_centroid改为calculate_centroid
  • 新增批次循环逻辑,每满20条自动生成新批次并重置质心计算基准
  • 移除了原代码中未定义的变量引用,优化了入参和变量命名的可读性
  • 增加了结果格式化输出,可直接查看每个批次的内容

内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 01:39:02