You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于权重与距离的聚类优化:按Score比例规划Agent每日零售路线

问题与优化方案

原问题分析

原聚类算法仅基于经纬度做约束性K-Means,存在以下问题:

  • 未考虑score列,聚类结果完全随机
  • 无法保证不同score样本的分配比例
  • 未控制不同score样本间的距离,可能出现路线跨度过大的情况

需求明确:

  1. 每个Agent的22个零售商严格按指定比例(如{3:0.7,2:0.2,1:0.1})分配不同score的样本
  2. 以最高score样本的聚类质心为基准,其他score样本的质心与基准距离需低于阈值(如1km),最多尝试5次选取最近质心

优化后的代码实现

import pandas as pd
import numpy as np
import math
from tqdm import tqdm
from sklearn.cluster import KMeans
from k_means_constrained import KMeansConstrained

def haversine_distance(lat1, lon1, lat2, lon2):
    """计算两点间的哈弗辛距离(单位:千米)"""
    # 转换为弧度
    lat1_rad = np.radians(lat1)
    lon1_rad = np.radians(lon1)
    lat2_rad = np.radians(lat2)
    lon2_rad = np.radians(lon2)
    
    # 哈弗辛公式计算实际距离
    dlat = lat2_rad - lat1_rad
    dlon = lon2_rad - lon1_rad
    a = np.sin(dlat/2)**2 + np.cos(lat1_rad) * np.cos(lat2_rad) * np.sin(dlon/2)**2
    c = 2 * np.arcsin(np.sqrt(a))
    r = 6371  # 地球平均半径(千米)
    return c * r

def clustering_with_score_constraint(df, route_size=22, score_ratio={3:0.7,2:0.2,1:0.1}, distance_threshold=1, max_attempts=5):
    df_with_cluster = pd.DataFrame()
    
    # 按Agent分组处理每个个体的路线规划
    for agent in tqdm(df['Agents'].unique()):
        temp = df.loc[df['Agents'] == agent].copy().reset_index(drop=True)
        agent_total = len(temp)
        if agent_total == 0:
            continue
        
        # 1. 计算各score需要抽取的样本数量,确保总数为route_size
        score_counts = {}
        total_required = route_size
        for score, ratio in score_ratio.items():
            score_counts[score] = round(total_required * ratio)
        
        # 修正总数误差,将差值补到占比最高的score类别
        total_calculated = sum(score_counts.values())
        if total_calculated != total_required:
            max_score = max(score_ratio, key=score_ratio.get)
            score_counts[max_score] += (total_required - total_calculated)
        
        selected_samples = pd.DataFrame()
        base_centroid = None
        
        # 2. 按score从高到低处理,优先确定基准质心
        sorted_scores = sorted(score_ratio.keys(), reverse=True)
        for idx, score in enumerate(sorted_scores):
            required_count = score_counts[score]
            score_samples = temp.loc[temp['score'] == score].copy()
            
            # 处理当前score无样本的情况,从其他高score类别补
            if len(score_samples) == 0:
                for s in sorted_scores:
                    if s != score and len(temp.loc[temp['score'] == s]) >= required_count:
                        score_samples = temp.loc[temp['score'] == s].copy()
                        break
                if len(score_samples) == 0:
                    score_samples = temp.copy()
            
            # 非最高score样本,需满足距离基准质心的阈值要求
            if idx > 0 and base_centroid is not None:
                # 计算每个样本到基准质心的距离
                score_samples['distance_to_base'] = score_samples.apply(
                    lambda x: haversine_distance(x['latitude'], x['longitude'], base_centroid[0], base_centroid[1]),
                    axis=1
                )
                # 筛选符合距离要求的样本
                filtered_samples = score_samples.loc[score_samples['distance_to_base'] <= distance_threshold]
                
                # 若数量不足,尝试最多max_attempt次选取最近的样本
                attempt = 0
                while len(filtered_samples) < required_count and attempt < max_attempts:
                    attempt += 1
                    filtered_samples = score_samples.sort_values('distance_to_base').head(required_count)
                
                # 仍不足则取所有可用样本
                if len(filtered_samples) < required_count:
                    filtered_samples = score_samples.copy()
                
                selected = filtered_samples.sample(n=min(required_count, len(filtered_samples)), random_state=42)
            else:
                # 最高score样本,直接抽取并计算基准质心
                selected = score_samples.sample(n=min(required_count, len(score_samples)), random_state=42)
                # 计算质心(单样本则直接用自身坐标)
                if len(selected) > 1:
                    kmeans = KMeans(n_clusters=1, random_state=42)
                    kmeans.fit(selected[['latitude', 'longitude']])
                    base_centroid = kmeans.cluster_centers_[0]
                else:
                    base_centroid = (selected['latitude'].iloc[0], selected['longitude'].iloc[0])
            
            selected_samples = pd.concat([selected_samples, selected])
        
        # 3. 确保最终样本数严格等于route_size
        if len(selected_samples) > route_size:
            selected_samples = selected_samples.sample(n=route_size, random_state=42)
        elif len(selected_samples) < route_size:
            # 从剩余样本中补充距离基准质心最近的点
            remaining = temp.loc[~temp.index.isin(selected_samples.index)].copy()
            remaining['distance_to_base'] = remaining.apply(
                lambda x: haversine_distance(x['latitude'], x['longitude'], base_centroid[0], base_centroid[1]),
                axis=1
            )
            remaining_sorted = remaining.sort_values('distance_to_base')
            need = route_size - len(selected_samples)
            selected_samples = pd.concat([selected_samples, remaining_sorted.head(need)])
        
        # 4. 标记聚类信息
        selected_samples['assigned_cluster'] = 0
        selected_samples['agent_cluster'] = f"{agent}_0"
        
        df_with_cluster = pd.concat([df_with_cluster, selected_samples])
    
    return df_with_cluster

关键优化点说明

  • 比例精确控制:严格按指定比例计算各score样本数量,自动修正总数误差,确保最终样本数为22
  • 基准质心约束:优先处理最高score样本并确定基准质心,后续样本必须满足距离阈值要求,最多尝试5次选取最近样本
  • 异常场景处理:针对某score无样本、数量不足的情况,自动从其他高score类别补全,优先保证核心比例要求
  • 确定性保证:设置固定random_state避免随机结果,确保每次运行输出一致
  • 真实距离计算:使用哈弗辛公式计算经纬度间的实际地表距离,保证距离阈值的准确性

使用示例

# 假设df包含Agents、latitude、longitude、score四列
result_df = clustering_with_score_constraint(
    df,
    route_size=22,
    score_ratio={3:0.7,2:0.2,1:0.1},
    distance_threshold=1,
    max_attempts=5
)

内容的提问来源于stack exchange,提问作者Esmael Maher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 00:35:23