You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中根据指定数据点查找最邻近匹配点

问题描述

现有数据集包含DATE_TIME、ID(含数字与字母)、VALUE1-VALUE4、MODEL、SOLD等多列。原代码通过按MODEL分组,利用scipy的distance_matrix计算SOLD值不同的组内VALUE1-VALUE4的最邻近点,但仅返回每组的一对匹配结果。需要修改代码,支持输入指定行索引或行数据,返回满足MODEL相同、SOLD不同条件下,VALUE1-VALUE4最邻近的匹配点。

修改后的实现代码

首先导入必要依赖库:

import numpy as np
import pandas as pd
from scipy.spatial import distance_matrix

定义核心函数,兼容两种输入方式:

def find_nearest_matching_row(input_data, df):
    # 处理输入:整数索引则取对应行,行数据直接使用
    if isinstance(input_data, int):
        target_row = df.loc[input_data]
    else:
        target_row = input_data
    
    # 提取目标行关键参数
    target_model = target_row['MODEL']
    target_sold = target_row['SOLD']
    target_values = target_row[['VALUE1', 'VALUE2', 'VALUE3', 'VALUE4']].to_numpy().reshape(1, -1)
    
    # 筛选符合条件的候选样本
    candidate_df = df[(df['MODEL'] == target_model) & (df['SOLD'] != target_sold)]
    
    if candidate_df.empty:
        return None  # 无匹配样本时返回None
    
    # 计算距离矩阵并找到最近点
    candidate_values = candidate_df[['VALUE1', 'VALUE2', 'VALUE3', 'VALUE4']].to_numpy()
    dist_matrix = distance_matrix(target_values, candidate_values)
    min_dist_idx = dist_matrix.argmin()
    
    return candidate_df.iloc[min_dist_idx]
使用示例
  1. 通过行索引查找:
# 查找索引为15的行的最邻近匹配点
nearest_row = find_nearest_matching_row(15, df)
print(nearest_row)
  1. 通过行数据查找:
# 取数据集某一行作为输入
target_row = df.loc[30]
nearest_row = find_nearest_matching_row(target_row, df)
print(nearest_row)
关键逻辑说明
  • 输入兼容:自动识别输入类型(索引/行数据),统一转换为目标行对象处理
  • 精准筛选:严格过滤出与目标行MODEL一致、SOLD相反的样本,确保匹配规则符合需求
  • 距离计算:借助distance_matrix计算欧氏距离,快速定位距离最小的匹配点
  • 边界处理:无符合条件样本时返回None,避免程序报错

内容的提问来源于stack exchange,提问作者dspractician

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:50:22