You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pivot后的DataFrame获取非NaN值?优化距离查询逻辑

解决两个问题的优化方案

一、处理pivot表.loc返回NaN的问题

你的pivot表应该是对称结构(比如距离类数据),只填充了三角区域,导致部分行列组合查询返回NaN。直接让矩阵对称化就能彻底解决:

import pandas as pd

# 让pivot表双向填充,缺失的一侧用另一侧的值补全
piv = piv.combine_first(piv.T)

之后不管用piv.loc["gamma", "c"]还是piv.loc["c", "gamma"],都会返回正确的非NaN值,无需额外判断。

如果不想修改原表,也可以写个轻量的取值函数:

def get_pivot_value(piv_table, row, col):
    val = piv_table.loc[row, col]
    if pd.isna(val):
        val = piv_table.loc[col, row]
    return val

调用get_pivot_value(piv, "gamma", "c")就能直接拿到目标值。

二、优化对称距离对的查询(替代try-except)

从长格式的reduced_distance_df查询(id1,id2)或(id2,id1)的距离,最简洁高效的方式是统一配对顺序,避免重复判断:

方法1:排序配对做索引(推荐,查询效率最高)

先预处理数据,把每个id对按固定顺序排序后作为索引:

# 生成排序后的配对元组
reduced_distance_df['sorted_pair'] = reduced_distance_df.apply(
    lambda x: tuple(sorted((x['id1'], x['id2']))), axis=1
)
# 用配对元组做索引,只保留距离列
distance_index = reduced_distance_df.set_index('sorted_pair')['distance']

之后查询时,不管输入顺序如何,先排序配对再查:

def get_distance(a, b):
    target_pair = tuple(sorted((a, b)))
    return distance_index.get(target_pair, None)  # 找不到可自定义默认值

# 示例调用,两种顺序结果一致
get_distance("gamma", "c")
get_distance("c", "gamma")

这种方式的查询是O(1)级别,比try-except或循环判断高效得多。

方法2:布尔筛选(无需预处理)

如果不想修改原数据,也可以直接用布尔条件筛选:

def get_distance(df, a, b):
    # 匹配两种顺序的id对
    match_mask = ((df['id1'] == a) & (df['id2'] == b)) | ((df['id1'] == b) & (df['id2'] == a))
    result = df.loc[match_mask, 'distance']
    return result.iloc[0] if not result.empty else None

这种方式不用预处理,但每次查询会遍历全表,数据量大时效率不如方法1。


内容的提问来源于stack exchange,提问作者Antonio Carnevali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 07:50:22