You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从pandas dataframe中提取x1/y1/x2/y2唯一且置信度最高的行

实现方法

方法1:排序后去重(简单直观)

先按置信度降序排序,保证每个坐标组中置信度最高的行排在最前,再按四个坐标字段去重,保留首行即可:

import pandas as pd

# 示例数据构造(可替换为自己的df导入代码)
data = {
    'x1': [238.288834, 238.288834, 238.288834, 248.977844, 248.977844, 15.0],
    'y1': [118.716125, 118.716125, 118.716125, 115.054123, 115.054123, 10.0],
    'x2': [300.878754, 300.878754, 300.878754, 321.307007, 321.307007, 2298.9],
    'y2': [137.672791, 137.672791, 137.672791, 141.315460, 141.315460, 187.0],
    'confidence': [0.885205, 0.881469, 0.879645, 0.876451, 0.872008, 0.70],
    'class': [0.0, 1.0, 5.0, 0.0, 1.0, 0.0]
}
df = pd.DataFrame(data)

# 核心处理代码
# 按置信度降序排序
df_sorted = df.sort_values('confidence', ascending=False)
# 按四个坐标列去重,保留第一个出现的行(即置信度最高的)
result = df_sorted.drop_duplicates(subset=['x1', 'y1', 'x2', 'y2'], keep='first')
# 重置索引(可选,按需求决定是否保留)
result = result.reset_index(drop=True)

方法2:分组取最大值索引(逻辑更直接)

直接对四个坐标字段分组,取每组置信度最大值对应的行索引,筛选即可:

result = df.loc[df.groupby(['x1', 'y1', 'x2', 'y2'])['confidence'].idxmax()].reset_index(drop=True)

两种方法输出结果都和预期一致,数据量较大时推荐使用方法2,不需要全局排序性能更优。

内容的提问来源于stack exchange,提问作者Data_User

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 07:45:02