You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字典值比较DataFrame元组并过滤匹配行的技术问询

问题描述

我有一个包含元组列Col1和Col3的DataFrame,示例如下:

Check Col1        Col2 Col3
0    NaN  (F123, 1)    R  (F123, 2)
0    NaN  (F123, 1)    R  (F123, 6)
0    NaN  (F123, 1)    R  (F123, 7)

现有字典:

df1d = {('F123', 1): 'R', ('F123', 2): 'O', ('F123', 6): 'R'}

我想实现两个需求:

  1. 遍历行比较Col1和Col3元组对应的字典值,若值相同则在Check列输出Bad,否则保持NaN。我尝试的代码如下,但没达到预期:
def check_color(dictionary, value, adj_values, df):
    # for x in adj_values:
    if tuple(x) in [k for k, v in dictionary.items() if v == dictionary[value]]:
        df.set_value('Check', 'Bad')
        return
for index, row in df.iterrows():
    check_color(df1d, row['Col1'], row['Col3'], df)

期望输出:

Check Col1        Col2 Col3
0    NaN  (F123, 1)    R  (F123, 2)
0    Bad  (F123, 1)    R  (F123, 6)
0    NaN  (F123, 1)    R  (F123, 7)
  1. 如何过滤掉Col1和Col3不匹配的行?我尝试了df[(df['ConnectorAndPin'] == df['Adj.'])](注:这里是笔误,实际要比较的是Col1和Col3)。

解决方案

一、实现Check列标记

你的原始代码存在几个问题:比如set_value已经被pandas弃用,用iterrows遍历行效率较低,同时没处理元组不在字典中的情况(比如示例里的(F123,7)不在df1d中,直接取dictionary[value]会报错)。

推荐用矢量化操作来实现,既高效又简洁:

import pandas as pd
import numpy as np

# 构造示例DataFrame
data = {
    'Check': [np.nan, np.nan, np.nan],
    'Col1': [('F123', 1), ('F123', 1), ('F123', 1)],
    'Col2': ['R', 'R', 'R'],
    'Col3': [('F123', 2), ('F123', 6), ('F123', 7)]
}
df = pd.DataFrame(data)

df1d = {('F123', 1): 'R', ('F123', 2): 'O', ('F123', 6): 'R'}

# 先获取Col1和Col3对应的字典值,不在字典中的会返回NaN
col1_values = df['Col1'].map(df1d)
col3_values = df['Col3'].map(df1d)

# 比较两个值是否相等且都不为NaN,满足则标记为'Bad'
df['Check'] = np.where(
    (col1_values == col3_values) & col1_values.notna() & col3_values.notna(),
    'Bad',
    np.nan
)

print(df)

运行后就能得到你想要的输出:

Check      Col1 Col2      Col3
0   NaN  (F123, 1)    R  (F123, 2)
1   Bad  (F123, 1)    R  (F123, 6)
2   NaN  (F123, 1)    R  (F123, 7)

二、过滤行

你提到的“过滤掉Col1和Col3不匹配的行”,分两种常见场景处理:

场景1:过滤掉Col1和Col3元组本身不相等的行(仅保留元组完全相同的行)

直接比较两列即可:

# 保留Col1和Col3元组相等的行
filtered_df = df[df['Col1'] == df['Col3']]

不过你的示例中没有元组相等的行,所以结果会是空DataFrame。

场景2:过滤掉Check列为Bad的行(去掉字典值相同的行)

如果需求是去掉标记为Bad的行,可以这样写:

# 过滤掉Check列为'Bad'的行
filtered_df = df[df['Check'].isna()]

运行后得到的结果是:

Check      Col1 Col2      Col3
0   NaN  (F123, 1)    R  (F123, 2)
2   NaN  (F123, 1)    R  (F123, 7)

内容的提问来源于stack exchange,提问作者MaxB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:15:50