You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于唯一值筛选DataFrame行?保留lot对应两shipment的cargotype唯一值数量相等的行

筛选满足同一Lot下各Shipment的Cargotype唯一值数量相等的行

完整解决方案代码

import pandas as pd

# 重建数据框
df = pd.DataFrame({
    "": [0,1,2,3,4,5,6,7,8,9,10],
    "lot": ["dfg", "dfg", "dfg","dfg","ghj","ghj","ghj","abc","abc","abc","abc"], 
    "shipment": ["a", "b", "a","b","c","d","d","e","f","e","e"], 
    "cargotype": ["adam", "chris", "bob","tom","chris","hanna","chris","charlie","king","su","min"], 
    "col4": [777, 775, 767,715,772,712,712, 123, 122, 121,120], 
    "col5": [13, 12, 13,12,14,12,12, 15, 16, 17,18], 
    "col6": [4, 3, 4,3, 5, 8,8, 7,7,0,0]
})

# 1. 计算每个(lot, shipment)组合的cargotype唯一值数量
nunique_counts = df.groupby(["lot", "shipment"])["cargotype"].nunique().reset_index()

# 2. 筛选出所有shipment的cargotype唯一值数量一致的lot
valid_lots = nunique_counts.groupby("lot")["cargotype"].nunique() == 1

# 3. 提取有效lot列表并过滤原始数据
finaldf = df[df["lot"].isin(valid_lots[valid_lots].index)]

print(finaldf)

输出结果

lot shipment cargotype  col4  col5  col6
0  dfg        a      adam   777    13     4
1  dfg        b     chris   775    12     3
2  dfg        a       bob   767    13     4
3  dfg        b       tom   715    12     3

分步说明

  • 统计分组唯一值:通过groupby(["lot", "shipment"])对数据分组,统计每组cargotype的唯一值数量,得到每个批次下不同运输单的唯一货型数。
  • 验证Lot有效性:对上述结果按lot二次分组,检查每个Lot下的唯一货型数是否只有一种(即所有运输单的唯一值数量相等)。
  • 过滤原始数据:用有效Lot列表过滤原始数据框,保留符合条件的所有行,同时保留col4-col6等无关列。

内容的提问来源于stack exchange,提问作者AAA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 15:10:25