You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现同ID组内任意行满足条件时整组统一标记值

Python实现按ID分组打标方案

处理表格类数据优先用pandas库实现,代码简洁运行效率高,核心逻辑是先按ID分组判断是否存在符合条件的记录,再将分组结果批量映射回全表,避免逐行遍历的冗余操作。

基于pandas的实现(推荐,适配绝大多数表格处理场景)

如果未安装pandas,先执行安装命令:
pip install pandas

完整可运行示例代码:

import pandas as pd

# 1. 加载表格数据
# 示例为手动构造的测试数据,实际使用时替换为自己的读取逻辑即可,比如pd.read_csv("数据文件.csv")、pd.read_excel("数据文件.xlsx")
df = pd.DataFrame([
    {"ID": 1, "social score": 620, "other_col": "记录1"},
    {"ID": 1, "social score": 0, "other_col": "记录2"},
    {"ID": 1, "social score": 580, "other_col": "记录3"},
    {"ID": 2, "social score": 650, "other_col": "记录4"},
    {"ID": 2, "social score": 710, "other_col": "记录5"},
])

# 2. 按ID分组,统计每个ID是否存在满足条件的记录
# 示例条件为social score == 0,可根据实际业务修改判断逻辑
id_abnormal_map = df.groupby("ID")["social score"].apply(lambda x: (x == 0).any()).to_dict()

# 3. 给全表所有行赋值class字段
df["class"] = df["ID"].map(id_abnormal_map).map({True: "abnormal", False: "normal"})

# 打印验证结果
print(df)

运行后输出结果完全匹配需求:ID=1的所有行class均为abnormal,ID=2的所有行class均为normal。

如果判断逻辑更复杂,只要修改lambda内的判断规则即可。比如要判断「social score小于300且逾期次数大于2」,可以把分组逻辑改成df.groupby("ID").apply(lambda x: ((x["social score"] < 300) & (x["overdue_cnt"] > 2)).any()).to_dict()。

原生Python实现(无第三方库依赖,适合极小数据量场景)

如果不想安装第三方依赖,可以用两次遍历的方式实现:第一次遍历统计每个ID是否存在异常记录,第二次遍历给全量行打标。

# 示例数据,格式为列表套字典,和原生读取csv、json的返回结构一致
table_data = [
    {"ID": 1, "social score": 620, "other_col": "记录1"},
    {"ID": 1, "social score": 0, "other_col": "记录2"},
    {"ID": 1, "social score": 580, "other_col": "记录3"},
    {"ID": 2, "social score": 650, "other_col": "记录4"},
    {"ID": 2, "social score": 710, "other_col": "记录5"},
]

# 第一次遍历:记录每个ID是否存在异常
id_status = {}
for row in table_data:
    current_id = row["ID"]
    # 已经标记为异常的ID无需重复判断
    if id_status.get(current_id) is True:
        continue
    # 示例判断条件为social score等于0,可按需修改
    if row["social score"] == 0:
        id_status[current_id] = True

# 第二次遍历:给所有行打class标签
for row in table_data:
    row["class"] = "abnormal" if id_status.get(row["ID"], False) else "normal"

# 打印验证结果
for row in table_data:
    print(row)

内容的提问来源于stack exchange,提问作者Eunsoo Ko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:57:19