Python实现同ID组内任意行满足条件时整组统一标记值
Python实现按ID分组打标方案
处理表格类数据优先用pandas库实现,代码简洁运行效率高,核心逻辑是先按ID分组判断是否存在符合条件的记录,再将分组结果批量映射回全表,避免逐行遍历的冗余操作。
基于pandas的实现(推荐,适配绝大多数表格处理场景)
如果未安装pandas,先执行安装命令:pip install pandas
完整可运行示例代码:
import pandas as pd # 1. 加载表格数据 # 示例为手动构造的测试数据,实际使用时替换为自己的读取逻辑即可,比如pd.read_csv("数据文件.csv")、pd.read_excel("数据文件.xlsx") df = pd.DataFrame([ {"ID": 1, "social score": 620, "other_col": "记录1"}, {"ID": 1, "social score": 0, "other_col": "记录2"}, {"ID": 1, "social score": 580, "other_col": "记录3"}, {"ID": 2, "social score": 650, "other_col": "记录4"}, {"ID": 2, "social score": 710, "other_col": "记录5"}, ]) # 2. 按ID分组,统计每个ID是否存在满足条件的记录 # 示例条件为social score == 0,可根据实际业务修改判断逻辑 id_abnormal_map = df.groupby("ID")["social score"].apply(lambda x: (x == 0).any()).to_dict() # 3. 给全表所有行赋值class字段 df["class"] = df["ID"].map(id_abnormal_map).map({True: "abnormal", False: "normal"}) # 打印验证结果 print(df)
运行后输出结果完全匹配需求:ID=1的所有行class均为abnormal,ID=2的所有行class均为normal。
如果判断逻辑更复杂,只要修改
lambda内的判断规则即可。比如要判断「social score小于300且逾期次数大于2」,可以把分组逻辑改成df.groupby("ID").apply(lambda x: ((x["social score"] < 300) & (x["overdue_cnt"] > 2)).any()).to_dict()。
原生Python实现(无第三方库依赖,适合极小数据量场景)
如果不想安装第三方依赖,可以用两次遍历的方式实现:第一次遍历统计每个ID是否存在异常记录,第二次遍历给全量行打标。
# 示例数据,格式为列表套字典,和原生读取csv、json的返回结构一致 table_data = [ {"ID": 1, "social score": 620, "other_col": "记录1"}, {"ID": 1, "social score": 0, "other_col": "记录2"}, {"ID": 1, "social score": 580, "other_col": "记录3"}, {"ID": 2, "social score": 650, "other_col": "记录4"}, {"ID": 2, "social score": 710, "other_col": "记录5"}, ] # 第一次遍历:记录每个ID是否存在异常 id_status = {} for row in table_data: current_id = row["ID"] # 已经标记为异常的ID无需重复判断 if id_status.get(current_id) is True: continue # 示例判断条件为social score等于0,可按需修改 if row["social score"] == 0: id_status[current_id] = True # 第二次遍历:给所有行打class标签 for row in table_data: row["class"] = "abnormal" if id_status.get(row["ID"], False) else "normal" # 打印验证结果 for row in table_data: print(row)
内容的提问来源于stack exchange,提问作者Eunsoo Ko
相关产品推荐
相关产品推荐

