基于条件与计算为Pandas DataFrame创建新列
按ID分组生成长度计算规则的实现方案
原始数据
import pandas as pd df = pd.DataFrame({ "ID": [1, 1, 2, 2, 2, 3, 3], "length": [0.7, 0.7, 0.8, 0.6, 0.6, 0.9, 0.9], "comment": ["typed", "handwritten", "typed", "typed", "handwritten", "handwritten", "handwritten"] })
原始数据预览:
| ID | length | comment |
|---|---|---|
| 1 | 0.7 | typed |
| 1 | 0.7 | handwritten |
| 2 | 0.8 | typed |
| 2 | 0.6 | typed |
| 2 | 0.6 | handwritten |
| 3 | 0.9 | handwritten |
| 3 | 0.9 | handwritten |
需求说明
对每个ID分组执行以下规则:
- 若组内存在相同length对应不同comment的情况,该组所有行统一使用
typed公式:5 × length - 若组内不存在上述冲突,则按每行的comment分别匹配公式:
- typed:
5 × length - handwritten:
7 × length
- typed:
实现代码
# 1. 标记每个ID-length组合是否存在comment冲突 df['has_conflict'] = df.groupby(['ID', 'length'])['comment'].transform('nunique') > 1 # 2. 标记整个ID组是否存在冲突(只要有一个length组冲突,全组统一规则) df['group_has_conflict'] = df.groupby('ID')['has_conflict'].transform('any') # 3. 生成对应计算规则的文本 def get_calc_rule(row): if row['group_has_conflict']: return f"5*{row['length']}" else: return f"5*{row['length']}" if row['comment'] == 'typed' else f"7*{row['length']}" df['Calculated Length'] = df.apply(get_calc_rule, axis=1) # 清理中间辅助列(可选) df = df.drop(['has_conflict', 'group_has_conflict'], axis=1)
最终结果
ID length comment Calculated Length 0 1 0.7 typed 5*0.7 1 1 0.7 handwritten 5*0.7 2 2 0.8 typed 5*0.8 3 2 0.6 typed 5*0.6 4 2 0.6 handwritten 7*0.6 5 3 0.9 handwritten 7*0.9 6 3 0.9 handwritten 7*0.9
代码逻辑说明
- 标记冲突单元:通过
groupby(['ID', 'length'])统计每个长度组下的comment类型数量,若数量大于1则标记为冲突 - 标记组冲突状态:对每个ID组,只要存在任意一个冲突的长度组,就标记该ID组需要统一使用typed公式
- 生成计算规则:根据组冲突状态选择对应规则,无冲突时则按每行的comment匹配对应公式
内容的提问来源于stack exchange,提问作者ukanafun
相关产品推荐
相关产品推荐

