Pandas如何筛选col2位数≤3的行并按规则生成计算列col3
需求说明
当前需要处理规模约2000行的pandas DataFrame,核心处理规则如下:
- 识别
col2列中数字位数为3位及以下的行 - 新增
col3列,赋值规则:- 若
col2位数≤3:col3取值为对应行col1与col2的乘积 - 若
col2位数>3:col3取值与对应行col2一致
- 若
示例数据集
构造的测试数据集代码如下:
import pandas as pd d = {'col1': [10000, 2000,300,4000,50000], 'col2': [10, 20000, 300, 4000, 100]} df = pd.DataFrame(data=d) print(df)
数据集输出:
col1 col2 0 10000 10 1 2000 20000 2 300 300 3 4000 4000 4 50000 100
注:示例中dtypes打印显示Area/Price为粘贴笔误,实际字段为col1、col2,类型均为int64,不影响逻辑实现
预期输出
处理完成后的目标结果如下:
col1 col2 col3 0 10000 10 100000 1 2000 20000 20000 2 300 300 90000 3 4000 4000 4000 4 50000 100 500000
输出字段类型要求:col1、col2、col3均为int64类型。
实现方案
使用numpy.where做条件判断即可,运行效率高,适配2000行规模数据集无性能压力,代码如下:
import numpy as np # 正整数场景下,小于1000即为3位及以下数字,判断效率最高 df['col3'] = np.where(df['col2'] < 1000, df['col1'] * df['col2'], df['col2'])
如果你的col2列可能存在0、负整数,可使用通用位数判断逻辑,兼容所有整数场景:
# 转字符串后去除负号再判断长度,兼容负数、0场景 cond = df['col2'].astype(str).str.lstrip('-').str.len() <= 3 df['col3'] = np.where(cond, df['col1'] * df['col2'], df['col2'])
运行后打印df和df.dtypes即可得到和预期完全一致的结果。
内容的提问来源于stack exchange,提问作者Bojan Pavlović
相关产品推荐
相关产品推荐

