如何在PySpark中基于其他列计算创建新列(百分比示例)
在PySpark DataFrame中添加百分比计算列
嘿,这事儿很简单,我来一步步教你怎么实现!首先先确认下你的原始DataFrame结构:
| some_id | one_col | other_col |
|---|---|---|
| xx1 | 11 | 177 |
| xx2 | 1613 | 2000 |
| xx4 | 0 | 12473 |
你需要的是新增一列,计算one_col占other_col的百分比,核心公式就是(one_col / other_col) * 100。下面是具体的实现代码,还特意考虑了除数为0的边界情况,避免报错:
首先导入PySpark的函数库:
from pyspark.sql import functions as F
然后用withColumn方法添加新列,同时处理other_col为0的异常场景:
# 假设你的DataFrame名为df df = df.withColumn( "percentage_col", # 当other_col不等于0时计算百分比,否则设为0(你也可以根据需求改成null或者其他默认值) F.when(F.col("other_col") != 0, (F.col("one_col") / F.col("other_col")) * 100).otherwise(0) )
执行完之后,你的DataFrame就会变成这样:
| some_id | one_col | other_col | percentage_col |
|---|---|---|---|
| xx1 | 11 | 177 | 6.214689265536723 |
| xx2 | 1613 | 2000 | 80.65 |
| xx4 | 0 | 12473 | 0.0 |
如果想要让百分比结果更整洁,比如保留两位小数,可以用F.round函数优化:
df = df.withColumn( "percentage_col", F.when(F.col("other_col") != 0, F.round((F.col("one_col") / F.col("other_col")) * 100, 2)).otherwise(0) )
优化后的结果会更直观:
| some_id | one_col | other_col | percentage_col |
|---|---|---|---|
| xx1 | 11 | 177 | 6.21 |
| xx2 | 1613 | 2000 | 80.65 |
| xx4 | 0 | 12473 | 0.0 |
内容的提问来源于stack exchange,提问作者Ivan Bilan
相关产品推荐
相关产品推荐

