如何在列表推导式中添加条件处理Pandas中的NaN值?
问题描述
现有如下Pandas数据集:
import pandas as pd import numpy as np df = pd.DataFrame({ 'name': ['John','William', 'Nancy', 'Susan', 'Robert', 'Lucy', 'Blake', 'Sally', 'Bruce', 'Mike', 'Jeff'], 'injury': ['right hand broken', 'lacerated left foot', 'foot broken', 'right foot fractured', '', 'sprained finger', 'chest pain', 'swelling in arm', 'laceration to arms, hands, and foot', np.NaN, 'swollen cheek'] })
数据集内容:
| name | injury | |
|---|---|---|
| 0 | John | right hand broken |
| 1 | William | lacerated left foot |
| 2 | Nancy | foot broken |
| 3 | Susan | right foot fractured |
| 4 | Robert | |
| 5 | Lucy | sprained finger |
| 6 | Blake | chest pain |
| 7 | Sally | swelling in arm |
| 8 | Bruce | laceration to arms, hands, and foot |
| 9 | Mike | NaN |
| 10 | Jeff | swollen cheek |
尝试将损伤信息简化为指定身体部位,使用以下代码:
selected_words = ["hand", "foot", "finger", "chest", "arms", "arm", "hands"] df["injury"] = ( df["injury"] .str.replace(",", "") .str.split(" ", expand=False) .apply(lambda x: ", ".join(set([i for i in x if i in selected_words]))) )
处理索引9的NaN值时抛出错误:
TypeError: 'float' object is not iterable
需求:
- 正确识别所有NaN值
- 空行或未包含指定身体部位的行(如索引10)输出NaN
期望输出:
| name | injury | |
|---|---|---|
| 0 | John | hand |
| 1 | William | foot |
| 2 | Nancy | foot |
| 3 | Susan | foot |
| 4 | Robert | NaN |
| 5 | Lucy | finger |
| 6 | Blake | chest |
| 7 | Sally | arm |
| 8 | Bruce | hand, foot, arm |
| 9 | Mike | NaN |
| 10 | Jeff | NaN |
尝试过的错误代码:
.apply(lambda x: ", ".join(set([i for i in x if i in selected_words and i is not np.nan else np.nan])))
解决方案
问题核心是NaN为float类型无法被迭代,且空字符串与NaN的空值逻辑不统一,修正后的代码如下:
import pandas as pd import numpy as np selected_words = ["hand", "foot", "finger", "chest", "arms", "arm", "hands"] # 先将空字符串转为NaN,统一空值处理逻辑 df["injury"] = df["injury"].replace("", np.nan) df["injury"] = ( df["injury"] # 仅对非NaN值执行字符串拆分操作 .apply(lambda x: x.replace(",", "").split(" ") if pd.notna(x) else x) .apply(lambda x: # 过滤指定部位、去重后拼接 ", ".join(set([word for word in x if word in selected_words])) if pd.notna(x) and len([word for word in x if word in selected_words]) > 0 else np.nan) )
代码说明
- 统一空值:将原数据中的空字符串转为NaN,避免后续区分两种空值情况;
- 跳过NaN处理:通过
pd.notna(x)判断,仅对非NaN值执行字符串替换、拆分操作,NaN直接返回; - 过滤与空值判断:拆分后的列表先过滤出目标身体部位,若过滤结果为空(无匹配部位),则返回NaN,否则拼接去重后的结果;
- 去重逻辑:使用
set自动去重,保证每个身体部位只出现一次。
执行后即可得到符合要求的输出结果。
内容的提问来源于stack exchange,提问作者equanimity
相关产品推荐
相关产品推荐

