如何为含列表与NaN的Pandas DataFrame生成指定规则的dummy列
问题:基于Pandas列表列生成条件哑变量
原始DataFrame
import pandas as pd import numpy as np temp = pd.DataFrame({'x': [['ab', 'bc'], ['hg'], np.nan]}) temp
输出结果:
| x | |
|---|---|
| 0 | [ab, bc] |
| 1 | [hg] |
| 2 | NaN |
需求
创建名为dummy的新列,规则如下:
- 若该行的列表元素中包含字母'a'(不区分大小写),值为1
- 若列表中不包含'a',值为0
- 若该行是NaN,保持NaN
预期结果:
| x | dummy | |
|---|---|---|
| 0 | [ab, bc] | 1 |
| 1 | [hg] | 0 |
| 2 | NaN | NaN |
尝试过的方法及问题
方法1:直接用
str.contains比较整个列表,结果全为0temp['dummy'] = np.where(temp.x.str.contains('a', case=False, na=False), 1, 0)问题:
str.contains会把整个列表当作一个字符串比较,无法识别列表内元素的'a'。方法2:转为字符串后调用
str.contains,但NaN被转为字符串'nan',导致na=False失效temp['dummy'] = np.where(temp.x.astype(str).str.contains('a', case=False, na=False), 1, 0)问题:第2行的NaN转为"nan"后,因不包含'a'被赋值为0,不符合需求。
方法3:用
all组合条件,触发Series布尔值歧义错误temp['dummy'] = np.where(all([temp.x.astype(str).str.contains('a', case=False, na=False) , temp.x.astype(str) != 'nan']), 1, 0)错误:
ValueError: The truth value of a Series is ambiguous.方法4:列表推导式处理NaN时触发迭代错误
temp['dummy'] = [1 if all(['a' in y , y != np.nan]) else 0 for y in temp.x ]错误:
TypeError: argument of type 'float' is not iterable方法5:可行但代码繁琐,先占位NaN再赋值
temp['dummy'] = np.nan # placeholder temp['dummy'][temp.x.notnull()] = np.where(temp[temp.x.notnull()].x.astype(str).str.contains('a', case=False, na=False), 1, 0)
简洁解决方案
方案1:apply结合类型判断与元素检查
直接对每行元素做判断,先确认是列表再检查是否有元素包含'a':
temp['dummy'] = temp['x'].apply( lambda y: 1 if isinstance(y, list) and any('a' in s.lower() for s in y) else 0 if isinstance(y, list) else np.nan )
方案2:str.join拼接列表后用str.contains
利用Pandas的字符串方法自动处理NaN,拼接列表后检查包含关系:
temp['dummy'] = temp['x'].str.join(',')\ .str.contains('a', case=False)\ .map({True:1, False:0})\ .astype('Int64')
解释:str.join(',')对NaN会返回NaN,str.contains保持NaN,映射为1/0后用Int64类型保留空值。
方案3:np.where结合notnull与apply
用np.where区分NaN行和非NaN行,非NaN行再检查列表元素:
temp['dummy'] = np.where( temp['x'].notnull(), temp['x'].apply(lambda lst: 1 if any('a' in s.lower() for s in lst) else 0), np.nan )
内容的提问来源于stack exchange,提问作者Saeed
相关产品推荐
相关产品推荐

