You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为含列表与NaN的Pandas DataFrame生成指定规则的dummy列

问题:基于Pandas列表列生成条件哑变量

原始DataFrame

import pandas as pd
import numpy as np

temp = pd.DataFrame({'x': [['ab', 'bc'], ['hg'], np.nan]})
temp

输出结果:

x
0[ab, bc]
1[hg]
2NaN

需求

创建名为dummy的新列,规则如下:

  • 若该行的列表元素中包含字母'a'(不区分大小写),值为1
  • 若列表中不包含'a',值为0
  • 若该行是NaN,保持NaN

预期结果:

xdummy
0[ab, bc]1
1[hg]0
2NaNNaN

尝试过的方法及问题

  • 方法1:直接用str.contains比较整个列表,结果全为0

    temp['dummy'] = np.where(temp.x.str.contains('a', case=False, na=False), 1, 0)
    

    问题:str.contains会把整个列表当作一个字符串比较,无法识别列表内元素的'a'。

  • 方法2:转为字符串后调用str.contains,但NaN被转为字符串'nan',导致na=False失效

    temp['dummy'] = np.where(temp.x.astype(str).str.contains('a', case=False, na=False), 1, 0)
    

    问题:第2行的NaN转为"nan"后,因不包含'a'被赋值为0,不符合需求。

  • 方法3:用all组合条件,触发Series布尔值歧义错误

    temp['dummy'] = np.where(all([temp.x.astype(str).str.contains('a', case=False, na=False) , temp.x.astype(str) != 'nan']), 1, 0)
    

    错误:ValueError: The truth value of a Series is ambiguous.

  • 方法4:列表推导式处理NaN时触发迭代错误

    temp['dummy'] = [1 if all(['a' in y , y != np.nan]) else 0 for y in temp.x ]
    

    错误:TypeError: argument of type 'float' is not iterable

  • 方法5:可行但代码繁琐,先占位NaN再赋值

    temp['dummy'] = np.nan # placeholder
    temp['dummy'][temp.x.notnull()] = np.where(temp[temp.x.notnull()].x.astype(str).str.contains('a', case=False, na=False), 1, 0)
    

简洁解决方案

方案1:apply结合类型判断与元素检查

直接对每行元素做判断,先确认是列表再检查是否有元素包含'a':

temp['dummy'] = temp['x'].apply(
    lambda y: 1 if isinstance(y, list) and any('a' in s.lower() for s in y) 
              else 0 if isinstance(y, list) 
              else np.nan
)

方案2:str.join拼接列表后用str.contains

利用Pandas的字符串方法自动处理NaN,拼接列表后检查包含关系:

temp['dummy'] = temp['x'].str.join(',')\
                         .str.contains('a', case=False)\
                         .map({True:1, False:0})\
                         .astype('Int64')

解释:str.join(',')对NaN会返回NaN,str.contains保持NaN,映射为1/0后用Int64类型保留空值。

方案3:np.where结合notnull与apply

用np.where区分NaN行和非NaN行,非NaN行再检查列表元素:

temp['dummy'] = np.where(
    temp['x'].notnull(),
    temp['x'].apply(lambda lst: 1 if any('a' in s.lower() for s in lst) else 0),
    np.nan
)

内容的提问来源于stack exchange,提问作者Saeed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 21:55:04