You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于不同列的条件创建Pandas DataFrame并解决歧义报错

问题解决:创建符合条件的control_df DataFrame

需求说明

需要创建名为control_df的DataFrame,满足以下任一条件:

  • (i) mrna_assignment列包含子串control //
  • (ii) probeset_type列包含子串control

原错误代码及报错

错误代码

import pandas as pd
import numpy as np

df.columns = df.iloc[0]
df = df[1:].set_index("Gene Symbol") # set Gene Symbol as row index
df.sort_index() # Sort by row index
df

## Samples are "control" if:
## (i) "mrna_assignment" column contains substring "control //"; OR
## (ii) "probeset_type" column contains substring "control"
control_df = df[df["mrna_assignment"].str.contains("control //")] or df[df["category"].str.contains("control->")]
control_df

报错信息

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-167-516e7c765379> in <module>()
      3 ## (i) "mrna_assignment" column contains substring "control //"; OR
      4 ## (ii) "probeset_type" column contains substring "control"
----> 5 control_df = df[df["mrna_assignment"].str.contains("control //")] or df[df["category"].str.contains("control->")]
      6 control_df

/usr/local/lib/python3.7/dist-packages/pandas/core/generic.py in __nonzero__(self)
   1536     def __nonzero__(self):
   1537         raise ValueError(
-> 1538             f"The truth value of a {type(self).__name__} is ambiguous. "
   1539             "Use a.empty, a.bool(), a.item(), a.any() or a.all()."
   1540         )

ValueError: The truth value of a DataFrame is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().

问题分析

  1. 逻辑判断错误:Python原生的or不能直接用于Pandas DataFrame的布尔筛选,需使用Pandas支持的逻辑或运算符|组合布尔索引。
  2. 列名不匹配需求:需求要求检查probeset_type列,但代码错误使用了category列。
  3. 空值处理缺失:str.contains遇到NaN值会返回NaN,导致布尔索引报错,需将NaN转为False避免问题。
  4. 排序未生效:df.sort_index()返回新的排序后DataFrame,未赋值的话原df不会改变,需添加赋值操作。

修正后的代码

import pandas as pd
import numpy as np

# 初始化数据:设置列名、切片有效数据、设置行索引、执行排序
df.columns = df.iloc[0]
df = df[1:].set_index("Gene Symbol")
df = df.sort_index()  # 赋值使排序生效

# 构建两个筛选条件,同时处理空值
cond1 = df["mrna_assignment"].str.contains("control //", na=False)
cond2 = df["probeset_type"].str.contains("control", na=False)

# 组合条件并筛选得到目标DataFrame
control_df = df[cond1 | cond2]
control_df

代码说明

  • na=False:在str.contains中直接将空值对应的结果设为False,避免布尔索引报错。
  • cond1 | cond2:用Pandas逻辑或运算符组合两个条件,筛选出满足任一条件的行。
  • 修正列名:将错误的category替换为需求中的probeset_type。
  • 排序添加赋值:确保排序操作实际作用于df。

内容的提问来源于stack exchange,提问作者melolilili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 19:21:35