You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas query方法处理DataFrame中的字符串整数及异常值

问题

原始代码如下:

import numpy as np
import pandas as pd
numbers = ["1", "2", "3", "4", "5", "6", "Missing Value", "7"]
df = pd.DataFrame(numbers, columns=["Numbers"])
new_df = df.query("Numbers > 4")
print(new_df)

运行后报错:TypeError: '>' not supported between instances of 'str' and 'int',原因是Numbers列的元素均为字符串类型,直接与整数4比较会触发类型不匹配错误。

要求:必须使用query方法,不能将整列转换为整数类型,同时要处理"Missing Value"这类无法转为整数的字符串。尝试过new_df = df.query("int(Numbers) > 4")但无效,求解决办法。

解决方法

在query中使用pd.to_numeric并指定errors='coerce'参数,将可转换的字符串转为数值,不可转换的转为NaN,NaN在比较时会自动被排除:

import numpy as np
import pandas as pd
numbers = ["1", "2", "3", "4", "5", "6", "Missing Value", "7"]
df = pd.DataFrame(numbers, columns=["Numbers"])
new_df = df.query("pd.to_numeric(Numbers, errors='coerce') > 4")
print(new_df)

说明

  • pd.to_numeric(Numbers, errors='coerce')会对Numbers列的每个元素做转换:能转成数字的字符串转为对应整数,无法转换的内容(如"Missing Value")转为NaN。
  • 在比较运算>4中,NaN参与的比较结果为False,因此这部分不符合条件的行会被自动过滤,最终得到的结果只包含数值大于4的有效字符串行。

内容的提问来源于stack exchange,提问作者bravesheeptribe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 08:54:17