如何在Pandas DataFrame中统计特定字符串与正负整数的数量
混合类型Pandas列的统计方案
方案一:基于元素类型筛选统计
针对你的场景(列中仅包含整数和'Hello'字符串),可以通过判断元素类型来分别统计:
- 统计'Hello'的数量:
helloCount = len(df[df['A1'] == 'Hello'])
- 统计非负整数数量:先筛选出整数类型的行,再判断数值≥0
posNums = len(df[df['A1'].apply(lambda x: isinstance(x, int)) & (df['A1'] >= 0)])
- 统计负整数数量:同样先筛选整数类型,再判断数值<0
negNums = len(df[df['A1'].apply(lambda x: isinstance(x, int)) & (df['A1'] < 0)])
方案二:数值转换后统计(更通用)
如果列中可能存在其他非数字字符串,推荐用pd.to_numeric将非数值转为NaN,再基于NaN区分统计:
import pandas as pd # 将A1列转为数值,非数值自动设为NaN numeric_col = pd.to_numeric(df['A1'], errors='coerce') # 统计Hello数量 helloCount = (df['A1'] == 'Hello').sum() # 统计非负数字数量(排除NaN) posNums = (numeric_col >= 0).sum() # 统计负数字数量(排除NaN) negNums = (numeric_col < 0).sum()
示例验证
用你的测试数据运行方案二的代码:
data = {'A1': [1, 'Hello', -8, 'Hello']} df = pd.DataFrame(data) numeric_col = pd.to_numeric(df['A1'], errors='coerce') helloCount = (df['A1'] == 'Hello').sum() posNums = (numeric_col >= 0).sum() negNums = (numeric_col < 0).sum() print(f"posNums={posNums},negNums={negNums},helloCount={helloCount}")
输出:posNums=1,negNums=1,helloCount=2,完全符合预期。
方案对比
- 方案一:逻辑直接,适合元素类型明确只有整数和'Hello'的场景;
- 方案二:兼容性更强,能处理包含其他非数字字符串的混合列。
内容的提问来源于stack exchange,提问作者Petra Enis
相关产品推荐
相关产品推荐

