如何使用Pandas统计CSV文件指定列中纯整数的数量
Hey there! Let's break this down clearly since you're new to Pandas—regex matching and counting can feel a bit confusing at first, but we'll get it sorted.
First, great call using ^ and $ in your regex! Those anchors ensure you're matching the entire string instead of just parts of it, which is exactly what we need for pure integers. Here are two straightforward ways to count those values:
方法1:用str.fullmatch()(推荐)
This method checks if the entire string matches the regex pattern, returning a boolean Series (True for pure integers, False otherwise). Since True is treated as 1 and False as 0 in Pandas, we can just sum the results to get the count.
import pandas as pd # 读取你的CSV数据集 df = pd.read_csv('your_dataset.csv') # 替换成你要统计的列名 target_column = 'your_column_name' # 统计纯整数数量:用fullmatch匹配全数字字符串,na=False处理缺失值 pure_int_count = df[target_column].str.fullmatch(r'\d+', na=False).sum() print(f"指定列中的纯整数数量为: {pure_int_count}")
r'\d+':匹配至少一个数字(比\d*更合适,因为\d*会匹配空字符串)na=False:把缺失值(NaN)标记为False,避免干扰计数结果
方法2:延续你原来的str.extract()思路
If you want to stick with the extract approach you started with, you can count how many non-null values you get from the extraction (since non-matches return NaN):
# 用extract提取纯整数,然后统计非空值的数量 extracted = df[target_column].str.extract(r'(^\d+$)') pure_int_count = extracted.notna().sum()[0] # [0]是因为extract返回的是DataFrame类型 print(f"指定列中的纯整数数量为: {pure_int_count}")
This works because extract(r'(^\d+$)') returns a DataFrame where cells with pure integers hold the number string, and all non-matching cells have NaN. notna() converts those to booleans, then summing gives the total count of pure integers.
Either method will get you the number you need—fullmatch is just more concise and efficient for this specific task.
内容的提问来源于stack exchange,提问作者Girish venkata

