按client_ip分组统计DataFrame中4xx响应码次数的代码错误排查
分组统计4xx响应码的报错原因及解决方法
问题场景
现有如下结构的DataFrame:
client_ip http_response_code index 2022-07-23 05:10:10+00:00 172.19.0.1 300 2022-07-23 06:13:26+00:00 192.168.0.1 400 ... ... ...
需求是按client_ip分组,统计http_response_code列中所有4xx响应码(以4开头的整数)的出现次数。
尝试了以下代码后报错:
df.groupby('client_ip')['http_response_code'].apply(lambda x: (str(x).startswith(str(4))).sum())
报错信息:
AttributeError: 'bool' object has no attribute 'sum'
但统计特定响应码400的代码能正常运行:
df.groupby('client_ip')['http_response_code'].apply(lambda x: (x==400).sum())
错误原因
核心问题出在str(x)的处理逻辑上:
- 这里的
x是分组后的整个Series对象,不是单个响应码元素。str(x)会把整个Series转换成一段描述性字符串(比如类似"0 300\n1 400\nName: http_response_code, dtype: int64"),而非对每个元素单独转字符串。 - 后续调用
startswith('4'),是判断这段整体字符串是否以"4"开头,结果返回一个单个布尔值(显然不成立),而布尔值没有sum()方法,因此触发报错。
反观x==400的写法,是对Series中的每个元素逐一做比较,得到一个布尔类型的Series,这类Series支持sum()方法(统计其中True的数量),所以能正常运行。
正确解法
推荐两种高效的实现方式:
方式1:字符串逐元素处理
对分组后的Series每个元素单独转字符串,再判断是否以"4"开头:
df.groupby('client_ip')['http_response_code'].apply(lambda x: x.astype(str).str.startswith('4').sum())
方式2:数值范围判断(更高效)
利用4xx响应码的数值范围(400≤code<500)直接判断,避免字符串转换的性能开销:
df.groupby('client_ip')['http_response_code'].apply(lambda x: ((x >= 400) & (x < 500)).sum())
也可以用更简洁的agg写法:
df.groupby('client_ip')['http_response_code'].agg(lambda x: ((x >=400) & (x <500)).sum())
内容的提问来源于stack exchange,提问作者Kosmylo
相关产品推荐
相关产品推荐

