Pandas如何统计Series每个字符串元素的最大连续空格数
统计Pandas Series内每个字符串的最长连续空格长度
你提到的ser.str.count(' ')仅能统计空格总数量,无法识别连续空格的长度,要实现需求可以通过正则匹配连续空格片段后取长度最大值,代码如下:
实现代码
import pandas as pd ser = pd.Series([ 'the quick brown fox', 'the quick pink fox', 'a quick brown fox', 'the jumpy brown fox ', 'the quick brown animal', ]) # 匹配所有连续空格段,计算每段长度后取最大值 res = ser.str.findall(r' +').apply(lambda item: max(len(space_seq) for space_seq in item)) print(res)
运行输出和预期完全一致:
0 1 1 2 2 3 3 4 4 2 dtype: int64
实现逻辑
ser.str.findall(r' +')会逐元素匹配所有1个及以上的连续空格片段,返回值是列表格式,每个列表存储当前字符串里所有连续空格的子串,例如第三个元素a quick brown fox匹配得到的列表为[' ', ' ', ' ']- 后续通过
apply遍历每个列表,计算每个空格子串的长度,取最大值即为当前字符串的最长连续空格长度
如果需要兼容字符串完全不含空格的边界场景,可以给max函数设置默认返回值0,避免空列表触发报错:
res = ser.str.findall(r' +').apply(lambda item: max((len(space_seq) for space_seq in item), default=0))
内容的提问来源于stack exchange,提问作者ignoring_gravity
相关产品推荐
相关产品推荐

