如何基于指定字符数字格式过滤DataFrame行并解决re.findall报错
问题原因分析
re.findall()仅接收单个字符串/字节类型的入参,你传入的是pandas Series列对象,哪怕列的dtype是string,本质也是序列而非单个字符串,因此触发类型错误- 你写的正则表达式存在语法错误:数字匹配的结尾符号写错,
[0-9]{4]应该改为[0-9]{4},需使用右大括号闭合长度限制 - 示例数据的字典定义不完整,缺少右大括号,也没有转成pandas DataFrame对象
实现方案
遍历实现(小数据量适用)
import pandas as pd import re from collections import defaultdict # 补全示例DataFrame df_dict = {'a':[1,2,4,5,6], 'b':[7, 8, 9,10, 11], 'target':[ 'ABC1234','ABC123', '123ABC', '7KZA23', 'XYZ9876']} df = pd.DataFrame(df_dict) # 定义格式和对应正则规则 format_patterns = { 'ABC1234': r'^[A-Z]{3}[0-9]{4}$', # 3大写字母+4数字 'ABC123': r'^[A-Z]{3}[0-9]{3}$', # 3大写字母+3数字 '123ABC': r'^[0-9]{3}[A-Z]{3}$' # 3数字+3大写字母 } count_res = defaultdict(int) # 逐行匹配计数 for val in df['target'].astype('string'): matched = False for fmt_name, pat in format_patterns.items(): if re.fullmatch(pat, val): count_res[fmt_name] += 1 matched = True break if not matched: count_res['any_other_format'] += 1 # 输出最终字典 print(dict(count_res))
运行输出结果:
{'ABC1234': 2, 'ABC123': 1, '123ABC': 1, 'any_other_format': 1}
向量化实现(大数据量适用)
用pandas内置的字符串匹配方法,无需遍历,效率更高:
count_res = {} # 统计指定格式匹配数 count_res['ABC1234'] = df['target'].str.fullmatch(r'^[A-Z]{3}[0-9]{4}$').sum() count_res['ABC123'] = df['target'].str.fullmatch(r'^[A-Z]{3}[0-9]{3}$').sum() count_res['123ABC'] = df['target'].str.fullmatch(r'^[0-9]{3}[A-Z]{3}$').sum() # 其他格式=总数减去已匹配的总和 count_res['any_other_format'] = len(df) - sum(count_res.values()) print(count_res)
内容的提问来源于stack exchange,提问作者dimension_dweller
相关产品推荐
相关产品推荐

