You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于指定字符数字格式过滤DataFrame行并解决re.findall报错

问题原因分析
  • re.findall() 仅接收单个字符串/字节类型的入参,你传入的是pandas Series列对象,哪怕列的dtype是string,本质也是序列而非单个字符串,因此触发类型错误
  • 你写的正则表达式存在语法错误:数字匹配的结尾符号写错,[0-9]{4] 应该改为 [0-9]{4},需使用右大括号闭合长度限制
  • 示例数据的字典定义不完整,缺少右大括号,也没有转成pandas DataFrame对象
实现方案

遍历实现(小数据量适用)

import pandas as pd
import re
from collections import defaultdict

# 补全示例DataFrame
df_dict = {'a':[1,2,4,5,6], 'b':[7, 8, 9,10, 11], 'target':[ 'ABC1234','ABC123', '123ABC', '7KZA23', 'XYZ9876']}
df = pd.DataFrame(df_dict)

# 定义格式和对应正则规则
format_patterns = {
    'ABC1234': r'^[A-Z]{3}[0-9]{4}$',  # 3大写字母+4数字
    'ABC123': r'^[A-Z]{3}[0-9]{3}$',   # 3大写字母+3数字
    '123ABC': r'^[0-9]{3}[A-Z]{3}$'    # 3数字+3大写字母
}

count_res = defaultdict(int)

# 逐行匹配计数
for val in df['target'].astype('string'):
    matched = False
    for fmt_name, pat in format_patterns.items():
        if re.fullmatch(pat, val):
            count_res[fmt_name] += 1
            matched = True
            break
    if not matched:
        count_res['any_other_format'] += 1

# 输出最终字典
print(dict(count_res))

运行输出结果:

{'ABC1234': 2, 'ABC123': 1, '123ABC': 1, 'any_other_format': 1}

向量化实现(大数据量适用)

用pandas内置的字符串匹配方法,无需遍历,效率更高:

count_res = {}
# 统计指定格式匹配数
count_res['ABC1234'] = df['target'].str.fullmatch(r'^[A-Z]{3}[0-9]{4}$').sum()
count_res['ABC123'] = df['target'].str.fullmatch(r'^[A-Z]{3}[0-9]{3}$').sum()
count_res['123ABC'] = df['target'].str.fullmatch(r'^[0-9]{3}[A-Z]{3}$').sum()
# 其他格式=总数减去已匹配的总和
count_res['any_other_format'] = len(df) - sum(count_res.values())

print(count_res)

内容的提问来源于stack exchange,提问作者dimension_dweller

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 04:21:01