You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则提取OCR处理的银行PDF脱敏账号问题

Fixing Account Number Extraction from OCR'd Bank Statements

Let's break down why your current regex isn't working and get you the correct account number extraction.

What's Wrong with the Original Regex?

Your regex r'^[A-Z].*([0-9]{4}$)' is trying to match lines starting with an uppercase letter, then capture the last 4 digits of the line. But this doesn't target the specific structure of your account number: 8 X's followed by 4 digits (e.g., XXXXXXXX1002). It also only captures the final 4 digits, not the full 12-character account string you need.

Correct Regex Approach

Since you know the account number is a fixed 12-character string with exactly 8 leading X's and 4 trailing digits, we can write a regex that directly matches this pattern:

import re

splits = [
    'ACCOUNT TYPE A/C. BALANCE (I) | FIXED DEPOSITS (LINKED) BAL. (II) TOTAL BALANCE (I+II) NOMINATION',
    'Savings A/c XXXXXXXX1002 4,58,636.08 0.00 4,58,636.08 Not Registered',
    'TOTAL 4,58,636.08 0.00 4,58,636.08'
]

account_number = None
for line in splits:
    # Match exactly 8 X's followed by 4 digits
    match = re.search(r'XXXXXXXX\d{4}', line)
    if match:
        account_number = match.group()
        break  # Stop searching once we find the account

print(account_number)  # Output: XXXXXXXX1002

Key Improvements:

  • The regex r'XXXXXXXX\d{4}' directly targets the unique structure of your desensitized account number, ensuring we only capture the exact string you need.
  • Using re.search() looks for the pattern anywhere in the line (since the account number isn't at the start or end of the line in your example).
  • We use match.group() to get the full matched string, not just a subset.

Flexible Alternative (If X Count Varies)

If for some reason the number of leading X's isn't always 8 (but the total length is 12 with last 4 digits), you can use this more flexible regex:

match = re.search(r'X{8}\d{4}', line)
# Or, if X count could vary but total length is 12:
match = re.search(r'X+\d{4}', line)

This should reliably extract your account number from the OCR'd lines.

内容的提问来源于stack exchange,提问作者Subhojyoti Lahiri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:58:12