Python正则提取OCR处理的银行PDF脱敏账号问题
Let's break down why your current regex isn't working and get you the correct account number extraction.
What's Wrong with the Original Regex?
Your regex r'^[A-Z].*([0-9]{4}$)' is trying to match lines starting with an uppercase letter, then capture the last 4 digits of the line. But this doesn't target the specific structure of your account number: 8 X's followed by 4 digits (e.g., XXXXXXXX1002). It also only captures the final 4 digits, not the full 12-character account string you need.
Correct Regex Approach
Since you know the account number is a fixed 12-character string with exactly 8 leading X's and 4 trailing digits, we can write a regex that directly matches this pattern:
import re splits = [ 'ACCOUNT TYPE A/C. BALANCE (I) | FIXED DEPOSITS (LINKED) BAL. (II) TOTAL BALANCE (I+II) NOMINATION', 'Savings A/c XXXXXXXX1002 4,58,636.08 0.00 4,58,636.08 Not Registered', 'TOTAL 4,58,636.08 0.00 4,58,636.08' ] account_number = None for line in splits: # Match exactly 8 X's followed by 4 digits match = re.search(r'XXXXXXXX\d{4}', line) if match: account_number = match.group() break # Stop searching once we find the account print(account_number) # Output: XXXXXXXX1002
Key Improvements:
- The regex
r'XXXXXXXX\d{4}'directly targets the unique structure of your desensitized account number, ensuring we only capture the exact string you need. - Using
re.search()looks for the pattern anywhere in the line (since the account number isn't at the start or end of the line in your example). - We use
match.group()to get the full matched string, not just a subset.
Flexible Alternative (If X Count Varies)
If for some reason the number of leading X's isn't always 8 (but the total length is 12 with last 4 digits), you can use this more flexible regex:
match = re.search(r'X{8}\d{4}', line) # Or, if X count could vary but total length is 12: match = re.search(r'X+\d{4}', line)
This should reliably extract your account number from the OCR'd lines.
内容的提问来源于stack exchange,提问作者Subhojyoti Lahiri

