关于构建ICD-10编码Python正则表达式的技术咨询
ICD-10 Coding Regular Expression Solution
Got it, let's fix that regex for you. The key missing piece in your ending section was ensuring the dot is required if any trailing characters exist, and wrapping that entire optional segment correctly.
Complete Regular Expression
Here's the final regex that meets all your requirements:
import re icd10_pattern = re.compile(r'^[A-Z][0-9]{2}(?:\.[A-Z0-9]{1,4})?$')
If you want to accept lowercase letters (e.g., a00.0), add the re.IGNORECASE flag:
icd10_pattern = re.compile(r'^[A-Z][0-9]{2}(?:\.[A-Z0-9]{1,4})?$', re.IGNORECASE)
Breakdown of the Pattern
Let's walk through each part to make sure it makes sense:
^: Anchors the match to the start of the string (prevents partial matches at the beginning)[A-Z]: Matches the required initial uppercase letter[0-9]{2}: Matches exactly two digits (aligns with your original[A-Z][0-9][0-9]segment)(?:\.[A-Z0-9]{1,4})?: This is the optional ending segment:?:: Marks this as a non-capturing group (we don't need to extract this part separately)\.: The required dot—this only exists if trailing characters are present[A-Z0-9]{1,4}: Matches 1 to 4 alphanumeric characters (letters or digits)?: Makes the entire ending group optional (so codes without the trailing segment are valid)
$: Anchors the match to the end of the string (prevents extra characters after the code)
Test Cases
Let's verify with your examples and edge cases:
Valid Codes
These should all return True:
valid_codes = [ "A00", # No trailing segment "A00.0", # .0 "A00.A9", # .A9 "A00.A9A9", # .A9A9 "A00.ZZZZ", # .ZZZZ "A00.9999", # .9999 "B12.3X" # Mixed digits and letters ] for code in valid_codes: print(f"{code}: {icd10_pattern.match(code) is not None}")
Invalid Codes
These should all return False:
invalid_codes = [ "A0", # Too short initial segment (only 2 characters) "A000", # Extra digit without a dot "A00.", # Dot with no characters after "A00.A9A9A", # More than 4 characters after dot "a00.0", # Lowercase initial letter (without IGNORECASE flag) "A00-0" # Hyphen instead of dot ] for code in invalid_codes: print(f"{code}: {icd10_pattern.match(code) is not None}")
Notes
- You mentioned you know some ICD-10 codes have specific range restrictions, but since you don't need those, this pattern keeps it general as requested.
- The non-capturing group (
?:) is optional if you want to capture the trailing segment—just remove the?:if you need to extract that part later.
内容的提问来源于stack exchange,提问作者DanielBell99
相关产品推荐
相关产品推荐

