You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于构建ICD-10编码Python正则表达式的技术咨询

ICD-10 Coding Regular Expression Solution

Got it, let's fix that regex for you. The key missing piece in your ending section was ensuring the dot is required if any trailing characters exist, and wrapping that entire optional segment correctly.

Complete Regular Expression

Here's the final regex that meets all your requirements:

import re

icd10_pattern = re.compile(r'^[A-Z][0-9]{2}(?:\.[A-Z0-9]{1,4})?$')

If you want to accept lowercase letters (e.g., a00.0), add the re.IGNORECASE flag:

icd10_pattern = re.compile(r'^[A-Z][0-9]{2}(?:\.[A-Z0-9]{1,4})?$', re.IGNORECASE)

Breakdown of the Pattern

Let's walk through each part to make sure it makes sense:

  • ^: Anchors the match to the start of the string (prevents partial matches at the beginning)
  • [A-Z]: Matches the required initial uppercase letter
  • [0-9]{2}: Matches exactly two digits (aligns with your original [A-Z][0-9][0-9] segment)
  • (?:\.[A-Z0-9]{1,4})?: This is the optional ending segment:
    • ?:: Marks this as a non-capturing group (we don't need to extract this part separately)
    • \.: The required dot—this only exists if trailing characters are present
    • [A-Z0-9]{1,4}: Matches 1 to 4 alphanumeric characters (letters or digits)
    • ?: Makes the entire ending group optional (so codes without the trailing segment are valid)
  • $: Anchors the match to the end of the string (prevents extra characters after the code)

Test Cases

Let's verify with your examples and edge cases:

Valid Codes

These should all return True:

valid_codes = [
    "A00",          # No trailing segment
    "A00.0",        # .0
    "A00.A9",       # .A9
    "A00.A9A9",     # .A9A9
    "A00.ZZZZ",     # .ZZZZ
    "A00.9999",     # .9999
    "B12.3X"        # Mixed digits and letters
]

for code in valid_codes:
    print(f"{code}: {icd10_pattern.match(code) is not None}")

Invalid Codes

These should all return False:

invalid_codes = [
    "A0",           # Too short initial segment (only 2 characters)
    "A000",         # Extra digit without a dot
    "A00.",         # Dot with no characters after
    "A00.A9A9A",    # More than 4 characters after dot
    "a00.0",        # Lowercase initial letter (without IGNORECASE flag)
    "A00-0"         # Hyphen instead of dot
]

for code in invalid_codes:
    print(f"{code}: {icd10_pattern.match(code) is not None}")

Notes

  • You mentioned you know some ICD-10 codes have specific range restrictions, but since you don't need those, this pattern keeps it general as requested.
  • The non-capturing group (?:) is optional if you want to capture the trailing segment—just remove the ?: if you need to extract that part later.

内容的提问来源于stack exchange,提问作者DanielBell99

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 15:38:09