技术需求:从格式不规范文本中提取符合特定规则的编号
Alright, let's break down how to solve this problem—you've got messy, unstructured product text and need to pull out specific standardized numbers, right? Based on your sample text, I'll walk you through a practical, code-based solution using regular expressions, which is the go-to tool for this kind of text extraction.
Step 1: Define Your Target Number Rules
First, let's lock in the patterns we care about from your sample:
- NDC Drug Codes: Follow the format
XXX-XXXX-XX(3 digits, 4 digits, 2 digits separated by hyphens), almost always preceded by the "NDC" keyword - Product Line Item Identifiers: Lowercase letter followed by a right parenthesis (like
a),b),c)), used to tag different variants of the same product
Step 2: Use Regular Expressions to Extract the Data
Regex is perfect here because it can flexibly match patterns even when the surrounding text is inconsistent. Below are Python examples you can copy-paste and adapt:
Example 1: Extract NDC Codes
import re # Your sample text (replace with your actual dataset) raw_text = """Walgreens常规强度抗酸液(Alumina Magnesia Simethicone Antacid & Anti Gas)薄荷味:a)12盎司瓶装(NDC 0363-0073-02);b)26盎司瓶装(NDC 0363-0073-26),由Walgreens CO(地址:伊利诺伊州迪尔菲尔德市Wilmot路200号,邮编60015)分销;IDPN(Intradialytic Parenteral Nutrition - 添加氨基酸的透析液):a)490mL袋装;b)500mL袋装;c)590mL袋装,Pentec Health出品……""" # Regex pattern to catch NDC codes after the "NDC " prefix ndc_regex = re.compile(r'NDC (\d{3}-\d{4}-\d{2})') extracted_ndc = ndc_regex.findall(raw_text) print("Extracted NDC Codes:") for code in extracted_ndc: print(f"- {code}")
Running this will output:
Extracted NDC Codes: - 0363-0073-02 - 0363-0073-26
Example 2: Extract Line Item IDs + Corresponding Product Details
If you need to pull both the a)/b) identifiers and the product specs they're tied to, use this pattern:
# Regex to match line item IDs and their associated content (stops at semicolon or end of text) item_regex = re.compile(r'([a-z]\))([^;]+)(?:;|$)') extracted_items = item_regex.findall(raw_text) print("\nExtracted Product Items:") for item_id, details in extracted_items: print(f"{item_id} {details.strip()}")
Output:
Extracted Product Items: a) 12盎司瓶装(NDC 0363-0073-02) b) 26盎司瓶装(NDC 0363-0073-26) a) 490mL袋装 b) 500mL袋装 c) 590mL袋装
Step 3: Adapt to Edge Cases
Your text might have variations, so tweak the regex to cover them:
- If NDC codes sometimes omit hyphens: Use
r'NDC (\d{3}-?\d{4}-?\d{2})'(the?makes hyphens optional) - If "NDC" is followed by a colon or dash instead of space: Adjust to
r'NDC[ :-](\d{3}-\d{4}-\d{2})' - To extract US zip codes (like 60015 in your sample): Use
r'\b\d{5}\b'
Pro Tips
- For large datasets, wrap this logic in a script that reads text files line-by-line to avoid memory issues
- Test your regex on a small sample of your messiest text first to make sure it catches all edge cases
内容的提问来源于stack exchange,提问作者Ollie Beumkes

