You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术需求:从格式不规范文本中提取符合特定规则的编号

Alright, let's break down how to solve this problem—you've got messy, unstructured product text and need to pull out specific standardized numbers, right? Based on your sample text, I'll walk you through a practical, code-based solution using regular expressions, which is the go-to tool for this kind of text extraction.

Solution: Extract Targeted Numbers from Unstructured Product Text

Step 1: Define Your Target Number Rules

First, let's lock in the patterns we care about from your sample:

  • NDC Drug Codes: Follow the format XXX-XXXX-XX (3 digits, 4 digits, 2 digits separated by hyphens), almost always preceded by the "NDC" keyword
  • Product Line Item Identifiers: Lowercase letter followed by a right parenthesis (like a), b), c)), used to tag different variants of the same product

Step 2: Use Regular Expressions to Extract the Data

Regex is perfect here because it can flexibly match patterns even when the surrounding text is inconsistent. Below are Python examples you can copy-paste and adapt:

Example 1: Extract NDC Codes

import re

# Your sample text (replace with your actual dataset)
raw_text = """Walgreens常规强度抗酸液(Alumina Magnesia Simethicone Antacid & Anti Gas)薄荷味:a)12盎司瓶装(NDC 0363-0073-02);b)26盎司瓶装(NDC 0363-0073-26),由Walgreens CO(地址:伊利诺伊州迪尔菲尔德市Wilmot路200号,邮编60015)分销;IDPN(Intradialytic Parenteral Nutrition - 添加氨基酸的透析液):a)490mL袋装;b)500mL袋装;c)590mL袋装,Pentec Health出品……"""

# Regex pattern to catch NDC codes after the "NDC " prefix
ndc_regex = re.compile(r'NDC (\d{3}-\d{4}-\d{2})')
extracted_ndc = ndc_regex.findall(raw_text)

print("Extracted NDC Codes:")
for code in extracted_ndc:
    print(f"- {code}")

Running this will output:

Extracted NDC Codes:
- 0363-0073-02
- 0363-0073-26

Example 2: Extract Line Item IDs + Corresponding Product Details

If you need to pull both the a)/b) identifiers and the product specs they're tied to, use this pattern:

# Regex to match line item IDs and their associated content (stops at semicolon or end of text)
item_regex = re.compile(r'([a-z]\))([^;]+)(?:;|$)')
extracted_items = item_regex.findall(raw_text)

print("\nExtracted Product Items:")
for item_id, details in extracted_items:
    print(f"{item_id} {details.strip()}")

Output:

Extracted Product Items:
a) 12盎司瓶装(NDC 0363-0073-02)
b) 26盎司瓶装(NDC 0363-0073-26)
a) 490mL袋装
b) 500mL袋装
c) 590mL袋装

Step 3: Adapt to Edge Cases

Your text might have variations, so tweak the regex to cover them:

  • If NDC codes sometimes omit hyphens: Use r'NDC (\d{3}-?\d{4}-?\d{2})' (the ? makes hyphens optional)
  • If "NDC" is followed by a colon or dash instead of space: Adjust to r'NDC[ :-](\d{3}-\d{4}-\d{2})'
  • To extract US zip codes (like 60015 in your sample): Use r'\b\d{5}\b'

Pro Tips

  • For large datasets, wrap this logic in a script that reads text files line-by-line to avoid memory issues
  • Test your regex on a small sample of your messiest text first to make sure it catches all edge cases

内容的提问来源于stack exchange,提问作者Ollie Beumkes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:21:20