You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python、Spacy从结构化邮件提取特定人员个人信息遇实体识别问题

Solution for Extracting Structured Personal Info from Email Content

Hey there! Let's work through this problem together. The core issue here is that spaCy's pre-trained NER models are optimized for unstructured natural language, not the clean, key-value formatted text in your email. Since your content follows a clear Label: Value pattern, we don't need to force spaCy's NER to do the job—there are simpler, more reliable approaches to get the data you need.

Option 1: Use Regular Expressions (Most Reliable)

Regex is perfect for this structured format because it can directly match label-value pairs, even when values have spaces (like "Bill Smith"). Here's a working code snippet:

import re

# Your email content
email_text = """Hi Alex, Please find my personal details below, Name: Bill Smith Company: Apple Insurance number: AB 654321C Phone: +447677679999 Birth date: 21-11-1990 City: California Regards, Bill"""

# Regex pattern to capture each label and its corresponding value
# Matches labels (like "Company") followed by ":", then captures everything until the next capitalized label or "Regards"
pattern = r'([A-Za-z\s]+):\s*(.*?)(?=\s+[A-Z][a-z]+:| Regards)'
matches = re.findall(pattern, email_text, re.DOTALL)

# Convert matches into a clean dictionary
personal_info = {key.strip(): value.strip() for key, value in matches}

# Extract the specific fields you need
target_fields = {
    "工作单位": personal_info["Company"],
    "居住城市": personal_info["City"],
    "联系方式": personal_info["Phone"],
    "出生日期": personal_info["Birth date"]
}

print(target_fields)

Output:

{
    '工作单位': 'Apple',
    '居住城市': 'California',
    '联系方式': '+447677679999',
    '出生日期': '21-11-1990'
}

Option 2: Simple String Parsing (For Super Clean Format)

If your email format never varies, you can split the text directly to isolate the details section, then parse each pair manually:

email_text = """Hi Alex, Please find my personal details below, Name: Bill Smith Company: Apple Insurance number: AB 654321C Phone: +447677679999 Birth date: 21-11-1990 City: California Regards, Bill"""

# Isolate the details section between "below," and "Regards"
details_part = email_text.split("below,")[1].split("Regards")[0].strip()

# Split into individual entries and build the info dictionary
entries = details_part.split(" ")
personal_info = {}
current_key = None

for item in entries:
    if ":" in item:
        current_key = item.split(":")[0].strip()
        # Capture the first part of the value if it's in the same item
        value_part = item.split(":")[1].strip()
        if value_part:
            personal_info[current_key] = value_part
    elif current_key:
        # Append to the current key's value if it's a multi-word entry
        personal_info[current_key] += " " + item

# Extract your target fields
target_fields = {
    "工作单位": personal_info["Company"],
    "居住城市": personal_info["City"],
    "联系方式": personal_info["Phone"],
    "出生日期": personal_info["Birth date"]
}

print(target_fields)

Option 3: Use spaCy Matcher (If You Must Use spaCy)

If you want to stick with spaCy, forget about the default NER—use the Matcher to define custom rules for your structured labels. This lets you explicitly tell spaCy what to look for:

import spacy
from spacy.matcher import Matcher

nlp = spacy.load("en_core_web_sm")
doc = nlp(email_text)

matcher = Matcher(nlp.vocab)

# Define patterns for each field we care about
patterns = [
    # Match "Company: [Value]"
    [{"LOWER": "company"}, {"ORTH": ":"}, {"IS_ALPHA": True, "OP": "+"}],
    # Match "City: [Value]"
    [{"LOWER": "city"}, {"ORTH": ":"}, {"IS_ALPHA": True, "OP": "+"}],
    # Match "Phone: [Value]" (handles numbers and symbols like +)
    [{"LOWER": "phone"}, {"ORTH": ":"}, {"LIKE_NUM": True, "OP": "+"}],
    # Match "Birth date: [Value]"
    [{"LOWER": "birth"}, {"LOWER": "date"}, {"ORTH": ":"}, {"LIKE_NUM": True, "OP": "+"}]
]

# Add patterns to the matcher
for pattern in patterns:
    matcher.add("PERSONAL_DETAILS", [pattern])

# Find matches and extract the info
matches = matcher(doc)
target_fields = {}
for match_id, start, end in matches:
    span = doc[start:end]
    # Split into label and value
    label = span[:2].text.replace(":", "").strip()
    value = span[2:].text.strip()
    # Map to your desired field names
    if label == "Company":
        target_fields["工作单位"] = value
    elif label == "City":
        target_fields["居住城市"] = value
    elif label == "Phone":
        target_fields["联系方式"] = value
    elif label == "Birth date":
        target_fields["出生日期"] = value

print(target_fields)

Key Takeaway

Since your email uses a strict structured format, regex or simple string parsing will always be faster and more accurate than relying on spaCy's NER (which is designed for unstructured text). Save spaCy for when you're dealing with free-form natural language where patterns aren't this clear!

内容的提问来源于stack exchange,提问作者Leanda De Araujo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:04:34