使用Python、Spacy从结构化邮件提取特定人员个人信息遇实体识别问题
Hey there! Let's work through this problem together. The core issue here is that spaCy's pre-trained NER models are optimized for unstructured natural language, not the clean, key-value formatted text in your email. Since your content follows a clear Label: Value pattern, we don't need to force spaCy's NER to do the job—there are simpler, more reliable approaches to get the data you need.
Option 1: Use Regular Expressions (Most Reliable)
Regex is perfect for this structured format because it can directly match label-value pairs, even when values have spaces (like "Bill Smith"). Here's a working code snippet:
import re # Your email content email_text = """Hi Alex, Please find my personal details below, Name: Bill Smith Company: Apple Insurance number: AB 654321C Phone: +447677679999 Birth date: 21-11-1990 City: California Regards, Bill""" # Regex pattern to capture each label and its corresponding value # Matches labels (like "Company") followed by ":", then captures everything until the next capitalized label or "Regards" pattern = r'([A-Za-z\s]+):\s*(.*?)(?=\s+[A-Z][a-z]+:| Regards)' matches = re.findall(pattern, email_text, re.DOTALL) # Convert matches into a clean dictionary personal_info = {key.strip(): value.strip() for key, value in matches} # Extract the specific fields you need target_fields = { "工作单位": personal_info["Company"], "居住城市": personal_info["City"], "联系方式": personal_info["Phone"], "出生日期": personal_info["Birth date"] } print(target_fields)
Output:
{ '工作单位': 'Apple', '居住城市': 'California', '联系方式': '+447677679999', '出生日期': '21-11-1990' }
Option 2: Simple String Parsing (For Super Clean Format)
If your email format never varies, you can split the text directly to isolate the details section, then parse each pair manually:
email_text = """Hi Alex, Please find my personal details below, Name: Bill Smith Company: Apple Insurance number: AB 654321C Phone: +447677679999 Birth date: 21-11-1990 City: California Regards, Bill""" # Isolate the details section between "below," and "Regards" details_part = email_text.split("below,")[1].split("Regards")[0].strip() # Split into individual entries and build the info dictionary entries = details_part.split(" ") personal_info = {} current_key = None for item in entries: if ":" in item: current_key = item.split(":")[0].strip() # Capture the first part of the value if it's in the same item value_part = item.split(":")[1].strip() if value_part: personal_info[current_key] = value_part elif current_key: # Append to the current key's value if it's a multi-word entry personal_info[current_key] += " " + item # Extract your target fields target_fields = { "工作单位": personal_info["Company"], "居住城市": personal_info["City"], "联系方式": personal_info["Phone"], "出生日期": personal_info["Birth date"] } print(target_fields)
Option 3: Use spaCy Matcher (If You Must Use spaCy)
If you want to stick with spaCy, forget about the default NER—use the Matcher to define custom rules for your structured labels. This lets you explicitly tell spaCy what to look for:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_sm") doc = nlp(email_text) matcher = Matcher(nlp.vocab) # Define patterns for each field we care about patterns = [ # Match "Company: [Value]" [{"LOWER": "company"}, {"ORTH": ":"}, {"IS_ALPHA": True, "OP": "+"}], # Match "City: [Value]" [{"LOWER": "city"}, {"ORTH": ":"}, {"IS_ALPHA": True, "OP": "+"}], # Match "Phone: [Value]" (handles numbers and symbols like +) [{"LOWER": "phone"}, {"ORTH": ":"}, {"LIKE_NUM": True, "OP": "+"}], # Match "Birth date: [Value]" [{"LOWER": "birth"}, {"LOWER": "date"}, {"ORTH": ":"}, {"LIKE_NUM": True, "OP": "+"}] ] # Add patterns to the matcher for pattern in patterns: matcher.add("PERSONAL_DETAILS", [pattern]) # Find matches and extract the info matches = matcher(doc) target_fields = {} for match_id, start, end in matches: span = doc[start:end] # Split into label and value label = span[:2].text.replace(":", "").strip() value = span[2:].text.strip() # Map to your desired field names if label == "Company": target_fields["工作单位"] = value elif label == "City": target_fields["居住城市"] = value elif label == "Phone": target_fields["联系方式"] = value elif label == "Birth date": target_fields["出生日期"] = value print(target_fields)
Key Takeaway
Since your email uses a strict structured format, regex or simple string parsing will always be faster and more accurate than relying on spaCy's NER (which is designed for unstructured text). Save spaCy for when you're dealing with free-form natural language where patterns aren't this clear!
内容的提问来源于stack exchange,提问作者Leanda De Araujo

