Displacy无法识别自定义NER实体的技术求助
自定义NER实体Displacy不高亮的解决方案
核心问题与修复步骤
1. 管道加载时机错误
你先调用nlp(cleaned_text)处理文本,之后才添加entity_ruler管道。spaCy的管道只会作用于添加管道之后处理的文本,所以之前生成的doc不会包含自定义实体。
修复: 先添加entity_ruler管道,再处理文本。
2. Entity Ruler模式不匹配测试文本
你的pattern设置不符合实际文本结构:
- BUN的pattern要求前面有一个
Xxxxx形状的词(首字母大写、后面小写),但测试文本里"BUN"是独立出现的,没有前置词; - SerumUricAcid的pattern多了一个
Xxxxx形状的词,实际文本是"Serum Uric Acid"三个词,不需要前置额外词汇。
修复: 调整pattern为匹配实际文本的结构:
ner_patterns = [ {"label": "BUN", "pattern": [{"LOWER": "bun"}]}, {"label": "SerumUricAcid", "pattern": [{"LOWER": "serum"}, {"LOWER": "uric"}, {"LOWER": "acid"}]}, {"label": "SerumPotassium", "pattern": [{"LOWER": "serum"}, {"LOWER": "potassium"}]} ]
3. Displacy配置语法错误
options的ents列表中,"BUN"后面的引号未闭合,导致语法错误,Displacy无法识别自定义实体标签。
修复: 修正引号:
options = {"ents": ["BUN", "SerumUricAcid", "SerumPotassium"], "colors": colors}
修正后的完整代码
import spacy from spacy import displacy from IPython.display import HTML # Load the spaCy English model nlp = spacy.load("en_core_web_sm") # Add the custom NER patterns to the entity ruler FIRST ner_patterns = [ {"label": "BUN", "pattern": [{"LOWER": "bun"}]}, {"label": "SerumUricAcid", "pattern": [{"LOWER": "serum"}, {"LOWER": "uric"}, {"LOWER": "acid"}]}, {"label": "SerumPotassium", "pattern": [{"LOWER": "serum"}, {"LOWER": "potassium"}]} ] ruler = nlp.add_pipe("entity_ruler", after='ner') ruler.add_patterns(ner_patterns) # 测试文本(替换为你的cleaned_text) cleaned_text = """BIOCHEMISTRY KIDNEY FUNCTION TEST (KFT) TEST VALUE UNIT REFERENCE BUN 10.27 mg/dl 7.9 - 20 Serum Urea 22 mg/dl 13 - 40 Serum Creatinine H 0.9 mg/dl 0.5 - 0.8 Serum Calcium 9.0 mg/dl 8.8 - 10.6 Serum Potassium 3.9 mmol/L 3.5 - 5.1 Serum Sodium L 132 mmol/L 136 - 146 Serum Uric Acid 5 mg/dl 2.6 - 6 Urea / Creatinine Ratio 24.44 BUN / Creatinine Ratio 11.41 ~~~ End of report ~~~ LABSMART SAMPLE REPORT Patient Name: Mrs. Dummy Registered on: 09/08/2022 11:35 AM 1001 Age / Sex: 34 YRS / F Collected on: 09/08/2022 Referred By: Dr. Self Received on: 12/08/2022 Reg. no. / UHID: 1001 / Reported on: 09/08/2022 11:35 AM Investigations: Kidney Function Test (KFT) Page 1 of 1 Mr. Sachin Sharma DMLT, Lab Incharge Dr. A. K. Asthana MBBS, MD Pathologist""" # Process the cleaned text AFTER adding the pipeline doc = nlp(cleaned_text) # Extracted values extracted_values = {} # Process the tokens to extract values for token in doc: # 调整匹配逻辑,对应自定义标签 if token.ent_type_ in ["BUN", "SerumUricAcid", "SerumPotassium"]: # 找到实体后的数值 next_token = token.nbor() while next_token and not next_token.like_num: next_token = next_token.nbor() if next_token and next_token.like_num: extracted_values[token.ent_type_] = next_token.text # Print the extracted values print("Extracted values:") for label, value in extracted_values.items(): print(f"{label}: {value}") # Prepare entity highlighting with linear gradients colors = { "BUN": "linear-gradient(90deg, #aa9cfc, #fc9ce7)", "SerumUricAcid": "linear-gradient(90deg, #ffabab, #ffd8a8)", "SerumPotassium": "linear-gradient(90deg, #a8e6cf, #dcedc1)" } options = {"ents": ["BUN", "SerumUricAcid", "SerumPotassium"], "colors": colors} # Render the displaCy visualization with highlighting html = displacy.render(doc, style="ent", options=options) # Display the HTML HTML(html)
额外优化点
- 提取数值的逻辑改为基于实体类型判断,更贴合自定义NER的结果,避免匹配无关的"Serum"等单独词汇;
- 添加了循环跳过非数值词,确保找到实体对应的第一个数值。
内容的提问来源于stack exchange,提问作者Ujjwal
相关产品推荐
相关产品推荐

