You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用spaCy关联参数与对应数值?Regex匹配问题求助

用spaCy提取科研文献中的参数与对应数值

需求描述

需要从科研文本中提取参数名称及其对应数值,例如从以下句子中:

sentence = "The feed rate, aspirator rate, inlet and outlet temperature and air flow rate were approximately 3l/hr, 100%, 120C, 90C, and 357l/hr, respectively."

提取出:feed rate对应3l/hr、aspirator rate对应100%、inlet temperature对应120C、outlet temperature对应90C、air flow rate对应357l/hr。

原代码使用spaCy entity ruler未达到预期效果,无法匹配feed rate和outlet temperature,现解决以下两个问题:


问题1:提取带特殊字符的单位

原正则错误地将ml/min放在字符集[]中,导致被解析为单个字符匹配,而非整体单位。正确处理方式:

  • 直接用括号包裹单位组合,如(ml/min|l/hr|L/hr|mL/min),匹配完整单位
  • 兼容数值与单位连写(如357l/hr)和分开写(如3 ml/min)两种情况,设计对应的匹配模式

问题2:覆盖参数写法变体

针对inlet temperature、inlet air temperature这类变体,需要设计灵活的正则匹配结构:

  • 用OP: "?"标记可选修饰词(如air、gas),允许其出现0次或1次
  • 将多token参数拆分为多个字典元素,让spaCy识别连续token组成的实体

修正后的完整代码

import spacy

# 使用英文小模型,提升分词和多token实体识别能力
nlp = spacy.load("en_core_web_sm")
ruler = nlp.add_pipe("entity_ruler", before="ner")

patterns = [
    # 匹配温度类参数:兼容inlet temperature、inlet air temperature等变体
    {"label": "PARAMETER", "pattern": [
        {"LOWER": {"REGEX": r"^(inlet|outlet)$"}},
        {"LOWER": {"REGEX": r"^(air|gas)?$"}, "OP": "?"},
        {"LOWER": "temperature"}
    ]},
    # 匹配其他固定参数
    {"label": "PARAMETER", "pattern": [{"LOWER": "feed"}, {"LOWER": "rate"}]},
    {"label": "PARAMETER", "pattern": [{"LOWER": "aspirator"}, {"LOWER": "rate"}]},
    {"label": "PARAMETER", "pattern": [{"LOWER": "air"}, {"LOWER": "flow"}, {"LOWER": "rate"}]},
    
    # 匹配温度数值
    {"label": "TEMP_VALUE", "pattern": [{"TEXT": {"REGEX": r"\d+(C|K|F)"}}]},
    # 匹配百分比数值(对应aspirator rate)
    {"label": "PERCENT_VALUE", "pattern": [{"TEXT": {"REGEX": r"\d{1,5}%"}}]},
    # 匹配流量数值:兼容连写和分写两种格式
    {"label": "FLOW_VALUE", "pattern": [{"TEXT": {"REGEX": r"\d+(\.\d+)?(ml/min|l/hr|L/hr|mL/min)"}}]},
    {"label": "FLOW_VALUE", "pattern": [{"TEXT": {"REGEX": r"\d+"}}, {"TEXT": {"REGEX": r"(ml/min|l/hr|L/hr|mL/min)"}}]}
]

ruler.add_patterns(patterns)

text = "The feed rate, aspirator rate, inlet and outlet temperature and air flow rate were approximately 3 ml/min, 100%, 120C and 90C and 357 l/hr, respectively."

doc = nlp(text)

# 输出识别结果
for ent in doc.ents:
    print(f"{ent.label_}: {ent.text}")

代码说明

  1. 单位提取优化:
    • 用分组正则匹配完整单位,避免字符集解析错误
    • 新增两种流量匹配规则,覆盖数值与单位连写、分写的常见格式
  2. 参数变体兼容:
    • 通过OP: "?"实现可选修饰词的匹配,轻松覆盖inlet temperature、inlet air temperature等变体
    • 拆分多token参数为多个匹配项,确保spaCy正确识别连续token组成的实体

若需要进一步关联参数与对应数值,可利用文本中的respectively提示的顺序对应关系,或结合spaCy的依赖分析功能实现映射。

内容的提问来源于stack exchange,提问作者faz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 06:20:42