如何用Spacy从杂乱文本中匹配国家并关联对应建议?
问题:从混合建议文本中提取国家及对应建议
需要从一段包含多国建议的文本中,提取对应的国家列表和建议内容,示例输入如下:
Consider implementing the recommendations of the Special Rapporteur on violence against women and CEDAW (India) (Thailand), and France recommends to strengthen measures to increase the participation by ethnic minority women in line with CEDAW recommendations, and consider intensifying human rights education (Ghana).
期望输出为两个结构化列表:
states = [["india", "thailand"], ["ghana"]] recommendations = ["Consider implementing the recommendations of the Special Rapporteur on violence against women and CEDAW", "France recommends to strengthen measures to increase the participation by ethnic minority women in line with CEDAW recommendations, and consider intensifying human rights education"]
当前自定义分句方法无法覆盖边缘情况,包括:括号内多国家并列(如(Country1, Country2 and Country3))、括号缺失、国家拼写错误、国家名称变体(如iran/islamic republic of iran/iran (islamic republic of))。以下是基于Spacy的解决方案:
一、用Spacy Matcher定位国家标记
首先用Matcher匹配括号内的国家,同时处理括号内多国家并列的情况:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) # 匹配括号内的内容:单个国家、多个逗号分隔/and连接的国家 pattern1 = [{"ORTH": "("}, {"IS_ALPHA": True, "OP": "+"}, {"ORTH": ")"}] pattern2 = [{"ORTH": "("}, {"IS_ALPHA": True, "OP": "+"}, {"ORTH": ","}, {"IS_ALPHA": True, "OP": "+"}, {"ORTH": ")"}] pattern3 = [{"ORTH": "("}, {"IS_ALPHA": True, "OP": "+"}, {"LOWER": "and"}, {"IS_ALPHA": True, "OP": "+"}, {"ORTH": ")"}] matcher.add("COUNTRY_PAREN", [pattern1, pattern2, pattern3]) doc = nlp(your_input_text) matches = matcher(doc) # 提取匹配到的括号内文本,拆分国家 country_groups = [] for match_id, start, end in matches: span = doc[start+1:end-1] # 去掉括号 # 拆分逗号和and分隔的国家 countries = [c.strip().lower() for c in span.text.replace(" and ", ",").split(",")] country_groups.append(countries)
二、国家名称归一化处理
针对拼写变体和别名,构建映射字典,结合Spacy的NER实体识别补全:
# 国家名称映射字典,覆盖常见变体 country_normalize = { "islamic republic of iran": "iran", "iran (islamic republic of)": "iran", "iran, islamic republic of": "iran", "usa": "united states", # 补充更多变体 } # 结合NER识别无括号的国家(如示例中的France) for ent in doc.ents: if ent.label_ == "GPE": normalized = country_normalize.get(ent.text.lower(), ent.text.lower()) # 结合上下文判断该国家所属的建议片段 # 处理拼写错误:用模糊匹配映射到标准国家名 from fuzzywuzzy import process standard_countries = ["india", "thailand", "ghana", "iran", ...] for country in raw_countries: match, score = process.extractOne(country, standard_countries) if score > 80: # 设置匹配阈值 normalized_country = match
三、建议文本分割逻辑
分割建议文本时,以括号国家标记或NER识别的国家为分界点:
- 收集所有国家标记的位置(括号匹配位置、NER实体位置)
- 按位置切片文本,去除末尾的国家标记部分
- 清理切片后的文本(去除多余逗号、空格)
示例代码片段:
# 收集所有国家标记的结束位置 split_positions = [end for match_id, start, end in matches] # 补充文本首尾位置 split_positions = [0] + sorted(split_positions) + [len(doc)] # 切片生成建议文本 recommendations = [] for i in range(len(split_positions)-1): start = split_positions[i] end = split_positions[i+1] # 切片后去掉括号内容和多余符号 text = doc[start:end].text.strip() text = text.rsplit("(", 1)[0].strip().rstrip(",") if text: recommendations.append(text)
四、边缘情况处理
- 括号缺失:依赖Spacy的GPE实体识别,结合上下文判断(如"France recommends..."中的France)
- 多国家并列:通过拆分逗号和"and"提取所有关联国家
- 拼写错误:用模糊匹配将近似字符串映射到标准国家名
内容的提问来源于stack exchange,提问作者Charles Menguy
相关产品推荐
相关产品推荐

