如何编写Regex模式提取Hearst模式中的下位词-上位词对
问题根源
你之前写的正则(NP_[\w.]*(, NP_[\w.]*)*,? (and)? other NP_[\w.]*)匹配失败的核心原因是:NP_[\w.]*仅能匹配NP_开头后接字母、数字、下划线、小数点的内容,而你需要提取的下位词片段包含空格、or、the等不在\w范围内的字符,因此无法命中目标片段。
可行实现方案
核心思路分两步:
- 先从句子中匹配到「所有并列下位词 + and other + 上位词」的完整结构,分别提取下位词整体串和上位词
- 按分隔符拆分下位词串得到单个下位词,再按需求组装为列表或元组对
参考实现(Python)
import re # 匹配核心结构的正则:分组1是所有下位词的整体串,分组2是上位词 core_pattern = re.compile(r'((?:\s*NP_\S+[^,and]*[,and]*)+)\s+other\s+(NP_\w+)') # 拆分下位词的分隔符正则:仅按逗号、and拆分,保留or连接的完整项 split_pattern = re.compile(r'\s*,\s*|\s+and\s+') # 测试句子 test_sentences = [ "NP_kimmel faces NP_dui , NP_fleeing or NP_evading_police , and other NP_possible_charges .", "The NP_network has asked NP_big_bang_theory_co-creator_bill prady to mastermind the NP_revival , which would see the NP_return of NP_kermit the NP_frog , NP_miss_piggy , NP_fozzie_bear and other NP_old_favorites ." ] for idx, sent in enumerate(test_sentences, 1): res = core_pattern.search(sent) if not res: continue hyponym_str, hypernym = res.groups() # 拆分得到单个下位词,过滤空串 hyponym_list = [item.strip() for item in split_pattern.split(hyponym_str.strip()) if item.strip()] # 输出列表格式 print(f"第{idx}句列表输出:{hyponym_list + [hypernym]}") # 输出元组对格式 print(f"第{idx}句元组对输出:") for hyponym in hyponym_list: print(f"({hyponym}, {hypernym})") print("-" * 40)
运行结果
第1句列表输出:['NP_dui', 'NP_fleeing or NP_evading_police', 'NP_possible_charges'] 第1句元组对输出: (NP_dui, NP_possible_charges) (NP_fleeing or NP_evading_police, NP_possible_charges) ---------------------------------------- 第2句列表输出:['NP_kermit the NP_frog', 'NP_miss_piggy', 'NP_fozzie_bear', 'NP_old_favorites'] 第2句元组对输出: (NP_kermit the NP_frog, NP_old_favorites) (NP_miss_piggy, NP_old_favorites) (NP_fozzie_bear, NP_old_favorites) ----------------------------------------
内容的提问来源于stack exchange,提问作者Atif
相关产品推荐
相关产品推荐

