如何用regex仅提取包含and other结构的Hearst模式匹配结果
可行解决方案
你需要将原有正则模式调整为强制匹配包含and other的NP并列结构,修改后的完整可运行代码如下:
import re annotated = ["NP_eliot_spitzer will preview his NP_first_executive_budget in a NP_speech on NP_friday_afternoon , NP_laying out his NP_plan to rein in NP_medicaid and other NP_health_care_spending , NP_setting the NP_stage for a NP_battle with politically NP_powerful_health_care_providers and NP_unions .","NP_galloway 's NP_account is among about 150 NP_new_complaints that have emerged from 44 NP_secure_state_schools , NP_halfway_houses and NP_residential_youth_care_programs in NP_texas as a NP_result of NP_several__overlapping_inquiries into NP_accusations of NP_sexual_abuse and other NP_mistreatment there "] # 修改后的正则:仅匹配用and other连接的NP结构,兼容连接词前后可能存在的逗号、多余空格 p = r'NP_[\w.]+\s*,?\s*and other\s+NP_[\w.]+' for ann in annotated: for match in re.findall(p, ann): matches = re.findall(r'NP_[\w.]*\b', match) matches = [w.replace('NP_', '').replace('_', ' ') for w in matches] print(matches)
运行后输出结果完全符合你的需求:
['medicaid', 'health care spending'] ['sexual abuse', 'mistreatment']
内容的提问来源于stack exchange,提问作者Atif
相关产品推荐
相关产品推荐

