正则表达式提取Citation内容失败求助:如何正确提取目标引用文本?
正则提取租赁协议引用内容失败的解决方案
问题描述
需要从租赁协议文本中提取每个**Citation:**后引号内的内容,但运行以下代码返回空列表:
import re text = """1. **Fact:** The Tenant is obligated to obtain all consents required by law and provide copies to the Landlord upon request. **Citation:** "to obtain all consents which the law requires and give copies to the Landlord on request;" 2. **Fact:** The Tenant must pay all rates, taxes, and outgoings related to the Property, including any imposed after the date of the Lease. **Citation:** "To pay all rates, taxes and outgoings relating to the Property, including any which are imposed after the date of this Lease (even if of a novel nature)." 3. **Fact:** The Tenant is required to pay VAT on any payments made under the Lease and on any payments made by the Landlord that the Tenant agrees to reimburse. **Citation:** "To pay VAT on any payment made by the Tenant under this Lease and (except to the extent that the Landlord can reclaim it) on any payment made by the Landlord where the Tenant agrees to reimburse the Landlord" """ regex = r"\*\*Citation\:\s*?\"(?P<citation>.*?)\"" citations = re.findall(regex, text, re.MULTILINE | re.DOTALL) print(citations) # 输出空列表
错误原因
原正则表达式遗漏了**Citation:**末尾的两个星号**,实际文本中Citation:是被两个星号包裹的(即**Citation:**),但原正则只匹配了开头的**Citation:,导致无法定位到目标内容的起始位置,自然提取不到任何结果。
另外,\s*?的非贪婪空格匹配在这里没必要,改为贪婪的\s*更稳妥;同时用[^\"]*替代.*?能更精准匹配引号内的内容,避免不必要的跨行匹配(虽然本例中内容都是单行)。
修正后的代码
import re text = """1. **Fact:** The Tenant is obligated to obtain all consents required by law and provide copies to the Landlord upon request. **Citation:** "to obtain all consents which the law requires and give copies to the Landlord on request;" 2. **Fact:** The Tenant must pay all rates, taxes, and outgoings related to the Property, including any imposed after the date of the Lease. **Citation:** "To pay all rates, taxes and outgoings relating to the Property, including any which are imposed after the date of this Lease (even if of a novel nature)." 3. **Fact:** The Tenant is required to pay VAT on any payments made under the Lease and on any payments made by the Landlord that the Tenant agrees to reimburse. **Citation:** "To pay VAT on any payment made by the Tenant under this Lease and (except to the extent that the Landlord can reclaim it) on any payment made by the Landlord where the Tenant agrees to reimburse the Landlord" """ # 修正正则,补上末尾的**,并优化匹配逻辑 regex = r"\*\*Citation:\*\*\s*\"(?P<citation>[^\"]*)\"" citations = re.findall(regex, text) # 输出格式化结果 print("[") for idx, cite in enumerate(citations): if idx < len(citations)-1: print(f' "{cite}",') else: print(f' "{cite}"') print("]")
输出结果
[ "to obtain all consents which the law requires and give copies to the Landlord on request;", "To pay all rates, taxes and outgoings relating to the Property, including any which are imposed after the date of this Lease (even if of a novel nature).", "To pay VAT on any payment made by the Tenant under this Lease and (except to the extent that the Landlord can reclaim it) on any payment made by the Landlord where the Tenant agrees to reimburse the Landlord" ]
内容的提问来源于stack exchange,提问作者user1753640
相关产品推荐
相关产品推荐

