如何在Python中提取‘Pickup Location’等特定关键词后的内容?
问题原因
你之前的代码仅截取了External Identifier:之后的全部文本,未限制到该行结束,导致输出包含后续所有无关内容,不符合预期。
正确实现方案
下面提供两种实用的提取方法:
方法1:逐行遍历提取
将文本按行拆分,逐个匹配目标关键词,提取冒号后的内容:
text = """Pickup Location: ccirc:USD-COPLEY :USD Copley Circulation External Identifier: 580978 Requested Bar code: 31822044004612 Barcode: Title: Liberalism and the postcolony : thinking the state in 20th-century Philippines / Library: Geisel Library Location: Books Floor5 Call Number: DS685 .C57 2017 Author: Claudio, Lisandro E., UC San Diego Library Contact Us""" # 指定要提取的关键词 target_keys = ['Pickup Location', 'External Identifier', 'Location'] extracted = {} # 按行处理文本 for line in text.split('\n'): line = line.strip() if not line: continue # 遍历关键词匹配 for key in target_keys: if line.startswith(f"{key}:"): # 分割冒号并提取值,去除多余空格 value = line.split(':', 1)[1].strip() extracted[key] = value break # 打印结果 for k, v in extracted.items(): print(f"{k}: {v}")
方法2:正则表达式匹配
用正则精准匹配每行的关键词与对应值:
import re text = """Pickup Location: ccirc:USD-COPLEY :USD Copley Circulation External Identifier: 580978 Requested Bar code: 31822044004612 Barcode: Title: Liberalism and the postcolony : thinking the state in 20th-century Philippines / Library: Geisel Library Location: Books Floor5 Call Number: DS685 .C57 2017 Author: Claudio, Lisandro E., UC San Diego Library Contact Us""" # 正则匹配模式,匹配目标关键词及其后的值 pattern = r'^(Pickup Location|External Identifier|Location):\s*(.*)$' matches = re.findall(pattern, text, re.MULTILINE) # 整理成字典 extracted = {key: value.strip() for key, value in matches} # 打印结果 for k, v in extracted.items(): print(f"{k}: {v}")
运行结果
两种方法都会输出:
- Pickup Location: ccirc:USD-COPLEY :USD Copley Circulation
- External Identifier: 580978
- Location: Books Floor5
内容的提问来源于stack exchange,提问作者Kaung Min Khant
相关产品推荐
相关产品推荐

