如何用Python提取指定关键词对应的关联ID?
Python提取关键词对应ID的实现
问题分析
需要从给定长字符串中,找出包含指定关键词的片段对应的ID。字符串中ID的格式有两种:id:数字 和 id=数字,每个ID后跟随对应的文本片段。
代码实现
import re def get_matching_ids(text, keyword): # 匹配所有ID及对应文本片段的正则表达式 pattern = r'id[:=](\d+),\s*(.*?)(?=,\s*id[:=]|$)' matches = re.findall(pattern, text) matching_ids = [] # 遍历匹配结果,检查关键词是否存在于文本中 for id_num, content in matches: if keyword.lower() in content.lower(): matching_ids.append(id_num) return matching_ids def format_output(ids_list): # 按需求格式化输出结果 if not ids_list: return "无匹配ID" elif len(ids_list) == 1: return ids_list[0] else: return '和'.join(ids_list) if len(ids_list) == 2 else ', '.join(ids_list[:-1]) + '和' + ids_list[-1] # 示例字符串 source_string = "id:0, an apple a day keeps the doctor away, id=1, oranges are orange not red, id=3, fox jumps over the dog, id=4, fox ate grapes" # 测试用例 print(f"搜索apple时,输出对应的ID:{format_output(get_matching_ids(source_string, 'apple'))}") print(f"搜索orange时,输出对应的ID:{format_output(get_matching_ids(source_string, 'orange'))}") print(f"搜索fox时,输出对应的ID:{format_output(get_matching_ids(source_string, 'fox'))}")
代码说明
- 正则匹配:用
id[:=](\d+),\s*(.*?)(?=,\s*id[:=]|$)精准捕获每个ID及其对应的文本片段,兼容两种ID格式。 - 关键词检查:将关键词和文本都转为小写,实现不区分大小写的匹配(若需要区分,去掉
.lower()即可)。 - 结果格式化:根据匹配到的ID数量,输出符合要求的格式(单个ID直接输出,多个ID用“和”连接)。
内容的提问来源于stack exchange,提问作者Lalas M
相关产品推荐
相关产品推荐

