Python正则如何捕获Categories : 后所有逗号分隔子串及起止索引
问题原因
你当前使用的正则r"(?<=Categories : )([A-Za-z ]+(?:,)?)+"中,捕获组处于重复量词+的作用范围内,属于重复捕获场景,Python标准正则库只会保留捕获组最后一次匹配的结果,因此最终仅能拿到最后一个子串Mid Size。
实现代码
推荐先定位到Categories : 的起始位置,再单独处理后续的分类项,逻辑更清晰,也能准确计算每个子串在原文档中的起止索引:
import re def extract_categories(doc_tex): # 先匹配到目标行的分类内容段 prefix_pattern = r"Categories : (.+)" prefix_match = re.search(prefix_pattern, doc_tex) if not prefix_match: return [] # 记录分类内容段在原文档中的起始偏移量 content_offset = prefix_match.start(1) category_content = prefix_match.group(1) result = [] # 匹配所有分类项,自动跳过项前后的空格与分隔逗号 for item_match in re.finditer(r"[A-Za-z ]+?(?=\s*,|$)", category_content): raw_item = item_match.group() # 去掉子串前后多余空格 cleaned_item = raw_item.strip() # 计算在原文档中的真实起止索引 real_start = content_offset + item_match.start() + (len(raw_item) - len(raw_item.lstrip())) real_end = real_start + len(cleaned_item) result.append({ "content": cleaned_item, "start_index": real_start, "end_index": real_end }) return result # 测试示例 doc = "Categories : Turbo Prop , Very Light , Light , Mid Size" print(extract_categories(doc))
运行上述代码即可得到所有分类子串及其对应的起止索引,示例输出如下:
[ {'content': 'Turbo Prop', 'start_index': 13, 'end_index': 23}, {'content': 'Very Light', 'start_index': 25, 'end_index': 35}, {'content': 'Light', 'start_index': 37, 'end_index': 42}, {'content': 'Mid Size', 'start_index': 44, 'end_index': 52} ]
内容的提问来源于stack exchange,提问作者Tommy
相关产品推荐
相关产品推荐

