You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则如何捕获Categories : 后所有逗号分隔子串及起止索引

问题原因

你当前使用的正则r"(?<=Categories : )([A-Za-z ]+(?:,)?)+"中,捕获组处于重复量词+的作用范围内,属于重复捕获场景,Python标准正则库只会保留捕获组最后一次匹配的结果,因此最终仅能拿到最后一个子串Mid Size。

实现代码

推荐先定位到Categories : 的起始位置,再单独处理后续的分类项,逻辑更清晰,也能准确计算每个子串在原文档中的起止索引:

import re

def extract_categories(doc_tex):
    # 先匹配到目标行的分类内容段
    prefix_pattern = r"Categories : (.+)"
    prefix_match = re.search(prefix_pattern, doc_tex)
    if not prefix_match:
        return []
    
    # 记录分类内容段在原文档中的起始偏移量
    content_offset = prefix_match.start(1)
    category_content = prefix_match.group(1)
    result = []

    # 匹配所有分类项,自动跳过项前后的空格与分隔逗号
    for item_match in re.finditer(r"[A-Za-z ]+?(?=\s*,|$)", category_content):
        raw_item = item_match.group()
        # 去掉子串前后多余空格
        cleaned_item = raw_item.strip()
        # 计算在原文档中的真实起止索引
        real_start = content_offset + item_match.start() + (len(raw_item) - len(raw_item.lstrip()))
        real_end = real_start + len(cleaned_item)
        result.append({
            "content": cleaned_item,
            "start_index": real_start,
            "end_index": real_end
        })
    return result

# 测试示例
doc = "Categories : Turbo Prop , Very Light , Light , Mid Size"
print(extract_categories(doc))

运行上述代码即可得到所有分类子串及其对应的起止索引,示例输出如下:

[
    {'content': 'Turbo Prop', 'start_index': 13, 'end_index': 23},
    {'content': 'Very Light', 'start_index': 25, 'end_index': 35},
    {'content': 'Light', 'start_index': 37, 'end_index': 42},
    {'content': 'Mid Size', 'start_index': 44, 'end_index': 52}
]

内容的提问来源于stack exchange,提问作者Tommy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 03:06:03