You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复Python正则,正确提取类中多行定义属性的注释?

修复Python类属性注释提取逻辑

问题根源

原正则存在两个核心缺陷:

  1. 用[^\)]*匹配字段的括号参数,无法处理括号内嵌套其他括号的情况(比如from_name("knife")中的括号),导致多行定义的字段(如backpack)无法被识别。
  2. 因第一个问题导致部分字段完全无法被捕获,自然也提取不到对应的注释。

修复方案

重新设计正则表达式,使其支持多行字段定义和括号嵌套,同时兼容带/不带注释的字段:

import re

code_block = '''class Player(Schema):
    score = fields.Float()
    """
    Total points from killing zombies and finding treasures
    """

    name = fields.String()
    age = fields.Int()

    backpack = fields.Nested(
        PlayerBackpackInventoryItem,
        missing=[PlayerBackpackInventoryItem.from_name("knife")],
    )
    """
    Collection of items that a player can store in their backpack
    """
'''


def parse_schema_comments(code):
    # 正则说明:
    # (\w+) 捕获字段名
    # fields\.\w+ 匹配fields开头的字段类型
    # (?P<parens>\((?:[^()]|(?P=parens))*\)) 递归匹配括号内容,支持嵌套和多行
    # (?:\s*"""\s*(.*?)\s*""")? 可选匹配多行三重引号注释,忽略前后空白
    pattern = r'(\w+)\s*=\s*fields\.\w+(?P<parens>\((?:[^()]|(?P=parens))*\))(?:\s*"""\s*(.*?)\s*""")?'
    matches = re.findall(pattern, code, re.DOTALL)
    
    result = []
    for match in matches:
        field_name, _, comment = match
        # 处理无注释的情况,统一转为空字符串
        comment = comment.strip() if comment else ""
        result.append((field_name, comment))
    
    return result


parsed_comments = parse_schema_comments(code_block)
print(parsed_comments)

输出结果

运行后将得到预期的正确结果:

[('score', 'Total points from killing zombies and finding treasures'), ('name', ''), ('age', ''), ('backpack', 'Collection of items that a player can store in their backpack')]

内容的提问来源于stack exchange,提问作者gameveloster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 11:02:09