You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别字符串形式Python函数中多行字符串的行索引?

获取Python多行字符串的行索引列表:正则与静态分析方案

问题描述

我有一段以字符串形式存储的Python函数代码:

def foo():
  print("hello world")
  x = 1
  for i in range(10):
    x += i
    print(x)

  multilinestring = '''hello
world
over
multiple
lines'''
  print(multilinestring)

  secondmultilinestring = '''hello
world
over
multiple
lines
again'''
  print(secondmultilinestring)

需要提取所有属于多行字符串内容的行索引列表,示例中期望结果为[7, 8, 9, 10, 11, 14, 15, 16, 17, 18, 19](行索引从0开始计数)。要求处理逻辑需满足:

  • 忽略通过反斜杠转义的三引号(如\"\"\"或\''')
  • 忽略注释#之后的三引号
  • 同时支持"""和'''两种多行字符串格式

方案一:正则表达式实现

正则可以实现该需求,但需要处理多种边缘情况,核心思路是先排除注释内容,再匹配多行字符串的起止区间,同时跳过转义的三引号。

正则表达式设计

使用以下正则(开启re.DOTALL和re.MULTILINE模式):

(?<!\\)#.*$|
(?<!\\)(\"\"\"|\''')((?:\\.|[^\\])*?)(?<!\\)\1

解释:

  • (?<!\\)#.*$:匹配未被转义的#及其后的注释内容,后续处理时跳过这些部分
  • (?<!\\)(\"\"\"|\'''):匹配未被反斜杠转义的三引号起始标记("""或''')
  • ((?:\\.|[^\\])*?):非贪婪匹配多行字符串内容,其中\\.匹配转义字符(包括转义的引号),[^\\]匹配非转义的普通字符
  • (?<!\\)\1:匹配未被转义的结束三引号,\1引用起始的引号类型

Python代码示例

import re

code = """def foo():
  print("hello world")
  x = 1
  for i in range(10):
    x += i
    print(x)

  multilinestring = '''hello
world
over
multiple
lines'''
  print(multilinestring)

  secondmultilinestring = '''hello
world
over
multiple
lines
again'''
  print(secondmultilinestring)"""

lines = code.splitlines()
line_indices = set()

# 先移除注释内容,避免干扰字符串匹配
code_no_comments = re.sub(r'(?<!\\)#.*$', '', code, flags=re.MULTILINE)

# 匹配所有多行字符串
pattern = re.compile(r'(?<!\\)(\"\"\"|\''')((?:\\.|[^\\])*?)(?<!\\)\1', flags=re.DOTALL)
for match in pattern.finditer(code_no_comments):
    # 计算匹配内容的起始和结束位置对应的行号
    start_pos = match.start()
    end_pos = match.end()
    # 统计起始位置前的换行符数量,得到起始行号
    start_line = code_no_comments[:start_pos].count('\n')
    # 统计内容部分的换行符数量,得到结束行号
    content_lines = match.group(2).count('\n')
    end_line = start_line + content_lines + 1
    # 把区间内的行号加入集合
    for line_num in range(start_line + 1, end_line + 1):
        line_indices.add(line_num)

# 转为排序后的列表
result = sorted(line_indices)
print(result)  # 输出: [7, 8, 9, 10, 11, 14, 15, 16, 17, 18, 19]

正则方案的局限性

  • 无法处理极端复杂的转义嵌套(虽然Python本身不允许多行字符串嵌套,但极端转义场景可能漏判)
  • 若代码中有字符串内包含注释符号的情况,可能出现误判(已通过先移除注释规避大部分场景)

方案二:静态分析工具(AST)实现

使用Python内置的ast模块进行抽象语法树解析,是最可靠的方案,因为它直接理解Python语法,不会出现正则的边缘情况问题。

Python代码示例

import ast

code = """def foo():
  print("hello world")
  x = 1
  for i in range(10):
    x += i
    print(x)

  multilinestring = '''hello
world
over
multiple
lines'''
  print(multilinestring)

  secondmultilinestring = '''hello
world
over
multiple
lines
again'''
  print(secondmultilinestring)"""

line_indices = set()

# 解析代码为AST树
tree = ast.parse(code)

# 遍历AST节点
for node in ast.walk(tree):
    # 处理Python3.8及以上的Constant节点(包含字符串)
    if isinstance(node, ast.Constant) and isinstance(node.value, str):
        if '\n' in node.value:
            # 转换为0-based索引
            start_line = node.lineno - 1
            end_line = node.end_lineno - 1
            for line_num in range(start_line, end_line + 1):
                line_indices.add(line_num)
    # 处理Python3.8以下的Str节点
    elif isinstance(node, ast.Str) and '\n' in node.s:
        start_line = node.lineno - 1
        end_line = node.end_lineno - 1
        for line_num in range(start_line, end_line + 1):
            line_indices.add(line_num)

# 转为排序后的列表
result = sorted(line_indices)
print(result)  # 输出: [7, 8, 9, 10, 11, 14, 15, 16, 17, 18, 19]

AST方案的优势

  • 完全遵循Python语法规则,不会误判转义字符、注释内的引号等情况
  • 能准确识别所有合法的多行字符串,包括复杂嵌套表达式中的多行字符串
  • 无需维护复杂的正则表达式,代码可读性和可维护性更高

内容的提问来源于stack exchange,提问作者Yetiowner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 13:55:15