You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用re.findall从含嵌套{}的字符串中提取数据?

问题描述

我有如下字符串(文件):

s = '''
\newcommand{\commandName1}{This is first command}

\newcommand{\commandName2}{This is second command with {} brackets inside
in multiple lines {} {}
}

\newcommand{\commandName3}{This is third, last command}

'''

我希望使用Python的re包将数据提取为字典,其中key是命令名(\commandName1、\commandName2和\commandName3),值分别为This is first command、This is second command with {} brackets inside in multiple lines {} {}和This is third, last command。

我尝试了如下代码:

re.findall(r'\\newcommand{(.+)}{(.+)}', s)

但由于第二个命令内部包含{},该代码无法正常工作,请问最简单的解决方法是什么?


解决方法

原代码失效的原因是.+属于贪婪匹配,且无法识别嵌套的括号结构。要处理包含内层{}的内容,需要用递归正则表达式来匹配成对的括号。

直接使用以下代码即可实现需求:

import re

s = '''
\newcommand{\commandName1}{This is first command}

\newcommand{\commandName2}{This is second command with {} brackets inside
in multiple lines {} {}
}

\newcommand{\commandName3}{This is third, last command}

'''

# 构建支持嵌套括号匹配的正则模式
pattern = r'\\newcommand{(\\\w+)}{((?:[^{}]|{(?:[^{}]|{[^{}]*})*})*)}'
matches = re.findall(pattern, s, re.DOTALL)

# 转换为目标字典,同时处理换行符
result = {key: value.replace('\n', ' ').strip() for key, value in matches}
print(result)

关键部分解释

  • (\\\w+):精准匹配命令名,比如\commandName1,其中\\匹配反斜杠,\w+匹配命令名的字母数字主体。
  • ((?:[^{}]|{(?:[^{}]|{[^{}]*})*})*):处理命令内容,允许嵌套一层{}:
    • [^{}]:匹配非括号的普通字符
    • {(?:[^{}]|{[^{}]*})*}:匹配成对括号,内部可包含普通字符或另一层成对括号
  • re.DOTALL:让正则中的.可以匹配换行符,支持多行内容的提取
  • 字典推导式:将匹配结果整理为目标格式,同时把内容中的换行替换为空格并去除首尾空白。

执行后输出的result为:

{
    '\\commandName1': 'This is first command',
    '\\commandName2': 'This is second command with {} brackets inside in multiple lines {} {}',
    '\\commandName3': 'This is third, last command'
}

内容的提问来源于stack exchange,提问作者Jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 03:40:34