You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3.12正则将指定代码字符串转为C++原始字符串格式

使用Python正则表达式将类C代码中的指定字符串转为C原始字符串格式

需求说明

用Python 3.12.0的正则表达式,将代码源文件中**指定变量(docstring和some_detailed_notes)**的字符串,转换为C++原始字符串格式,要求:

  • 完全保留原字符串内容(包括缩进、特殊字符、转义双引号\")
  • 仅处理目标变量的字符串,不修改其他变量的字符串
  • 可选支持带括号包裹的字符串声明转换

示例输入文件

文件myclassA.txt中的类C++代码:

class myClassA(myBaseClass) {

  docstring =  // End-of-line-comments possible
"
This is my class description docstring stored in a string variable inherited from
myBaseClass.
The content of this string, INCLUDING INDENTATION, MUST NOT be changed.
";

  // This variable is declared as string has a mismatched indentation:
string some_detailed_notes = "Some string like this is also possible.
  Also with
    some really
   strange indentation
  and `inline code`, ```code blocks``` $\text{LaTeX}$ $$\frac{1}{2}$$
:::{hint}
some admonitions
:::
and all kind of special characters such as:
\"
(...)
{...}
[...]
;.,
etc.

which MUST be preserved.";

  string some_other_string = "
    do NOT touch this, only docstring and some_detailed_notes
  ";
}

预期输出

仅转换docstring和some_detailed_notes的字符串为C++原始格式,其他字符串保持不变:

class myClassA(myBaseClass) {

  docstring =  // End-of-line-comments possible
R""""(
This is my class description docstring stored in a string variable inherited from
myBaseClass.
The content of this string, INCLUDING INDENTATION, MUST NOT be changed.
)"""";

  // This variable is declared as string has a mismatched indentation:
string some_detailed_notes = R""""(Some string like this is also possible.
  Also with
    some really
   strange indentation
  and `inline code`, ```code blocks``` $\text{LaTeX}$ $$\frac{1}{2}$$
:::{hint}
some admonitions
:::
and all kind of special characters such as:
\"
(...)
{...}
[...]
;.,
etc.

which MUST be preserved.)"""";

  string some_other_string = "
    do NOT touch this, only docstring and some_detailed_notes
  ";
}

已尝试方案的问题

之前的正则表达式错误地将整个变量声明(包括注释和赋值符号)包裹进了原始字符串标记,且未正确匹配some_detailed_notes的赋值语句,导致转换结果完全不符合预期:

import pathlib
import re

fp = pathlib.Path('myclassA.txt').resolve(True)
txt = fp.read_text(encoding='utf-8')

cpp_rawstr_start = 'R""""('
cpp_rawstr_end = ')""""'

txt_sub = re.sub(
    r'([ ]*(docstring|some_detailed_notes)\s*=[\s\t\n]*")(?:(?!")(?:\\.|[^\\]))*";',
    fr'{cpp_rawstr_start}\n\g<0>\n{cpp_rawstr_end}',
    txt,
    count=2,
    flags=re.DOTALL
)

额外需求支持

对于带括号包裹的字符串声明:

docstring = (
"
Enclosed in parentheses
"
);

转换为(可选保留括号):

docstring =
R""""(
Enclosed in parentheses
)"""";

解决方案

以下代码可正确处理所有场景,包括带注释、多行字符串、带括号的声明:

import pathlib
import re

def convert_to_raw(match):
    # 提取变量声明前缀、字符串内容和结尾分号部分
    prefix = match.group(1)
    content = match.group(3)
    suffix = match.group(4) or ''
    # 拼接为C++原始字符串格式
    return f"{prefix}R\"\"\"\"({content})\"\"\"\"{suffix}"

# 读取文件内容
file_path = pathlib.Path('myclassA.txt').resolve(True)
content = file_path.read_text(encoding='utf-8')

# 正则模式:匹配目标变量的赋值语句,支持带括号和注释的情况
# 分组1: 变量声明前缀(行首空格、变量名、=、可选左括号及注释)
# 分组2: 匹配目标变量名(用于确认,实际未使用)
# 分组3: 字符串内容(未转义双引号之间的所有内容)
# 分组4: 结尾分号及可选右括号
pattern = r'^(\s*(docstring|some_detailed_notes)\s*=\s*(?:\(\s*)?(?:.*?\s*)?)"((?:\\.|[^"\\])*)"(;\s*)?(?:\)\s*)?'
# 使用多行模式和DOTALL模式,确保匹配跨行内容
flags = re.MULTILINE | re.DOTALL

# 执行替换
updated_content = re.sub(pattern, convert_to_raw, content)

# 保存结果
file_path.write_text(updated_content, encoding='utf-8')

代码说明

  1. 正则模式:
    • re.MULTILINE让^匹配每行开头,确保正确定位变量声明行
    • re.DOTALL让.匹配换行符,支持跨行字符串内容
    • 分组设计精准分离变量声明前缀、字符串内容和结尾部分,避免误匹配其他内容
  2. 回调函数:负责将提取的部分拼接为符合要求的C++原始字符串格式
  3. 兼容性:自动处理带注释、多行字符串、带括号的声明场景,完全保留原字符串内容

内容的提问来源于stack exchange,提问作者JE_Muc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 07:40:08