如何用Python 3.12正则将指定代码字符串转为C++原始字符串格式
使用Python正则表达式将类C代码中的指定字符串转为C原始字符串格式
需求说明
用Python 3.12.0的正则表达式,将代码源文件中**指定变量(docstring和some_detailed_notes)**的字符串,转换为C++原始字符串格式,要求:
- 完全保留原字符串内容(包括缩进、特殊字符、转义双引号
\") - 仅处理目标变量的字符串,不修改其他变量的字符串
- 可选支持带括号包裹的字符串声明转换
示例输入文件
文件myclassA.txt中的类C++代码:
class myClassA(myBaseClass) { docstring = // End-of-line-comments possible " This is my class description docstring stored in a string variable inherited from myBaseClass. The content of this string, INCLUDING INDENTATION, MUST NOT be changed. "; // This variable is declared as string has a mismatched indentation: string some_detailed_notes = "Some string like this is also possible. Also with some really strange indentation and `inline code`, ```code blocks``` $\text{LaTeX}$ $$\frac{1}{2}$$ :::{hint} some admonitions ::: and all kind of special characters such as: \" (...) {...} [...] ;., etc. which MUST be preserved."; string some_other_string = " do NOT touch this, only docstring and some_detailed_notes "; }
预期输出
仅转换docstring和some_detailed_notes的字符串为C++原始格式,其他字符串保持不变:
class myClassA(myBaseClass) { docstring = // End-of-line-comments possible R""""( This is my class description docstring stored in a string variable inherited from myBaseClass. The content of this string, INCLUDING INDENTATION, MUST NOT be changed. )""""; // This variable is declared as string has a mismatched indentation: string some_detailed_notes = R""""(Some string like this is also possible. Also with some really strange indentation and `inline code`, ```code blocks``` $\text{LaTeX}$ $$\frac{1}{2}$$ :::{hint} some admonitions ::: and all kind of special characters such as: \" (...) {...} [...] ;., etc. which MUST be preserved.)""""; string some_other_string = " do NOT touch this, only docstring and some_detailed_notes "; }
已尝试方案的问题
之前的正则表达式错误地将整个变量声明(包括注释和赋值符号)包裹进了原始字符串标记,且未正确匹配some_detailed_notes的赋值语句,导致转换结果完全不符合预期:
import pathlib import re fp = pathlib.Path('myclassA.txt').resolve(True) txt = fp.read_text(encoding='utf-8') cpp_rawstr_start = 'R""""(' cpp_rawstr_end = ')""""' txt_sub = re.sub( r'([ ]*(docstring|some_detailed_notes)\s*=[\s\t\n]*")(?:(?!")(?:\\.|[^\\]))*";', fr'{cpp_rawstr_start}\n\g<0>\n{cpp_rawstr_end}', txt, count=2, flags=re.DOTALL )
额外需求支持
对于带括号包裹的字符串声明:
docstring = ( " Enclosed in parentheses " );
转换为(可选保留括号):
docstring = R""""( Enclosed in parentheses )"""";
解决方案
以下代码可正确处理所有场景,包括带注释、多行字符串、带括号的声明:
import pathlib import re def convert_to_raw(match): # 提取变量声明前缀、字符串内容和结尾分号部分 prefix = match.group(1) content = match.group(3) suffix = match.group(4) or '' # 拼接为C++原始字符串格式 return f"{prefix}R\"\"\"\"({content})\"\"\"\"{suffix}" # 读取文件内容 file_path = pathlib.Path('myclassA.txt').resolve(True) content = file_path.read_text(encoding='utf-8') # 正则模式:匹配目标变量的赋值语句,支持带括号和注释的情况 # 分组1: 变量声明前缀(行首空格、变量名、=、可选左括号及注释) # 分组2: 匹配目标变量名(用于确认,实际未使用) # 分组3: 字符串内容(未转义双引号之间的所有内容) # 分组4: 结尾分号及可选右括号 pattern = r'^(\s*(docstring|some_detailed_notes)\s*=\s*(?:\(\s*)?(?:.*?\s*)?)"((?:\\.|[^"\\])*)"(;\s*)?(?:\)\s*)?' # 使用多行模式和DOTALL模式,确保匹配跨行内容 flags = re.MULTILINE | re.DOTALL # 执行替换 updated_content = re.sub(pattern, convert_to_raw, content) # 保存结果 file_path.write_text(updated_content, encoding='utf-8')
代码说明
- 正则模式:
re.MULTILINE让^匹配每行开头,确保正确定位变量声明行re.DOTALL让.匹配换行符,支持跨行字符串内容- 分组设计精准分离变量声明前缀、字符串内容和结尾部分,避免误匹配其他内容
- 回调函数:负责将提取的部分拼接为符合要求的C++原始字符串格式
- 兼容性:自动处理带注释、多行字符串、带括号的声明场景,完全保留原字符串内容
内容的提问来源于stack exchange,提问作者JE_Muc
相关产品推荐
相关产品推荐

