如何将带注释的CMake头文件转换为TSV/CSV文件?
用Python解析CMake头文件生成结构化CSV/TSV
针对你提到的复杂场景,下面是一个Python脚本,能处理无关注释、多cmakedefine对应单注释、特殊字符转义等情况,输出包含cmakedefine名称、简要描述、详细描述的结构化表格。
import re import csv from typing import List, Tuple def parse_cmake_header(file_path: str) -> List[Tuple[str, str, str]]: in_comment_block = False current_comments = [] results = [] # 匹配cmakedefine/define行,提取宏名称 cmake_def_pattern = re.compile(r'^\s*#\s*(?:cmakedefine|define)\s+(\w+)\b') # 匹配多行注释的起始/结束 comment_start_pattern = re.compile(r'^\s*/\*') comment_end_pattern = re.compile(r'\*/\s*$') # 匹配单行//注释 single_line_comment_pattern = re.compile(r'^\s*//\s*(.*)') with open(file_path, 'r', encoding='utf-8') as f: for line in f: stripped_line = line.strip() if not stripped_line: continue # 处理多行注释开头 if comment_start_pattern.match(stripped_line): in_comment_block = True content = stripped_line.replace('/*', '', 1).strip() if content: current_comments.append(content) continue # 处理多行注释内容/结尾 if in_comment_block: if comment_end_pattern.search(stripped_line): in_comment_block = False content = stripped_line.rsplit('*/', 1)[0].strip() if content: current_comments.append(content) else: current_comments.append(stripped_line) continue # 处理单行注释 single_line_match = single_line_comment_pattern.match(stripped_line) if single_line_match: content = single_line_match.group(1).strip() if not current_comments: current_comments.append(content) continue # 处理cmakedefine/define行 cmake_def_match = cmake_def_pattern.match(stripped_line) if cmake_def_match: macro_name = cmake_def_match.group(1) if current_comments: brief_desc = current_comments[0] detailed_desc = '\n'.join(current_comments[1:]).strip() if len(current_comments) > 1 else '' results.append((macro_name, brief_desc, detailed_desc)) else: results.append((macro_name, '', '')) # 连续cmakedefine共享同一注释,遇到非宏行再清空注释 continue # 遇到无关代码行,清空当前未绑定宏的注释 current_comments = [] return results def write_to_csv(data: List[Tuple[str, str, str]], output_path: str): def sanitize(s: str) -> str: if not s: return '' s = s.replace('\n', ' ') if ',' in s or '"' in s: s = f'"{s.replace('"', '""')}"' return s with open(output_path, 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['cmakedefine名称', '简要描述', '详细描述']) for row in data: writer.writerow([sanitize(col) for col in row]) def write_to_tsv(data: List[Tuple[str, str, str]], output_path: str): def sanitize(s: str) -> str: if not s: return '' return s.replace('\t', ' ').replace('\n', ' ') with open(output_path, 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f, delimiter='\t') writer.writerow(['cmakedefine名称', '简要描述', '详细描述']) for row in data: writer.writerow([sanitize(col) for col in row]) if __name__ == '__main__': # 替换为你的实际文件路径 input_file = 'your_cmake_header.h' csv_output = 'output.csv' tsv_output = 'output.tsv' parsed_data = parse_cmake_header(input_file) write_to_csv(parsed_data, csv_output) write_to_tsv(parsed_data, tsv_output) print(f"解析完成,已生成CSV: {csv_output} 和 TSV: {tsv_output}")
关键处理说明
- 注释块合并:自动识别单行
//和多行/* */注释,把连续的注释内容合并成一个块,避免拆分零散注释。 - 多宏共享注释:如果一个注释块后跟着多个连续的
cmakedefine,每个宏都会绑定这个注释,解决“单注释对应多宏”的场景。 - 无关注释过滤:没有对应
cmakedefine的注释会被自动丢弃,不会写入结果表格。 - 特殊字符兼容:针对CSV/TSV格式要求,自动处理逗号、引号、换行符等特殊字符,避免表格格式错乱。
使用方式
- 将脚本中的
your_cmake_header.h替换为你的CMake头文件路径。 - 运行脚本,会自动生成CSV和TSV两种格式的输出文件。
- 如需处理多个文件,可循环调用
parse_cmake_header函数,合并结果后再写入输出。
内容的提问来源于stack exchange,提问作者not2qubit
相关产品推荐
相关产品推荐

