如何使用Python和docutils提取RST文档requirement区块并存入字典
实现方案
步骤1:安装依赖
首先安装docutils库,用于解析RST文档结构:pip install docutils
步骤2:完整实现代码
from docutils.core import publish_doctree from docutils.nodes import directive_block def extract_requirements(rst_path: str) -> list[dict]: # 读取RST文件内容 with open(rst_path, 'r', encoding='utf-8') as f: rst_content = f.read() # 解析RST为文档树 doctree = publish_doctree(rst_content) requirements = [] # 遍历所有指令块节点 for node in doctree.traverse(directive_block): # 只筛选requirement类型的指令 if node.get('name') != 'requirement': continue req_item = {} # 提取requirement的ID req_item['id'] = node['args'][0].strip() # 提取所有自定义属性 for attr_name, attr_val in node.attributes.items(): if attr_name.startswith('c_'): req_item[attr_name] = attr_val # 提取requirement区块的文本内容 content_parts = [] for child in node.children: content_parts.append(child.astext().strip()) req_item['content'] = '\n'.join([p for p in content_parts if p]) requirements.append(req_item) return requirements # 测试运行 if __name__ == '__main__': reqs = extract_requirements('SWAP.rst') # 打印输出提取结果 for req in reqs: print(req) print('-'*50)
注意:如果你的RST文件中requirement的属性命名规则不是以
c_开头,可以自行修改代码中属性筛选的判断条件即可适配。
输出示例
运行上述代码后,会返回结构如下的列表:
[ { 'id': 'SWAT_FD_AR_Force.432', 'c_This_is_a': 'functional', 'c_Release': 'AR_V2.RC', 'c_Maturity': 'accepted', 'c_Implementation': 'implemented', 'content': 'text text text text\ntables' }, { 'id': 'SWAT_FD_AR_Force.231', 'c_This_is_a': 'non-functional', 'c_Release': 'AP_V1_RD1', 'c_Maturity': 'accepted', 'c_Implementation': 'implemented', 'content': 'text, tables ,etc' } ]
内容的提问来源于stack exchange,提问作者IlikeBlue LoL
相关产品推荐
相关产品推荐

