Python解析YAML时如何为字符串/列表添加行范围元数据?
为YAML标量、序列添加行号元数据
问题背景
现有YAML配置片段:
.deploy_container: tags: - gcp image: google/cloud-sdk services: - docker:find variables: PORT: "$APP_PORT" script: - cp $GCP_SERVICE_KEY gcloud-service-key.json # Google Cloud
基于SafeLoader实现的自定义加载器,仅为映射(字典)类型添加了起止行元数据:
from yaml import SafeLoader, MappingNode from typing import Hashable, Any class SafeLineLoader(SafeLoader): def construct_mapping(self, node: MappingNode, deep: bool = False) -> dict[Hashable, Any]: mapping = super().construct_mapping(node, deep=deep) mapping['__startline__'] = node.start_mark.line + 1 mapping['__endline__'] = node.end_mark.line + 1 return mapping
解析后得到的对象中,仅字典类型包含行号信息,字符串、列表等类型没有:
{ ".deploy_container": { "tags": [ "gcp" ], "image": "google/cloud-sdk", "services": [ "docker:find" ], "variables": { "PORT": "$APP_PORT", "__startline__": 137, "__endline__": 138 }, "script": [ "cp $GCP_SERVICE_KEY gcloud-service-key.json" ], "__startline__": 131, "__endline__": 144 } }
需求:为字符串、列表等类型也添加起止行元数据,允许将这些类型转换为字典结构,例如将列表转换为包含内容和行号的字典:
"services": { "content": [ "docker:find" ], "__startline__": ..., "__endline__": ... }
解决方案
通过重写construct_sequence(处理序列/列表)和construct_scalar(处理标量/字符串、数字等)方法,将这些类型包装为包含内容和行号元数据的字典。
修改后的加载器代码:
from yaml import SafeLoader, MappingNode, SequenceNode, ScalarNode from typing import Hashable, Any, List class SafeLineLoader(SafeLoader): def construct_mapping(self, node: MappingNode, deep: bool = False) -> dict[Hashable, Any]: mapping = super().construct_mapping(node, deep=deep) mapping['__startline__'] = node.start_mark.line + 1 mapping['__endline__'] = node.end_mark.line + 1 return mapping def construct_sequence(self, node: SequenceNode, deep: bool = False) -> dict[str, Any]: sequence = super().construct_sequence(node, deep=deep) return { 'content': sequence, '__startline__': node.start_mark.line + 1, '__endline__': node.end_mark.line + 1 } def construct_scalar(self, node: ScalarNode) -> dict[str, Any]: scalar = super().construct_scalar(node) return { 'content': scalar, '__startline__': node.start_mark.line + 1, '__endline__': node.end_mark.line + 1 }
解析效果
使用上述加载器解析原始YAML后,会得到如下结构(示例行号对应原始实际位置):
{ ".deploy_container": { "tags": { "content": [ { "content": "gcp", "__startline__": 3, "__endline__": 3 } ], "__startline__": 2, "__endline__": 3 }, "image": { "content": "google/cloud-sdk", "__startline__": 4, "__endline__": 4 }, "services": { "content": [ { "content": "docker:find", "__startline__": 6, "__endline__": 6 } ], "__startline__": 5, "__endline__": 6 }, "variables": { "PORT": { "content": "$APP_PORT", "__startline__": 8, "__endline__": 8 }, "__startline__": 7, "__endline__": 8 }, "script": { "content": [ { "content": "cp $GCP_SERVICE_KEY gcloud-service-key.json # Google Cloud", "__startline__": 10, "__endline__": 10 } ], "__startline__": 9, "__endline__": 10 }, "__startline__": 1, "__endline__": 10 } }
说明
- 所有YAML节点类型(映射、序列、标量)都会被包装为包含
content字段(原始值)和行号元数据的字典 - 如果需要保留部分类型的原始结构(比如不需要包装数字类型),可以在
construct_scalar中添加类型判断,仅对字符串等目标类型进行包装 - 行号计算使用
node.start_mark.line + 1是因为YAML标记的行号从0开始,转换为用户习惯的1起始行号
内容的提问来源于stack exchange,提问作者Eliran Turgeman
相关产品推荐
相关产品推荐

