Python实现高度嵌套JSON扁平化 提取点分隔全路径方法
Elasticsearch Mapping 字段全路径提取实现
核心逻辑通过深度优先递归遍历嵌套JSON结构,遍历过程中实时拼接当前层级路径,匹配到ES mapping中带type属性的实际字段节点时,将完整路径存入结果集即可,完全适配多层嵌套场景。
可直接运行的代码
def extract_es_field_paths(json_obj, parent_path="", result=None): if result is None: result = [] for key, value in json_obj.items(): # 拼接当前节点路径 current_path = f"{parent_path}.{key}" if parent_path else key # 命中ES实际字段(带type属性),加入结果集,终止当前分支递归 if isinstance(value, dict) and "type" in value: result.append(current_path) # 遇到properties节点,递归遍历下一层字段 elif isinstance(value, dict) and "properties" in value: extract_es_field_paths(value["properties"], f"{current_path}.properties", result) # 处理外层普通嵌套节点(比如mappings层) elif isinstance(value, dict): extract_es_field_paths(value, current_path, result) return result # 测试示例 if __name__ == "__main__": sample_mapping = { "mappings": { "properties": { "class": { "properties": { "name": { "properties": { "firstname": {"type": "text"}, "lastname": {"type": "text"} } }, "age": {"type": "text "} } } } } } output = extract_es_field_paths(sample_mapping) print(output)
运行结果
执行后输出和预期完全一致:
[ "mappings.properties.class.properties.name.properties.firstname", "mappings.properties.class.properties.name.properties.lastname", "mappings.properties.class.properties.age" ]
扩展说明
- 该实现专门适配ES mapping结构,不会把
type、字段类型值这类属性内容拼入路径,也不会漏过任意深层嵌套的字段 - 如果需要处理通用嵌套JSON、提取所有叶子节点的全路径,只需要把终止条件修改为「当前节点不是字典类型/字典为空」即可,通用版本代码如下:
def extract_general_json_paths(json_obj, parent_path="", result=None): if result is None: result = [] for key, value in json_obj.items(): current_path = f"{parent_path}.{key}" if parent_path else key if isinstance(value, dict) and value: extract_general_json_paths(value, current_path, result) else: result.append(current_path) return result
- 相比
pandas.json_normalize、glom等固定逻辑的工具,自定义递归的方式可以灵活控制路径拼接规则、终止条件,适配各种特殊嵌套结构的提取需求。
内容的提问来源于stack exchange,提问作者Jennifer Crosby
相关产品推荐
相关产品推荐

