You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并两个JSON文件并保留唯一冗余对象(含嵌套结构处理)

解决嵌套JSON合并时避免空对象覆盖有效数据的问题

需求说明

需要合并两个结构相似的JSON文件,核心要求:

  • 保留biosample字段为第一个文件(File1.json)的内容
  • 合并wgs_qc_metrics下的aln_metrics和variant_metrics,空对象不能覆盖已有有效数据

输入文件

File1.json

{
    "biosample": {
        "id": "NA12878"
    },
    "wgs_qc_metrics": {
        "aln_metrics": {
            "insert_size_std_deviation": "98.2",
            "mad_autosome_coverage": "0",
            "mean_autosome_coverage": "0.00",
            "mean_insert_size": "447.3",
            "pct_autosomes_15x": "0.00",
            "pct_reads_mapped": "99.63",
            "pct_reads_properly_paired": "97.9",
            "yield_bp_q30": "1996315"
        },
        "variant_metrics": {}
    }
}

File2.json

{
    "biosample": {
        "id": "NA12878-chr14-AKT1"
    },
    "wgs_qc_metrics": {
        "aln_metrics": {},
        "variant_metrics": {
            "count_deletions": 848,
            "count_insertions": 850,
            "count_snvs": 8489,
            "ratio_heterozygous_homzygous_indel": 1.49,
            "ratio_heterozygous_homzygous_snv": 1.05,
            "ratio_insertion_deletion": 1.0,
            "ratio_transitions_transversions_snv": 2.13
        }
    }
}

预期输出

{
    "biosample": {
        "id": "NA12878"
    },
    "wgs_qc_metrics": {
        "aln_metrics": {
            "insert_size_std_deviation": "98.2",
            "mad_autosome_coverage": "0",
            "mean_autosome_coverage": "0.00",
            "mean_insert_size": "447.3",
            "pct_autosomes_15x": "0.00",
            "pct_reads_mapped": "99.63",
            "pct_reads_properly_paired": "97.9",
            "yield_bp_q30": "1996315"
        },
        "variant_metrics": {
            "count_deletions": 848,
            "count_insertions": 850,
            "count_snvs": 8489,
            "ratio_heterozygous_homzygous_indel": 1.49,
            "ratio_heterozygous_homzygous_snv": 1.05,
            "ratio_insertion_deletion": 1.0,
            "ratio_transitions_transversions_snv": 2.13
        }
    }
}

错误尝试分析

  1. Python字典解包:{**data1, **data2}或嵌套解包会直接覆盖同名键,导致File2的空aln_metrics替换File1的有效数据。
  2. jq的unique_by:该函数用于数组去重,不适用于对象合并场景,无法处理嵌套层级的覆盖问题。

解决方案

方案1:Python 针对性处理

针对已知的JSON结构,直接对目标字段进行非空判断合并:

import json

def merge_metrics(file1_path, file2_path, output_path):
    with open(file1_path, 'r') as f1:
        data1 = json.load(f1)
    with open(file2_path, 'r') as f2:
        data2 = json.load(f2)
    
    # 初始化合并结果,以file1的biosample为准
    merged = {
        "biosample": data1["biosample"],
        "wgs_qc_metrics": {}
    }
    
    # 处理aln_metrics:保留非空的那个
    merged_aln = data1["wgs_qc_metrics"]["aln_metrics"] if data1["wgs_qc_metrics"]["aln_metrics"] else data2["wgs_qc_metrics"]["aln_metrics"]
    merged["wgs_qc_metrics"]["aln_metrics"] = merged_aln
    
    # 处理variant_metrics:保留非空的那个
    merged_variant = data1["wgs_qc_metrics"]["variant_metrics"] if data1["wgs_qc_metrics"]["variant_metrics"] else data2["wgs_qc_metrics"]["variant_metrics"]
    merged["wgs_qc_metrics"]["variant_metrics"] = merged_variant
    
    with open(output_path, 'w') as f:
        json.dump(merged, f, indent=4)
        f.write("\n")

# 调用函数
merge_metrics("File1.json", "File2.json", "merged_result.json")

如果需要更通用的递归合并(支持多层嵌套,空对象不覆盖非空),可以使用以下函数:

import json

def merge_json(a, b):
    """递归合并两个字典,空对象不覆盖非空值"""
    if not isinstance(a, dict) or not isinstance(b, dict):
        return a if a != {} else b
    merged = dict(a)
    for key, value in b.items():
        if key in merged:
            merged[key] = merge_json(merged[key], value)
        else:
            merged[key] = value
    return merged

def merge_files(file1_path, file2_path, output_path):
    with open(file1_path, 'r') as f1:
        data1 = json.load(f1)
    with open(file2_path, 'r') as f2:
        data2 = json.load(f2)
    
    # 先合并所有字段,再强制替换biosample为file1的内容
    merged = merge_json(data1, data2)
    merged["biosample"] = data1["biosample"]
    
    with open(output_path, 'w') as f:
        json.dump(merged, f, indent=4)
        f.write("\n")

merge_files("File1.json", "File2.json", "merged_result.json")

方案2:jq 命令行工具

使用jq的递归合并和空值判断,实现需求:

jq -n '
    input as $file1
    | input as $file2
    | .biosample = $file1.biosample
    | .wgs_qc_metrics.aln_metrics = $file1.wgs_qc_metrics.aln_metrics // $file2.wgs_qc_metrics.aln_metrics
    | .wgs_qc_metrics.variant_metrics = $file1.wgs_qc_metrics.variant_metrics // $file2.wgs_qc_metrics.variant_metrics
' File1.json File2.json > merged_result.json

解释:

  • // 操作符表示如果左侧值为空/不存在,就使用右侧值
  • 强制指定biosample为File1的内容
  • 对每个metrics字段优先保留File1的非空数据,File1为空时才用File2的

内容的提问来源于stack exchange,提问作者user3214212

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 22:42:07