如何合并两个JSON文件并保留唯一冗余对象(含嵌套结构处理)
解决嵌套JSON合并时避免空对象覆盖有效数据的问题
需求说明
需要合并两个结构相似的JSON文件,核心要求:
- 保留
biosample字段为第一个文件(File1.json)的内容 - 合并
wgs_qc_metrics下的aln_metrics和variant_metrics,空对象不能覆盖已有有效数据
输入文件
File1.json
{ "biosample": { "id": "NA12878" }, "wgs_qc_metrics": { "aln_metrics": { "insert_size_std_deviation": "98.2", "mad_autosome_coverage": "0", "mean_autosome_coverage": "0.00", "mean_insert_size": "447.3", "pct_autosomes_15x": "0.00", "pct_reads_mapped": "99.63", "pct_reads_properly_paired": "97.9", "yield_bp_q30": "1996315" }, "variant_metrics": {} } }
File2.json
{ "biosample": { "id": "NA12878-chr14-AKT1" }, "wgs_qc_metrics": { "aln_metrics": {}, "variant_metrics": { "count_deletions": 848, "count_insertions": 850, "count_snvs": 8489, "ratio_heterozygous_homzygous_indel": 1.49, "ratio_heterozygous_homzygous_snv": 1.05, "ratio_insertion_deletion": 1.0, "ratio_transitions_transversions_snv": 2.13 } } }
预期输出
{ "biosample": { "id": "NA12878" }, "wgs_qc_metrics": { "aln_metrics": { "insert_size_std_deviation": "98.2", "mad_autosome_coverage": "0", "mean_autosome_coverage": "0.00", "mean_insert_size": "447.3", "pct_autosomes_15x": "0.00", "pct_reads_mapped": "99.63", "pct_reads_properly_paired": "97.9", "yield_bp_q30": "1996315" }, "variant_metrics": { "count_deletions": 848, "count_insertions": 850, "count_snvs": 8489, "ratio_heterozygous_homzygous_indel": 1.49, "ratio_heterozygous_homzygous_snv": 1.05, "ratio_insertion_deletion": 1.0, "ratio_transitions_transversions_snv": 2.13 } } }
错误尝试分析
- Python字典解包:
{**data1, **data2}或嵌套解包会直接覆盖同名键,导致File2的空aln_metrics替换File1的有效数据。 - jq的
unique_by:该函数用于数组去重,不适用于对象合并场景,无法处理嵌套层级的覆盖问题。
解决方案
方案1:Python 针对性处理
针对已知的JSON结构,直接对目标字段进行非空判断合并:
import json def merge_metrics(file1_path, file2_path, output_path): with open(file1_path, 'r') as f1: data1 = json.load(f1) with open(file2_path, 'r') as f2: data2 = json.load(f2) # 初始化合并结果,以file1的biosample为准 merged = { "biosample": data1["biosample"], "wgs_qc_metrics": {} } # 处理aln_metrics:保留非空的那个 merged_aln = data1["wgs_qc_metrics"]["aln_metrics"] if data1["wgs_qc_metrics"]["aln_metrics"] else data2["wgs_qc_metrics"]["aln_metrics"] merged["wgs_qc_metrics"]["aln_metrics"] = merged_aln # 处理variant_metrics:保留非空的那个 merged_variant = data1["wgs_qc_metrics"]["variant_metrics"] if data1["wgs_qc_metrics"]["variant_metrics"] else data2["wgs_qc_metrics"]["variant_metrics"] merged["wgs_qc_metrics"]["variant_metrics"] = merged_variant with open(output_path, 'w') as f: json.dump(merged, f, indent=4) f.write("\n") # 调用函数 merge_metrics("File1.json", "File2.json", "merged_result.json")
如果需要更通用的递归合并(支持多层嵌套,空对象不覆盖非空),可以使用以下函数:
import json def merge_json(a, b): """递归合并两个字典,空对象不覆盖非空值""" if not isinstance(a, dict) or not isinstance(b, dict): return a if a != {} else b merged = dict(a) for key, value in b.items(): if key in merged: merged[key] = merge_json(merged[key], value) else: merged[key] = value return merged def merge_files(file1_path, file2_path, output_path): with open(file1_path, 'r') as f1: data1 = json.load(f1) with open(file2_path, 'r') as f2: data2 = json.load(f2) # 先合并所有字段,再强制替换biosample为file1的内容 merged = merge_json(data1, data2) merged["biosample"] = data1["biosample"] with open(output_path, 'w') as f: json.dump(merged, f, indent=4) f.write("\n") merge_files("File1.json", "File2.json", "merged_result.json")
方案2:jq 命令行工具
使用jq的递归合并和空值判断,实现需求:
jq -n ' input as $file1 | input as $file2 | .biosample = $file1.biosample | .wgs_qc_metrics.aln_metrics = $file1.wgs_qc_metrics.aln_metrics // $file2.wgs_qc_metrics.aln_metrics | .wgs_qc_metrics.variant_metrics = $file1.wgs_qc_metrics.variant_metrics // $file2.wgs_qc_metrics.variant_metrics ' File1.json File2.json > merged_result.json
解释:
//操作符表示如果左侧值为空/不存在,就使用右侧值- 强制指定
biosample为File1的内容 - 对每个metrics字段优先保留File1的非空数据,File1为空时才用File2的
内容的提问来源于stack exchange,提问作者user3214212
相关产品推荐
相关产品推荐

