如何使用jq从包含子数组的JSON文件中提取子数组数据并关联上层ID输出为CSV格式
使用jq将JSON中子数组条目与上层id逐行对应输出
当然可以用jq实现你想要的效果!咱们先梳理下问题,再一步步解决。
原始JSON数据
{ "success": true, "status": { "http": { "code": 200, "message": "OK" } }, "result": [{ "id": "123456789", "start_date": "2021-01-01 08:17:39.989", "snippets": [{ "transcript": "yes", "matched_entry": null, "start": "2021-01-16 11:32:25.922" }, { "transcript": null, "matched_entry": null, "start": "2021-01-16 11:32:38.179" }] }, { "id": "987654321", "start_date": "2021-01-01 08:17:39.989", "duration_total": 301, "snippets": [{ "transcript": "yes", "matched_entry": null, "start": "2021-01-01 08:17:54.055" }, { "transcript": "something", "matched_entry": " meta entry", "start": "2021-01-01 08:18:11.028" }, { "transcript": "no", "matched_entry": null, "start": "2021-01-01 08:18:24.057" }] }] }
期望输出格式
123456789, yes , null, "2021-01-16 11:32:25.922" 123456789, null, null, "2021-01-16 11:32:38.179" 987654321, yes, null, "2021-01-01 08:17:54.055" 987654321, something, "meta entry", "2021-01-01 08:18:11.028" 987654321, no, null, "2021-01-01 08:18:24.057"
你的两次尝试问题分析
- 第一次尝试:你多次使用
.snippets[],这会触发jq的笛卡尔积逻辑——每个.snippets[]都会独立遍历数组,所以生成了所有字段的组合,而非一一对应。 - 第二次尝试:你把每个字段打包成数组,虽然字段对应上了,但没法拆分成你需要的单独行输出。
正确的jq解决方案
用下面的命令就能实现你的需求:
jq -rc '.result[] | .id as $id | .snippets[] | [$id, .transcript, .matched_entry, .start] | @csv' your_json_file.json
命令解释
.result[]:遍历顶层result数组中的每个对象.id as $id:把当前result对象的id存储为变量$id,方便后续在遍历snippets时引用.snippets[]:遍历当前result对象下的每个snippets条目[$id, .transcript, .matched_entry, .start]:将id与当前snippets条目的三个字段组合成一个数组@csv:把数组转换为CSV格式的字符串,自动处理字符串引号、null的显示,完美匹配你想要的输出格式
运行这个命令后,输出结果就和你期望的完全一致啦!
内容的提问来源于stack exchange,提问作者Chrno
相关产品推荐
相关产品推荐

