如何使用jq或Python脚本实现基于root_container键的JSON数据聚合与结构重构
Absolutely! You can solve this aggregation problem efficiently with either jq (a lightweight command-line JSON processor) or a Python script—both are well-suited to handle thousands of entries as you described. Below are step-by-step solutions for both approaches:
Using jq
jq is perfect for quick JSON transformations directly in the terminal. Here's a one-liner that handles the grouping and aggregation:
jq 'group_by(.root_container // .id) | map(if .[0].root_container then empty else .[0] + {properties: .[0].properties + {file_system: [.[1:][] | .properties.mount_point]}} end)' input.json
How it works:
group_by(.root_container // .id): Groups all entries together—computer entries use their ownidas the grouping key, while filesystem entries use theirroot_containervalue. This ensures all related entries for a single computer end up in the same group.map(...): Iterates over each grouped set of entries.if .[0].root_container then empty else ... end: Filters out groups that start with a filesystem entry (since we only care about keeping the computer entry as the base)..[] + {properties: .[0].properties + {file_system: [.[1:][] | .properties.mount_point]}}: Merges the original computer entry with a newfile_systemarray, populated with allmount_pointvalues from the filesystem entries in the same group.
Using Python Script
For more control (or if you need to integrate this into a larger workflow), a Python script is a great choice. It’s scalable for large datasets and easy to modify if your requirements change:
import json # Load the input JSON data with open('input.json', 'r') as input_file: raw_data = json.load(input_file) # Create a dictionary to store computer entries, keyed by their ID computer_map = {} # First pass: collect all computer entries and initialize the file_system array for entry in raw_data: # Identify computer entries by the absence of a root_container field if 'root_container' not in entry: entry['properties']['file_system'] = [] computer_map[entry['id']] = entry # Second pass: attach filesystem mount points to their parent computer for entry in raw_data: if 'root_container' in entry: parent_id = entry['root_container'] # Ensure the parent computer exists in our map before adding if parent_id in computer_map: computer_map[parent_id]['properties']['file_system'].append(entry['properties']['mount_point']) # Convert the dictionary values back to a list for the final JSON structure final_result = list(computer_map.values()) # Print the result or save to a file print(json.dumps(final_result, indent=2)) # with open('output.json', 'w') as output_file: # json.dump(final_result, output_file, indent=2)
How it works:
- First Pass: We iterate through all entries to collect computers, initializing an empty
file_systemarray for each. Storing them in a dictionary ensures O(1) lookups later. - Second Pass: We loop through entries again, this time targeting filesystem entries. We use their
root_containerto find the parent computer in our dictionary and append themount_pointto itsfile_systemarray. - Final Conversion: We convert the dictionary values to a list to match your desired output structure.
内容的提问来源于stack exchange,提问作者baggio10
相关产品推荐
相关产品推荐

