Python实现JSON指定元素转数组的可复用方法求助
Hey there! Let's tackle this problem step by step—since you're jumping back into Python and need to fix that BigQuery insertion issue, I've got a reusable function that should handle exactly what you need: converting specific nested JSON fields into arrays (whether they're single objects or already arrays, we'll make sure they're wrapped properly for BigQuery's repeated record requirements).
First, let's create a recursive function that can traverse nested JSON structures and convert your target fields to arrays. This works for both top-level and deeply nested fields (like your b and c examples):
import copy def convert_fields_to_arrays(data, target_paths): """ Convert specified nested fields in a JSON dictionary to arrays. If the field is already an array, it stays as-is; if it's a single object/value, it gets wrapped in a list. Args: data (dict): Input JSON data as a Python dictionary. target_paths (list): List of field paths to convert. Each path is a list of keys (e.g., ["a", "b", "c"]). Returns: dict: Modified JSON data with target fields converted to arrays. """ def traverse(current_node, path): # We've reached the target field—convert to array if needed if not path: if not isinstance(current_node, list): return [current_node] return current_node current_key = path[0] # Skip if the key doesn't exist in the current node if current_key not in current_node: return current_node # Recurse into the next level of the nested structure current_node[current_key] = traverse(current_node[current_key], path[1:]) return current_node # Make a deep copy to avoid modifying the original input data modified_data = copy.deepcopy(data) # Process each target path one by one for path in target_paths: modified_data = traverse(modified_data, path) return modified_data
Let's test this with your examples to make sure it works:
Example 1: Simple Top-Level Field
# Your first test case simple_input = { "persons": { "city": "Seattle", "name": "Brian" } } targets = [["persons"]] result = convert_fields_to_arrays(simple_input, targets) print(result) # Output: {"persons": [{"city": "Seattle", "name": "Brian"}]}
Example 2: Complex Nested Fields
# Your complex JSON structure complex_input = { "a" : { "b" : [ { "c" : { "foo" : "bar" }, "d" : { "foo" : "bar" }, "e" : "bar" }, { "c" : [ { "foo" : "bar" } ], "d" : { "foo" : "bar" }, "e" : "bar" } ] }, "f" : { "foo" : "bar" } } # Target paths: convert "b" (under "a") and "c" (under every item in "a"->"b") targets = [["a", "b"], ["a", "b", "c"]] result = convert_fields_to_arrays(complex_input, targets) print(result)
What This Does:
a.bis already an array, so it stays unchanged.- The first
c(a single object) gets wrapped into[{"foo": "bar"}]. - The second
cis already an array, so it remains as-is.
- Loading/Saving JSON: If your data is in a file, use
json.load()to read it into a dictionary, then pass it to the function, andjson.dump()to save the modified data. - Handling Missing Fields: The function skips keys that don't exist—if you want to throw an error instead, add a check like
raise KeyError(f"Key {current_key} not found")in the traverse function. - Deep Copy: We use
copy.deepcopy()to ensure your original JSON data isn't modified—remove this if you don't mind altering the input directly.
内容的提问来源于stack exchange,提问作者pegu

