如何在多层嵌套JSON中查找指定value值并获取关联数据
Dealing with large, deeply nested JSON responses to locate specific entries and their surrounding context is exactly the kind of problem regex falls short on—regex can find the string matches, but it can’t parse the JSON structure to give you meaningful关联数据. Here’s a practical, structure-aware approach using either recursive or iterative traversal to get exactly what you need.
Core Approach
We’ll traverse every node in the JSON, keeping track of:
- The path to the current node (so you know where the match is located)
- The parent context (the immediate parent object/list containing the match)
- Optional: sibling data (other entries in the same parent object/list)
This way, when we find a value key with the value "Not Answered", we can capture all relevant surrounding data along with the match itself.
Python Implementation
Recursive Traversal (Simple for Moderate Nesting)
Since your JSON has up to 7 levels of nesting, recursion is totally safe (Python’s default recursion depth limit is 1000, so 7 is nothing). This code will recursively check every dictionary and list, collecting matches and their context:
import json def find_not_answered(node, current_path="", parent_context=None, results=None): if results is None: results = [] # Handle dictionary nodes if isinstance(node, dict): for key, value in node.items(): # Check if we've found our target entry if key == "value" and value == "Not Answered": # Collect sibling data if needed sibling_data = {} if isinstance(parent_context, dict): # Grab all other key-value pairs in the parent dict sibling_data = {k: v for k, v in parent_context.items() if k != key} elif isinstance(parent_context, list): # Grab previous/next elements in the parent list try: idx = parent_context.index(node) sibling_data = { "previous": parent_context[idx-1] if idx > 0 else None, "next": parent_context[idx+1] if idx < len(parent_context)-1 else None } except ValueError: sibling_data = None # Add the match and its context to results results.append({ "match_path": f"{current_path}.value" if current_path else "value", "parent_context": parent_context, "sibling_data": sibling_data, "matched_entry": {key: value} }) # Recurse into child nodes new_path = f"{current_path}.{key}" if current_path else key find_not_answered(value, new_path, node, results) # Handle list nodes elif isinstance(node, list): for idx, item in enumerate(node): # Recurse into list elements, adding index to the path new_path = f"{current_path}[{idx}]" if current_path else f"[{idx}]" find_not_answered(item, new_path, node, results) return results # Example usage # Load your JSON response (replace with your file or API response) with open("large_response.json", "r") as f: json_data = json.load(f) # Get all matches with context matches = find_not_answered(json_data) # Print results for i, match in enumerate(matches, 1): print(f"Match {i}:") print(f"Path: {match['match_path']}") print(f"Parent Context: {json.dumps(match['parent_context'], indent=2)}") print(f"Sibling Data: {json.dumps(match['sibling_data'], indent=2)}") print("---")
Iterative Traversal (For Extra-Large JSON)
If you’re worried about recursion stack overflow (though 7 levels isn’t an issue), you can use an iterative approach with a stack to simulate recursion. This is more memory-efficient for extremely large datasets:
import json def find_not_answered_iterative(json_data): results = [] # Stack elements: (current_node, current_path, parent_context) stack = [(json_data, "", None)] while stack: node, current_path, parent_context = stack.pop() if isinstance(node, dict): for key, value in node.items(): if key == "value" and value == "Not Answered": # Collect sibling data sibling_data = {} if isinstance(parent_context, dict): sibling_data = {k: v for k, v in parent_context.items() if k != key} elif isinstance(parent_context, list): try: idx = parent_context.index(node) sibling_data = { "previous": parent_context[idx-1] if idx > 0 else None, "next": parent_context[idx+1] if idx < len(parent_context)-1 else None } except ValueError: sibling_data = None results.append({ "match_path": f"{current_path}.value" if current_path else "value", "parent_context": parent_context, "sibling_data": sibling_data, "matched_entry": {key: value} }) # Push child node to stack new_path = f"{current_path}.{key}" if current_path else key stack.append((value, new_path, node)) elif isinstance(node, list): for idx, item in enumerate(node): new_path = f"{current_path}[{idx}]" if current_path else f"[{idx}]" stack.append((item, new_path, node)) return results # Example usage with open("large_response.json", "r") as f: json_data = json.load(f) matches = find_not_answered_iterative(json_data)
Key Advantages Over Regex
- Structure-aware: Unlike regex, this method understands JSON’s nested structure, so it won’t accidentally match
"value": "Not Answered"inside a string value. - Context-rich: You can easily modify the code to capture any level of surrounding data—parent nodes, siblings, grandparent nodes, etc.—depending on your needs.
- Scalable: Works for any size of JSON (3000+ lines) and any nesting depth up to the recursion limit (or unlimited with the iterative approach).
内容的提问来源于stack exchange,提问作者tdlanghe

