Python解析非统一JSON:动态处理缺失键并拆分进程数据
Let’s walk through practical, adaptable solutions for your parsing challenges—since you can’t modify the API output, we’ll focus on defensive, flexible code that handles missing keys and inconsistent structures seamlessly.
1. Defensive Field Extraction for GRR Flows (Handling Missing Keys)
The core issue with varying GRR flow structures is avoiding crashes when keys like remoteAddress are missing. Instead of hardcoding direct key access, use a safe, dynamic extraction pattern that gracefully returns a default value (like None) when a key doesn’t exist.
Example in Python:
def extract_field(json_data, target_field): # Handle nested dicts and lists safely if isinstance(json_data, dict): # Use .get() to avoid KeyError return json_data.get(target_field, None) elif isinstance(json_data, list): # Recursively check each item in a list results = [] for item in json_data: results.append(extract_field(item, target_field)) return results else: # Not a dict/list—can't contain the target field return None # Test with a GRR flow that has remoteAddress grr_flow_1 = { "flow_id": "GRR-123", "client_details": {"remoteAddress": "192.168.1.100", "hostname": "WIN-PC"} } remote_addr_1 = extract_field(grr_flow_1, "remoteAddress") # Returns: "192.168.1.100" # Test with a GRR flow missing remoteAddress grr_flow_2 = { "flow_id": "GRR-456", "client_details": {"hostname": "LINUX-SRV"} } remote_addr_2 = extract_field(grr_flow_2, "remoteAddress") # Returns: None (no error thrown!)
For deeper nested fields (e.g., client_details.network.remoteAddress), extend the function to handle dot-separated paths:
def extract_nested_field(json_data, field_path): keys = field_path.split(".") current = json_data for key in keys: if isinstance(current, dict) and key in current: current = current[key] else: return None return current # Usage for nested fields remote_addr = extract_nested_field(grr_flow_1, "client_details.remoteAddress")
2. Splitting Two Processes’ Data
Assuming the API returns mixed process data (e.g., a list with entries from both processes), split them using unique identifiers or fields that distinguish each process.
Example: Split by Process ID
Suppose your API response looks like this:
{ "processes": [ {"proc_id": "PROCESS-A", "remoteAddress": "10.0.0.5", "status": "active"}, {"proc_id": "PROCESS-B", "remoteAddress": "10.0.0.6", "status": "idle"}, {"proc_id": "PROCESS-A", "remoteAddress": "10.0.0.7"} ] }
Split into separate groups with this code:
def split_process_data(json_data, target_proc_ids): proc_groups = {pid: [] for pid in target_proc_ids} processes = json_data.get("processes", []) for proc in processes: proc_id = proc.get("proc_id") if proc_id in proc_groups: proc_groups[proc_id].append(proc) return proc_groups # Split into PROCESS-A and PROCESS-B data split_data = split_process_data(api_response, ["PROCESS-A", "PROCESS-B"]) proc_a_data = split_data["PROCESS-A"] proc_b_data = split_data["PROCESS-B"]
If there’s no explicit process ID, split using unique fields that define each process:
def split_by_unique_fields(json_data, proc1_key, proc2_key): proc1_data = [] proc2_data = [] processes = json_data.get("processes", []) for proc in processes: if proc1_key in proc: proc1_data.append(proc) elif proc2_key in proc: proc2_data.append(proc) # Add logic here if you need to handle unclassified entries return proc1_data, proc2_data
3. Combine Both: Extract Target Fields + Split Processes
Put it all together to get only the fields you need (like remoteAddress) for each process group:
def get_proc_target_fields(api_response, proc_ids, target_field): split_procs = split_process_data(api_response, proc_ids) result = {} for pid, procs in split_procs.items(): result[pid] = [extract_field(proc, target_field) for proc in procs] return result # Get remoteAddress values for both processes remote_addrs_by_proc = get_proc_target_fields(api_response, ["PROCESS-A", "PROCESS-B"], "remoteAddress") # Output: {"PROCESS-A": ["10.0.0.5", "10.0.0.7"], "PROCESS-B": ["10.0.0.6"]}
Key Takeaways
- Never hardcode key access: Use
.get()or custom safe extraction functions to avoid crashes from missing fields. - Lean into dynamic identifiers: Split process data using unique IDs or fields instead of fixed positions in the JSON.
- Keep code modular: Break parsing into small, reusable functions (field extraction, splitting) so you can adjust logic as the API structure evolves.
内容的提问来源于stack exchange,提问作者Kenneth Van Mele

