如何将带筛选条件的JSON路径Python字符串转换为含字典的列表?
Convert JSON Path String to Dictionary-Embedded List for Spark RDD Lookup
Got it, let's tackle this problem step by step. You need to turn that path string into the structured list with embedded dictionaries for your Spark RDD lookups, right? Here's a straightforward way to do it in Python:
Step-by-Step Approach
- First split the original path string by
/to get your initial segments. - Loop through each segment:
- If the segment has
[and], split it into the base key (likeline) and the filter condition inside the brackets. - Parse the filter condition into a dictionary: split on
=to separate key and value, strip the quotes from the value, then build the key-value pair dict. - Add both the base key and the new dictionary to your result list.
- If the segment has no brackets, just add it directly to the list.
- If the segment has
Python Code Implementation
def parse_json_path(path_str): parsed_path = [] # Split the original path into segments segments = path_str.split('/') for seg in segments: # Check if the segment contains a filter condition if '[' in seg and ']' in seg: # Split into the base key and the filter part base_key, filter_raw = seg.split('[', 1) # Remove the closing bracket from the filter string filter_clean = filter_raw.rstrip(']') # Split into key and value, strip quotes from the value filter_key, filter_val = filter_clean.split('=', 1) filter_val = filter_val.strip('"') # Add both the base key and the filter dict to the result parsed_path.append(base_key) parsed_path.append({filter_key: filter_val}) else: # No filter, add the segment directly parsed_path.append(seg) return parsed_path # Test with your example string st = """data/policy/line[Type="BusinessOwners"]/risk/coverage[Type="FuelHeldForSale"]/id""" paths = parse_json_path(st) print(paths)
Output
When you run this code, you'll get exactly the target list you need:
['data', 'policy', 'line', {'Type': 'BusinessOwners'}, 'risk', 'coverage', {'Type': 'FuelHeldForSale'}, 'id']
Quick Extension Note
If you ever need to handle multiple filter conditions in one bracket (like line[Type="BusinessOwners",Status="Active"]), you can tweak the code to split the filter string by , and process each key-value pair individually. But for your current Spark RDD use case, the above code should work perfectly.
内容的提问来源于stack exchange,提问作者mdeonte001
相关产品推荐
相关产品推荐

