嵌套字典转CSV并动态生成表头:法院爬虫数据处理需求
Got it, let's solve this problem of turning your nested court crawler data into a CSV file with automatically generated columns based on all available keys in your dictionaries. The main challenge here is handling the nested structures (like pros and contact sub-dicts) and making sure we don't miss any fields across different court entries.
Step 1: Flatten the Nested Dictionaries
First, we need to convert each nested dictionary into a flat key-value structure, where nested keys are combined with their parent keys (using a separator like . for clarity). This makes it easy to map to CSV columns.
Here's a helper function to flatten nested dicts recursively:
def flatten_dict(nested_dict, parent_key='', sep='.'): items = [] for key, value in nested_dict.items(): # Build the full key (e.g., "contact.telephone - Pay a fine:") full_key = f"{parent_key}{sep}{key}" if parent_key else key if isinstance(value, dict): # Recursively flatten nested sub-dicts items.extend(flatten_dict(value, full_key, sep=sep).items()) else: # Add non-dict values directly items.append((full_key, value)) return dict(items)
Step 2: Prepare Your Data
Let's test this with your sample data first:
# Your sample court entry my_dict = { 'url': 'https://courttribunalfinder.service.gov.uk//courts/east-berkshire-magistrates-court-slough', 'court': "East Berkshire Magistrates' Court, Slough", 'pros': {'Crown Court location code': '1920'}, 'contact': { 'telephone - Pay a fine:': '0300 790 9901', 'email - Enquiries:': 'tv-berkshiremcenq@hmcts.gsi.gov.uk', 'telephone - Fine queries:': '0186...' } } # Flatten the single entry flat_court = flatten_dict(my_dict)
If you have multiple court entries (which you will from your crawler), flatten all of them:
# Example list of court data (replace with your crawler output) court_data = [my_dict, another_court_dict, third_court_dict] # Flatten every entry in the list all_flat_courts = [flatten_dict(court) for court in court_data]
Step 3: Collect All Dynamic Column Headers
Next, we need to gather every unique key from all flattened dictionaries to use as CSV columns. This ensures we don't miss any fields that might only exist in some court entries:
# Get all unique keys across all flattened entries all_headers = list({key for court in all_flat_courts for key in court.keys()})
Step 4: Write to CSV
Finally, use Python's built-in csv.DictWriter to write the data to a CSV file. This handles missing values automatically (they'll show up as empty cells in the CSV):
import csv # Write the data to a CSV file with open('court_details.csv', 'w', newline='', encoding='utf-8') as csv_file: # Initialize the writer with our dynamic headers writer = csv.DictWriter(csv_file, fieldnames=all_headers) # Write the header row writer.writeheader() # Write each flattened court entry as a row for court in all_flat_courts: writer.writerow(court)
Key Notes
- The separator (
.) in flattened keys helps you easily identify which original nested structure a field came from (e.g.,pros.Crown Court location codeclearly belongs to theprossub-dict). - Using
encoding='utf-8'ensures special characters (like the apostrophe in your sample court name) are handled correctly. newline=''prevents extra blank lines from appearing in the CSV file on Windows systems.
内容的提问来源于stack exchange,提问作者ron_g

