如何将CSV数据以字典形式读取并另存?附指定Commit数据详情
Alright, let's solve this problem step by step. You need to read a CSV file with sha and commit columns (where the commit column contains nested dictionary data), load it as a list of dictionaries, then write that data back to a new CSV file. Here's a practical Python solution using built-in modules:
Step 1: Import Required Modules
We'll use Python's built-in csv module for handling CSV files, and ast to safely parse the string-formatted dictionary in the commit column:
import csv import ast
Step 2: Read the Original CSV and Parse Data
This code reads the CSV, converts each row to a dictionary, and parses the commit string into an actual Python dictionary:
# Initialize a list to store our parsed data parsed_data = [] # Open the input CSV file with open('input.csv', mode='r', newline='', encoding='utf-8') as infile: # Use DictReader to automatically map rows to dictionaries using headers reader = csv.DictReader(infile) for row in reader: # Parse the commit column's string into a dictionary # ast.literal_eval is safer than eval() for parsing literal structures row['commit'] = ast.literal_eval(row['commit']) parsed_data.append(row)
Step 3: Write the Parsed Data to a New CSV
Now we'll write our list of dictionaries back to a new CSV. We need to convert the commit dictionary back to a string to store it in the CSV column:
# Open the output CSV file for writing with open('output.csv', mode='w', newline='', encoding='utf-8') as outfile: # Define the column headers (matches the original CSV) fieldnames = ['sha', 'commit'] writer = csv.DictWriter(outfile, fieldnames=fieldnames) # Write the header row first writer.writeheader() # Write each parsed row to the output CSV for item in parsed_data: writer.writerow({ 'sha': item['sha'], 'commit': str(item['commit']) })
Key Notes
- Safety with
ast.literal_eval: We use this instead ofeval()because it only evaluates safe literal Python structures (dictionaries, lists, strings, numbers), preventing execution of malicious code that could be hidden in the CSV. - JSON Alternative: If your
commitdata uses double quotes (valid JSON format), you can replaceast.literal_eval(row['commit'])withjson.loads(row['commit'])andstr(item['commit'])withjson.dumps(item['commit'])for strict JSON compliance.
内容的提问来源于stack exchange,提问作者mishi ahmad

