Python实现CSV转嵌套字典(不使用Pandas)
Problem
I have a CSV file and want to process it into a nested dictionary grouped by column values. The CSV format is as follows:
sample, date, depth, analyte, result 'ABC', '01/01/2018', '3', 'LEAD', 0.22 'ABC', '02/01/2018', '3', 'LEAD', 0.25 'ABC', '01/01/2018', '5', 'LEAD', 0.19 'ABC', '02/01/2018', '5', 'LEAD', 0.18 'ABC', '01/01/2018', '3', 'MERCURY', 0.97 'ABC', '02/01/2018', '3', 'MERCURY', 0.95 'ABC', '01/01/2018', '5', 'MERCURY', 0.34 'ABC', '02/01/2018', '5', 'MERCURY', 0.11 'DEF', '01/01/2018', '3', 'LEAD', 0.07 'DEF', '02/01/2018', '3', 'LEAD', 0.04 'DEF', '01/01/2018', '5', 'LEAD', 0.16 'DEF', '02/01/2018', '5', 'LEAD', 0.65 'DEF', '01/01/2018', '3', 'MERCURY', 0.03 'DEF', '02/01/2018', '3', 'MERCURY', 0.01 'DEF', '01/01/2018', '5', 'MERCURY', 0.11 'DEF', '02/01/2018', '5', 'MERCURY', 0.13
I want the final dictionary structure to be:
dictionary = {sample: {date: {depth: [analyte, result], [analyte, result] ... }}}
I need to access unique result sets like dictionary['ABC']['01/01/2018']['5'], which should return [['LEAD', 0.19], ['MERCURY', 0.34]]. I want to avoid Pandas and need a Pythonic solution; as a beginner, nested loops got confusing, so I'm seeking help.
Solution
No worries—we can use Python's built-in csv module and collections.defaultdict to handle this nested structure cleanly, without messy conditional checks. Let's break this down:
Step 1: Import Required Modules
We'll use csv to read the file and defaultdict to simplify building the nested dictionary:
import csv from collections import defaultdict
Step 2: Build the Nested Structure
The key here is using a 3-level nested defaultdict where each missing key automatically creates the next level. The innermost level is a list that we can append our analyte-result pairs to directly:
def build_nested_csv_data(csv_file_path): # Create a nested structure: sample -> date -> depth -> list of [analyte, result] nested_data = defaultdict(lambda: defaultdict(lambda: defaultdict(list))) with open(csv_file_path, 'r') as file: # Read CSV rows as dictionaries (access values by column name) reader = csv.DictReader(file) for row in reader: # Clean up values (remove surrounding single quotes) sample = row['sample'].strip("'") date = row['date'].strip("'") depth = row['depth'].strip("'") analyte = row['analyte'].strip("'") # Convert result string to float result = float(row['result']) # Append the analyte-result pair to the correct nested list nested_data[sample][date][depth].append([analyte, result]) # Optional: Convert defaultdicts to regular dicts if preferred def convert_to_regular_dict(d): if isinstance(d, defaultdict): return {key: convert_to_regular_dict(value) for key, value in d.items()} return d return convert_to_regular_dict(nested_data)
Step 3: Use the Function
Call the function with your CSV file path, then access the data exactly as you wanted:
# Replace 'your_data.csv' with your actual file path final_dict = build_nested_csv_data('your_data.csv') # Example access print(final_dict['ABC']['01/01/2018']['5']) # Output: [['LEAD', 0.19], ['MERCURY', 0.34]]
Explanation
csv.DictReader: Makes it easy to reference columns by name (likerow['sample']) instead of index, which makes the code readable and less error-prone.defaultdict: Eliminates the need to manually check if a key exists at each level (no moreif sample not in data: data[sample] = {}). It automatically creates the missing level when you access it.- Cleaning Values: The
.strip("'")removes the single quotes from your CSV string values (e.g.,'ABC'becomesABC). - Optional Conversion: The
convert_to_regular_dictfunction turns the defaultdict structure into standard Python dicts if you don't want the auto-creation behavior anymore.
Alternative: Without defaultdict (For Learning)
If you want to see how this works with regular dicts (to understand the explicit checks), here's a version that does the same thing without defaultdict:
def build_nested_data_no_defaultdict(csv_file_path): data = {} with open(csv_file_path, 'r') as file: reader = csv.DictReader(file) for row in reader: sample = row['sample'].strip("'") date = row['date'].strip("'") depth = row['depth'].strip("'") analyte = row['analyte'].strip("'") result = float(row['result']) # Check and create each level if missing if sample not in data: data[sample] = {} if date not in data[sample]: data[sample][date] = {} if depth not in data[sample][date]: data[sample][date][depth] = [] # Append the pair data[sample][date][depth].append([analyte, result]) return data
This is what you might have tried initially, but the defaultdict version is much cleaner and more Pythonic.
内容的提问来源于stack exchange,提问作者Ian B

