如何将样本特征数据整理为指定格式的Python字典?
Got it, let's walk through how to turn your raw data into the dictionary you need. It's simpler than it sounds—here's a step-by-step breakdown with Python code:
Step 1: Define Your Raw Data
First, let's structure your input into a list of lists (this is how you'd typically store this kind of data in Python):
raw_data = [ ['10', '0', '1915', '387', '1933', '402'], ['10', '0', '3350', '387', '3407', '391'], ['10', '0', '842', '505', '863', '521'], ['2', '29', '2986', '282', '3112', '300'], ['2', '29', '2753', '286', '2809', '297'] ]
Step 2: Build the Dictionary
You have two straightforward options here—either use a regular dictionary with a quick check for existing keys, or use collections.defaultdict to simplify the code:
Option 1: Regular Dictionary (No Extra Imports)
This works perfectly if you don't want to pull in any additional modules:
# Start with an empty dictionary to hold our final result result = {} for item in raw_data: # Create the key in "feature1-feature2" format using an f-string key = f"{item[0]}-{item[1]}" # Convert the last 4 string elements to integers and make a sublist values = list(map(int, item[2:])) # Add the sublist to the dictionary: append if the key already exists, create a new entry if not if key in result: result[key].append(values) else: result[key] = [values]
Option 2: Using collections.defaultdict (Cleaner Code)
This shortcut eliminates the need to check if the key exists, making the logic more concise:
from collections import defaultdict # Initialize a defaultdict that automatically creates empty lists for new keys result = defaultdict(list) for item in raw_data: key = f"{item[0]}-{item[1]}" values = list(map(int, item[2:])) # Just append directly—no conditional check needed! result[key].append(values) # (Optional) Convert back to a regular dictionary if you prefer that type result = dict(result)
Step 3: Verify the Result
When you print result, you'll get exactly the structure you requested:
print(result) # Output: # { # '10-0': [[1915, 387, 1933, 402], [3350, 387, 3407, 391], [842, 505, 863, 521]], # '2-29': [[2986, 282, 3112, 300], [2753, 286, 2809, 297]] # }
Quick Notes:
- If you want to keep the numeric values as strings instead of integers, just remove the
list(map(int, ...))part and useitem[2:]directly. - The f-string (
f"{item[0]}-{item[1]}") is the cleanest way to format your key, but you could also useitem[0] + "-" + item[1]if you prefer.
内容的提问来源于stack exchange,提问作者user9266899

