如何用Pandas读取含键值对列的CSV并提取键值?
Hey there! Since you're new to Python and Pandas, let's break this down step by step—you'll have this sorted in no time.
Step 1: Read Your CSV File
First, let's get your data into a Pandas DataFrame. Make sure you've imported Pandas first:
import pandas as pd # Replace 'your_file.csv' with your actual file path df = pd.read_csv('your_file.csv')
Step 2: Convert String Dictionaries to Real Python Dictionaries
The fruit column has strings that look like dictionaries, but they're just text right now. We need to turn them into actual Python dictionaries so we can work with their keys and values. The safest way to do this is using ast.literal_eval (avoid plain eval()—it's a security risk!):
import ast # Add a new column with actual dictionaries df['fruit_dict'] = df['fruit'].apply(ast.literal_eval)
Step 3: Extract Keys and Values
Now you have a column of real dictionaries—let's pull out the keys and values based on what you need.
Option 1: Get All Keys/Values as Lists in New Columns
If you want to keep each row intact and just add columns for the keys and values:
# Extract all keys into a new column (as a list) df['fruit_keys'] = df['fruit_dict'].apply(lambda x: list(x.keys())) # Extract all values into a new column (as a list of strings) df['fruit_values'] = df['fruit_dict'].apply(lambda x: list(x.values())) # Bonus: Convert the value strings (like "1,2,3,4") into lists of integers df['fruit_values_as_integers'] = df['fruit_dict'].apply( lambda x: [list(map(int, value.split(','))) for value in x.values()] )
Option 2: Expand Each Key-Value Pair into Its Own Row
If you want to flatten the data so each key-value pair gets its own row (super useful for analysis):
# Convert each dictionary into a list of (key, value) tuples df['fruit_items'] = df['fruit_dict'].apply(lambda x: list(x.items())) # Explode the list into separate rows expanded_df = df.explode('fruit_items').reset_index(drop=True) # Split the tuples into separate columns for key and value expanded_df[['fruit_key', 'fruit_value']] = pd.DataFrame( expanded_df['fruit_items'].tolist(), index=expanded_df.index ) # Clean up: Drop the original columns you don't need anymore expanded_df = expanded_df.drop(['fruit', 'fruit_dict', 'fruit_items'], axis=1)
Example Output
If your original data looks like this:
| id | fruit |
|---|---|
| 1 | {"apple": "1,2,3,4", "orange":"5,6,7,8"} |
| 2 | {"banana": "9,10", "grape":"11,12,13"} |
The expanded DataFrame will look like this:
| id | fruit_key | fruit_value |
|---|---|---|
| 1 | apple | 1,2,3,4 |
| 1 | orange | 5,6,7,8 |
| 2 | banana | 9,10 |
| 2 | grape | 11,12,13 |
Quick Note
Make sure the strings in your fruit column are valid dictionary syntax (matching quotes, no missing commas). If you get errors with ast.literal_eval, you might need to clean up any malformed entries first!
内容的提问来源于stack exchange,提问作者user804343

