Pandas技术问题:如何展开含不同键字典的DataFrame行?
Hey there! Let's work through this dictionary key mismatch problem you're stuck on. Since you’ve already got the hang of handling cases where the keys in column c are the same, I’ll break down practical approaches for when keys differ—using a super common pandas DataFrame scenario as an example, since that’s where this issue pops up most often.
First, let’s set up a sample input that mirrors your scenario:
import pandas as pd # Sample input DataFrame df = pd.DataFrame({ 'a': [1, 2], 'b': [3, 4], 'c': [{'x': 5, 'y': 6}, {'y': 7, 'z': 8}] })
#1: Expand All Unique Keys (Fill Missing Values with NaN)
This is the go-to approach if you want a wide-format output that includes every unique key from all dictionaries, filling in NaN for rows where a key doesn’t exist.
# Expand the dictionaries in column 'c' into separate columns expanded_c = pd.json_normalize(df['c']) # Merge the expanded columns back with the original DataFrame (dropping the original 'c' column) final_df = pd.concat([df.drop('c', axis=1), expanded_c], axis=1)
Result:
| a | b | x | y | z |
|---|---|---|---|---|
| 1 | 3 | 5 | 6 | NaN |
| 2 | 4 | NaN | 7 | 8 |
#2: Convert to Long-Format Key-Value Pairs
If you prefer a flexible, long-format output that preserves every key-value pair without forcing a unified column structure, this method works great:
# Convert each dictionary into a list of (key, value) tuples df['c'] = df['c'].apply(lambda x: list(x.items())) # Explode the list into separate rows, then split into key/value columns expanded_df = df.explode('c') expanded_df[['key', 'value']] = pd.DataFrame(expanded_df['c'].tolist(), index=expanded_df.index) # Clean up by dropping the original tuple column final_df = expanded_df.drop('c', axis=1)
Result:
| a | b | key | value |
|---|---|---|---|
| 1 | 3 | x | 5 |
| 1 | 3 | y | 6 |
| 2 | 4 | y | 7 |
| 2 | 4 | z | 8 |
#3: Keep Only Shared Common Keys
If you only care about keys that exist in all dictionaries in column c, first identify the common keys, then extract just those values:
# Find keys that are present in every non-null dictionary in column 'c' common_keys = set.intersection(*[set(d.keys()) for d in df['c'].dropna()]) # Extract only the common keys from each dictionary, then expand into columns expanded_common = df['c'].apply(lambda x: {k: x[k] for k in common_keys}).apply(pd.Series) # Merge back with original DataFrame final_df = pd.concat([df.drop('c', axis=1), expanded_common], axis=1)
Result (in our sample, only 'y' is a common key):
| a | b | y |
|---|---|---|
| 1 | 3 | 6 |
| 2 | 4 | 7 |
If your use case isn’t pandas (e.g., working with raw lists of dictionaries), the core ideas still apply: either standardize all keys with missing value fills, reshape to key-value pairs, or filter for shared keys. Adjust the code to fit your specific data structure!
内容的提问来源于stack exchange,提问作者shobhu

