如何使用Pandas将索引值转换为列并将嵌套字典转换为指定格式的DataFrame
Let’s tackle this two-part problem—first getting your nested dictionary into a clean, usable DataFrame, then converting index values to a dedicated column. I’ve dealt with messy nested dict structures in pandas plenty of times, so here’s a straightforward solution:
1. Convert the Nested Dictionary to the Expected DataFrame
Your original pd.DataFrame(my_dict['RuleSet']) call creates a DataFrame with columns 0, 1, 6 (the outer keys of your RuleSet dict), where each cell holds a nested dictionary. That’s almost certainly not the row-based structure you want.
Instead, use pd.json_normalize()—it’s built specifically for flattening nested JSON/dict data. Here’s how to apply it to your case:
import pandas as pd my_dict = {'RuleSet': {'0': {'RuleSetID': '0', 'RuleSetName': 'Allgemein', 'Rules': [{'RulesID': '10', 'RuleName': 'Gemeinde Seiten', 'GroupHits': '2', 'KeyWordGroups': ['100', '101', '102']}]}, '1': {'RuleSetID': '1', 'RuleSetName': 'Portale Berlin', 'Rules': [{'RulesID': '11', 'RuleName': 'Portale Berlin', 'GroupHits': '4', 'KeyWordGroups': ['100', '101', '102', '107']}]}, '6': {'RuleSetID': '6', 'RuleSetName': 'Zwangsvollstr. Berlin', 'Rules': [{'RulesID': '23', 'RuleName': 'Zwangsvollstr. Berlin', 'GroupHits': '1', 'KeyWordGroups': ['100', '101']}]}}} # Extract the core RuleSet entries (ignore the outer '0', '1', '6' keys) rule_set_data = list(my_dict['RuleSet'].values()) # Flatten the nested Rules list while retaining RuleSet metadata rules_pd = pd.json_normalize( rule_set_data, record_path='Rules', # Expand the nested Rules list into individual rows meta=['RuleSetID', 'RuleSetName'] # Keep these top-level fields as columns )
This will give you a clean, row-oriented DataFrame with columns:RulesID, RuleName, GroupHits, KeyWordGroups, RuleSetID, RuleSetName
Each row represents a single Rule, linked to its parent RuleSet—exactly the structured format you’re probably expecting.
2. Convert Index Values to a Column
There are two common approaches for this, depending on whether you want to keep the original index or reset it:
Option 1: Reset the index and save old values to a new column
Use reset_index() with the names parameter to name your new index column:
# Convert current index to a column named "original_index" and reset to default 0-based index rules_pd = rules_pd.reset_index(names='original_index')
Option 2: Keep the original index and add it as a new column
If you want to preserve the original index while adding its values as a column, assign it directly:
# Add the current index as a column without modifying the existing index rules_pd['original_index'] = rules_pd.index
Choose the option that fits your workflow—both work well, depending on whether you need to retain the original index order.
内容的提问来源于stack exchange,提问作者Daniel

