如何将Pickle加载的类JSON数据转换为指定格式的Pandas DataFrame
First, let's load the pickle data safely (using a with statement avoids leaving file handles open accidentally):
import pickle import pandas as pd # Load the pickle data properly with open("name_ethnicities.pkl", "rb") as f: data = pickle.load(f)
Next, we'll transform the dictionary into the structure you need. For each name (the key in your dict), we extract all best values from the list of dictionaries, then join them with commas:
# Process the data into a format suitable for a DataFrame processed_data = [] for name, ethnicity_entries in data.items(): # Extract all 'best' values and join with commas ethnicity_str = ", ".join([entry["best"] for entry in ethnicity_entries]) processed_data.append({"name": name, "ethnicity": ethnicity_str}) # Create the final DataFrame df = pd.DataFrame(processed_data)
If you prefer a more concise approach, you can build the DataFrame directly with a list comprehension:
df = pd.DataFrame( [ (name, ", ".join(entry["best"] for entry in ethnicity_entries)) for name, ethnicity_entries in data.items() ], columns=["name", "ethnicity"] )
Why pd.read_json didn't work?
pd.read_json is built to parse JSON-formatted files or strings into DataFrames. Your data is already a native Python dictionary loaded from a pickle file—it's not JSON data, which is why that method failed. We need to process the dictionary directly instead of treating it like JSON.
Running this code will produce exactly the output you're expecting:
| name | ethnicity |
|---|---|
| t creavalle | GreaterEuropean, British |
| uyŏng yi | Asian, GreaterEastAsian, EastAsian |
| temple orme | GreaterEuropean, British |
内容的提问来源于stack exchange,提问作者BKS

