从值为列表的字典列表创建DataFrame的高效方法咨询
Hey there! Great question—dealing with nested lists inside dictionaries when building pandas DataFrames can feel frustrating when your initial approach is slow. Your current pd.concat([pd.DataFrame(z) for z in test]) gets the job done, but looping to create tons of tiny DataFrames adds unnecessary overhead. Let’s walk through some far more efficient alternatives that’ll speed things up.
1. Use pd.json_normalize() (Most Recommended)
This is pandas’ built-in tool designed explicitly for handling nested JSON-like structures (which is exactly what your list of dictionaries with list values is). It’s vectorized, so it avoids the overhead of creating multiple small DataFrames.
Example Usage:
Suppose your input looks like this:
import pandas as pd test = [ {"user_id": 101, "purchases": [{"item": "apple", "price": 1.99}, {"item": "banana", "price": 0.99}]}, {"user_id": 102, "purchases": [{"item": "orange", "price": 1.49}, {"item": "grape", "price": 2.49}]} ]
You can flatten it in one line:
df = pd.json_normalize(test, record_path="purchases", meta="user_id")
This will automatically expand the nested list in purchases into rows, while keeping the user_id value aligned with each of its associated items. It’s way faster than your concat approach because it processes the data in bulk instead of iterating through each element.
2. Pre-Flatten the Data with a Python Loop (Simple & Fast)
If you prefer more control over the flattening process, you can first build a single list of flattened dictionaries, then convert it to a DataFrame in one go. This avoids creating multiple intermediate DataFrames, which is where most of your overhead comes from.
Example Usage:
flattened_data = [] for entry in test: # Extract scalar values (non-list entries) from the outer dictionary scalar_fields = {key: val for key, val in entry.items() if not isinstance(val, list)} # Iterate through the list value and merge scalar fields with each inner dictionary for item in entry["purchases"]: flattened_data.append({**scalar_fields, **item}) df = pd.DataFrame(flattened_data)
Python loops over lists are surprisingly efficient compared to creating pandas objects in a loop, so this method will outperform your concat approach by a significant margin, especially with large datasets.
3. For Simple Single-List Structures: Use explode() + pd.Series
If your dictionaries only have one list value and the rest are scalars, you can use explode() to unnest the list, then expand the nested dictionaries into columns. This is less flexible than the first two methods but works well for straightforward cases.
Example Usage:
# Create a temporary DataFrame with the original structure temp_df = pd.DataFrame(test) # Explode the list column to get one row per list item exploded_df = temp_df.explode("purchases") # Expand the nested dictionaries into separate columns df = exploded_df.join(pd.DataFrame(exploded_df["purchases"].tolist())) # Drop the original list column if needed df = df.drop("purchases", axis=1)
Why These Methods Are Better Than pd.concat
Each time you create a small DataFrame in your loop, pandas has to initialize metadata (like indexes, column objects, and data structures) for that tiny object. Multiply that by hundreds or thousands of entries, and the overhead adds up quickly. The methods above either process the data in bulk (json_normalize) or minimize pandas object creation (pre-flattening), which drastically reduces runtime.
内容的提问来源于stack exchange,提问作者Ryan Erwin

