DataFrame中apply函数效率低下,求高效提取列表首字典值的方法
Great question! The issue with your current apply(axis=1) approach is that it iterates over each row individually, which is inherently slow in pandas for large datasets like 50k rows. Instead, we can use vectorized operations—these are implemented in optimized C code and will drastically speed up your processing.
Here are two efficient solutions:
1. Using Pandas String Accessors (Simplest & Fastest)
Pandas allows you to access elements of lists and dictionaries directly with the .str accessor, which works vectorially across the entire column:
# Extract first dictionary from each list, then get values for k1 and k2 df['K1'] = df['A'].str[0].str.get('k1') df['K2'] = df['A'].str[0].str.get('k2')
This method avoids row-wise iteration entirely and should handle 50k rows in a fraction of a second. If some entries in column A might be empty lists, .str[0] will return NaN for those rows—you can handle this with .fillna() or drop them if needed.
2. Using pd.json_normalize (Good for Multiple Keys)
If you have more keys in your dictionaries and want to extract all of them at once, json_normalize is a clean option. First, we extract the first dictionary from each list, then normalize those dicts into columns:
import pandas as pd # Get the first dictionary from each row in column A first_dictionaries = df['A'].str[0] # Normalize the dicts into a DataFrame and rename columns to match your desired output normalized_df = pd.json_normalize(first_dictionaries).rename(columns={'k1': 'K1', 'k2': 'K2'}) # Join the normalized columns back to the original DataFrame df = df.join(normalized_df)
This is especially useful if you have more than just k1 and k2—it will automatically create columns for every key in the dictionaries.
Both of these methods are orders of magnitude faster than your original apply approach. For 50k rows, you should see processing times drop from over a minute to just a few milliseconds!
内容的提问来源于stack exchange,提问作者fricadelle

