如何加速/替代Python双层for循环?50万级字典数据处理优化求助
Hey there! Let's break down why your current code is running so slow and fix it with practical, pandas-friendly optimizations. Dealing with 500k+ documents each with 30+ fields is totally manageable—we just need to work with pandas instead of against it.
First, Why Your Current Code Is Slow
Let's pinpoint the bottlenecks:
- Row-by-row DataFrame appends:
AppendingDataframe.loc[len(AppendingDataframe)] = ExtraListis a major killer. Every time you do this, pandas has to reallocate memory for the entire DataFrame, turning an O(n) task into O(n²) time. For 500k rows, this adds up exponentially. - Unnecessary key iteration: You're looping through every key in each document, then checking if it's "USELESS". It's way more efficient to only iterate the keys you actually care about (the ones in
AllParameters). - Potential data pollution: Reusing
ExtraListwithout resetting it could leave leftover values from previous documents if a field is missing, leading to incorrect data.
Optimization 1: Batch Build a List of Dictionaries (Fastest Approach)
The best way to handle this in pandas is to collect all your processed data into a list first, then convert it to a DataFrame in one go. This avoids the overhead of modifying the DataFrame row-by-row.
import pandas as pd # Define target columns from your AllParameters mapping target_columns = list(AllParameters.values()) c_name_column = AllParameters["C_Name"] # Ensure the filename column is included if not already present if c_name_column not in target_columns: target_columns.append(c_name_column) processed_documents = [] filename_prefix = filename.split(".")[0] # Calculate once instead of every loop for doc in ListOfDocuments: doc_data = {} # Only iterate the keys we need (skip "USELESS" upfront) for original_key, df_column in AllParameters.items(): if original_key == "USELESS": continue # Use .get() to handle missing keys gracefully (default to pd.NA for missing values) doc_data[df_column] = doc.get(original_key, pd.NA) # Add the filename field doc_data[c_name_column] = filename_prefix processed_documents.append(doc_data) # Convert the entire list to a DataFrame in one step—this is lightning fast AppendingDataframe = pd.DataFrame(processed_documents)
Optimization 2: Speed Up with List/Dictionary Comprehensions
For even more speed, replace the nested for loops with comprehensions. They're optimized in Python and reduce the overhead of explicit loop logic.
import pandas as pd filename_prefix = filename.split(".")[0] c_name_column = AllParameters["C_Name"] # Use a list comprehension to build all document data in one go processed_documents = [ { # Dictionary comprehension for core fields (skip "USELESS") **{col: doc.get(key, pd.NA) for key, col in AllParameters.items() if key != "USELESS"}, # Add the filename field c_name_column: filename_prefix } for doc in ListOfDocuments ] AppendingDataframe = pd.DataFrame(processed_documents)
Optimization 3: Chunk Processing for Low-Memory Environments
If your machine doesn't have enough RAM to hold all 500k documents at once, split the work into chunks. This keeps memory usage manageable while still avoiding slow row-by-row appends.
import pandas as pd chunk_size = 10000 # Adjust based on your available RAM (10k is a safe starting point) total_docs = len(ListOfDocuments) filename_prefix = filename.split(".")[0] c_name_column = AllParameters["C_Name"] target_columns = list(AllParameters.values()) + [c_name_column] # Initialize empty DataFrame with correct columns AppendingDataframe = pd.DataFrame(columns=target_columns) for i in range(0, total_docs, chunk_size): # Grab a chunk of documents current_chunk = ListOfDocuments[i:i+chunk_size] # Process the chunk chunk_data = [ { **{col: doc.get(key, pd.NA) for key, col in AllParameters.items() if key != "USELESS"}, c_name_column: filename_prefix } for doc in current_chunk ] # Append the chunk to the DataFrame using pd.concat (way faster than .loc) AppendingDataframe = pd.concat([AppendingDataframe, pd.DataFrame(chunk_data)], ignore_index=True)
Key Takeaways
- Avoid row-by-row DataFrame modifications at all costs—pandas is designed for batch operations.
- Only iterate the keys you need instead of every key in each document.
- Handle missing values explicitly with
doc.get(key, pd.NA)to avoid data errors.
内容的提问来源于stack exchange,提问作者Jackdaw

