合并含空DataFrame的列表:保留行数并转为NA值的技术需求
Got it, let's fix this so you end up with exactly 5 rows in your final DataFrame—one for each entry in your original list, even the empty ones. The key is to replace those empty DataFrames with single-row DataFrames filled with NA values that match the column structure of your non-empty DataFrames.
Step-by-Step Approach
- Identify the column structure from your non-empty DataFrames (we'll assume all non-empty ones have the same columns, which matches your example).
- Loop through your DataFrame list: For each entry, if it's empty, create a new single-row DataFrame with
NAvalues for every column. If it's non-empty, keep it as-is. - Concatenate the processed list: Now when you concatenate, all 5 entries will contribute a row to the final result.
Example Code
First, let's replicate your DataFrame list to test with:
import pandas as pd # Simulate your original DataFrame list df1 = pd.DataFrame({ 0: ['102,000,000.00'], 1: ['2,000,000.00'], 2: ['1,400,000.00'], 3: ['0.00'] }) df2 = pd.DataFrame() # Empty df3 = pd.DataFrame() # Empty df4 = pd.DataFrame() # Empty df5 = pd.DataFrame({ 0: ['60,900,000.00'], 1: ['1,300,000.00'], 2: ['0.00'], 3: ['0.00'] }) dataframes = [df1, df2, df3, df4, df5]
Now process and concatenate:
# Get the column names from the first non-empty DataFrame non_empty_columns = next(df.columns for df in dataframes if not df.empty) # Process each DataFrame in the list processed_dataframes = [] for df in dataframes: if df.empty: # Create a single-row DataFrame filled with NA, matching the column structure na_row_df = pd.DataFrame([[pd.NA] * len(non_empty_columns)], columns=non_empty_columns) processed_dataframes.append(na_row_df) else: processed_dataframes.append(df) # Concatenate the processed list final_data = pd.concat(processed_dataframes, ignore_index=True) print(final_data)
Expected Output
0 1 2 3 0 102,000,000.00 2,000,000.00 1,400,000.00 0.00 1 <NA> <NA> <NA> <NA> 2 <NA> <NA> <NA> <NA> 3 <NA> <NA> <NA> <NA> 4 60,900,000.00 1,300,000.00 0.00 0.00
Why This Works
Your original pd.concat(dataframes) was dropping empty DataFrames entirely, which is the default behavior. By replacing each empty DataFrame with a valid (but NA-filled) single-row DataFrame, we ensure every entry in your list contributes exactly one row to the final result.
内容的提问来源于stack exchange,提问作者Hope

