如何高效实现含NaN的DataFrame中Identifier对应Name的全日期填充?
Great question! The loop approach you’re using works but is indeed inefficient for large datasets—repeated appends and copies are expensive operations in pandas. Instead, we can use vectorized operations like merge or groupby + explode to achieve the same result much faster.
Method 1: Using merge (Most Efficient)
This approach leverages a cross-join between your original data (without the Name column) and a lookup table of unique Identifier-Name pairs. Here’s how it works:
Create a lookup table of non-null, unique
Identifier-Namepairs:name_map = df.dropna(subset=['Name'])[['Identifier', 'Name']].drop_duplicates()This captures all valid Name values associated with each Identifier.
Merge with the original data (excluding the original
Namecolumn) to pair every row with each valid Name for its Identifier:final_df = df.drop(columns=['Name']).merge(name_map, on='Identifier', how='inner').drop_duplicates()
Testing this with your sample data:
name_mapwill have two rows:(1, 'CD')and(1, 'AD')- Merging with the original 6 rows (without Name) gives 12 rows total
drop_duplicates()removes the duplicate rows from the original 2019-05-07 entries, resulting in exactly your expected output.
Method 2: Using groupby + explode
Another clean option is to group by Identifier, attach the list of valid Names to each row in the group, then explode the list into separate rows:
final_df = ( df.groupby('Identifier') .apply(lambda grp: grp.assign(Name=grp['Name'].dropna().unique().tolist())) .reset_index(drop=True) .explode('Name') .drop_duplicates() )
We first assign all valid Names for the group to each row as a list, then explode the list to create a separate row for each Name. drop_duplicates() cleans up any redundant rows from the original data.
Why These Are Faster
Both methods avoid explicit loops and use pandas’ optimized internal operations. The merge approach is particularly efficient because it’s fully vectorized, making it ideal for very large datasets where performance is critical.
Content of the question originates from Stack Exchange, asked by Menno Van Dijk.

