You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效实现含NaN的DataFrame中Identifier对应Name的全日期填充?

Great question! The loop approach you’re using works but is indeed inefficient for large datasets—repeated appends and copies are expensive operations in pandas. Instead, we can use vectorized operations like merge or groupby + explode to achieve the same result much faster.

Method 1: Using merge (Most Efficient)

This approach leverages a cross-join between your original data (without the Name column) and a lookup table of unique Identifier-Name pairs. Here’s how it works:

  1. Create a lookup table of non-null, unique Identifier-Name pairs:

    name_map = df.dropna(subset=['Name'])[['Identifier', 'Name']].drop_duplicates()
    

    This captures all valid Name values associated with each Identifier.

  2. Merge with the original data (excluding the original Name column) to pair every row with each valid Name for its Identifier:

    final_df = df.drop(columns=['Name']).merge(name_map, on='Identifier', how='inner').drop_duplicates()
    

Testing this with your sample data:

  • name_map will have two rows: (1, 'CD') and (1, 'AD')
  • Merging with the original 6 rows (without Name) gives 12 rows total
  • drop_duplicates() removes the duplicate rows from the original 2019-05-07 entries, resulting in exactly your expected output.

Method 2: Using groupby + explode

Another clean option is to group by Identifier, attach the list of valid Names to each row in the group, then explode the list into separate rows:

final_df = (
    df.groupby('Identifier')
      .apply(lambda grp: grp.assign(Name=grp['Name'].dropna().unique().tolist()))
      .reset_index(drop=True)
      .explode('Name')
      .drop_duplicates()
)

We first assign all valid Names for the group to each row as a list, then explode the list to create a separate row for each Name. drop_duplicates() cleans up any redundant rows from the original data.

Why These Are Faster

Both methods avoid explicit loops and use pandas’ optimized internal operations. The merge approach is particularly efficient because it’s fully vectorized, making it ideal for very large datasets where performance is critical.

Content of the question originates from Stack Exchange, asked by Menno Van Dijk.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:23:52