如何基于列值执行查找并填充DataFrame的ParentCategoryId空列?
解决方案:用Pandas矢量化操作高效填充ParentCategoryId
核心思路是先建立SourceCategoryId到CategoryId的映射关系,再通过这个映射快速匹配SourceParentCategoryId对应的CategoryId,避免低效的逐行遍历。
具体步骤:
- 构建映射结构:从现有DataFrame中提取
SourceCategoryId作为键,CategoryId作为值,形成可快速查找的映射。 - 筛选目标行:仅处理
SourceParentCategoryId不为0的行,符合你提出的跳过逻辑。 - 批量填充:用映射关系批量替换
SourceParentCategoryId为对应的CategoryId,赋值给ParentCategoryId列。
代码实现
import pandas as pd # 构造样本数据(实际使用时直接读取你的数据源即可) data = { 'CategoryId': [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20], 'ParentCategoryId': [None, None, 9.0, 20.0, 4.0, None, None, None, None, None, None, None, None, None, None, None, None, None, None, None], 'SourceCategoryId': [100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,100], 'SourceParentCategoryId': [0,0,108,100,103,103,103,103,0,108,103,103,103,100,113,113,113,113,113,113] } df = pd.DataFrame(data) # 建立SourceCategoryId -> CategoryId的映射(Series更适配Pandas操作) source_category_map = df.set_index('SourceCategoryId')['CategoryId'] # 筛选需要填充的行并批量赋值 fill_mask = df['SourceParentCategoryId'] != 0 df.loc[fill_mask, 'ParentCategoryId'] = df.loc[fill_mask, 'SourceParentCategoryId'].map(source_category_map) # 查看结果 print(df)
关键说明:
- 效率优势:
map的矢量化操作比iterrows()或apply()逐行遍历快几个数量级,尤其适合大数据量场景。 - 缺失值处理:如果
SourceParentCategoryId的值在SourceCategoryId中不存在,map会返回NaN,可通过fillna()保留原有值或填充默认值,示例:df.loc[fill_mask, 'ParentCategoryId'] = df.loc[fill_mask, 'SourceParentCategoryId'].map(source_category_map).fillna(df['ParentCategoryId']) - 覆盖逻辑:若需要完全按规则覆盖现有
ParentCategoryId值,直接使用示例中的赋值语句即可。
内容的提问来源于stack exchange,提问作者Slavisha84
相关产品推荐
相关产品推荐

