You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中DataFrame与Series条件筛选的代码优化求助

优化Pandas循环判断代码的方案

需求回顾

  • 先把Ser1(id列)和Ser2(section列)合并成单个DataFrame(记为df1)
  • 对比df1和df2(含id、section列),生成result列:
    • 若df1的id不在df2的id中,result为False
    • 若id在df2中,但对应section不在该id的df2 section列表里,result为True
    • 其余情况(id和section都匹配df2),result为False

现有循环的问题

循环里每次都要切片df2并转成列表,数据量大时重复操作会导致效率极低,而且代码冗余。

优化方案

方法1:字典映射+列表推导(简洁高效)

先把df2的id和对应section集合做映射,再用列表推导生成结果:

import pandas as pd

# 1. 合并Ser1和Ser2为df1
df1 = pd.DataFrame({'id': Ser1, 'section': Ser2})

# 2. 给df2建立id到section集合的映射(集合查找比列表快N倍)
id_section_map = df2.groupby('id')['section'].agg(set).to_dict()

# 3. 生成result列
df1['result'] = [
    (id_val in id_section_map) and (section_val not in id_section_map[id_val])
    for id_val, section_val in zip(df1['id'], df1['section'])
]

方法2:全Pandas向量化操作(适合大型数据集)

用Pandas的分组、合并和向量化判断,避免Python层面的循环:

import pandas as pd

# 1. 合并Ser1和Ser2为df1
df1 = pd.DataFrame({'id': Ser1, 'section': Ser2})

# 2. 标记df1的id是否存在于df2中
df1['id_in_df2'] = df1['id'].isin(df2['id'].unique())

# 3. 给df2按id分组,生成每个id对应的section集合
df2_grouped = df2.groupby('id')['section'].agg(set).reset_index(name='df2_sections')

# 4. 合并df1和分组后的df2,保留每个id对应的section集合
df1_merged = df1.merge(df2_grouped, on='id', how='left')

# 5. 计算result:id存在且当前section不在对应集合里
df1_merged['result'] = df1_merged['id_in_df2'] & (
    df1_merged.apply(
        lambda x: x['section'] not in x['df2_sections'] if pd.notna(x['df2_sections']) else False,
        axis=1
    )
)

# 6. 整理成预期的结果格式
final_df = df1_merged[['id', 'section', 'result']]

验证结果

用你给出的示例数据测试,两种方法都会得到如下结果:

idsectionresult
1AFalse
2BFalse
2CTrue
3DFalse

内容的提问来源于stack exchange,提问作者Nebiros

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 23:10:43