You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Azure Databricks中用PySpark比较DataFrame,能否用for循环实现?

可以用for循环实现该需求

当然能用for loop完成这个任务,以下是具体的Python实现代码(基于pandas):

import pandas as pd

# 初始化第一个DataFrame
df1 = pd.DataFrame({
    'id': ['01', '02', '03', '04', '05'],
    'inUse': [True, True, False, True, False],
    'otherId': ['2304', '5654', '6785', None, '6785']
})

# 初始化第二个DataFrame
df2 = pd.DataFrame({
    'id': ['01', '02', '04', '05'],
    'otherId': ['2304', '5654', '37584', '6785']
})

# 从df1中提取有效的(id, otherId)组合(排除otherId为空的记录)
valid_pairs = set(df1.dropna(subset=['otherId']).apply(lambda row: (row['id'], row['otherId']), axis=1))

# 用for循环遍历df2,生成Presentin col的值
present_status = []
for _, row in df2.iterrows():
    current_pair = (row['id'], row['otherId'])
    present_status.append(current_pair in valid_pairs)

# 将结果列添加到df2
df2['Presentin col'] = present_status

# 查看最终结果
print(df2)

逻辑说明

  1. 先从第一个DataFrame中过滤掉otherId为空的记录,把剩下的id和otherId组成元组存入集合——集合的查找效率远高于列表,能提升循环速度。
  2. 遍历第二个DataFrame的每一行,把当前行的id和otherId组成元组,检查是否在第一步生成的有效集合中,将结果存入列表。
  3. 把这个结果列表作为新列添加到第二个DataFrame,就得到你想要的输出结构。

补充:虽然for loop能实现,但pandas本身更推荐向量化操作(比如用merge或isin),不过既然你明确问for loop的可行性,上面的代码完全满足需求。

内容的提问来源于stack exchange,提问作者Sanjay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 18:52:20