You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测df2的Origin-Destination组合是否在df1列表列中并拆分行?

解决方案

要实现需求,我们可以通过笛卡尔积匹配+条件过滤的方式高效完成,避免for循环的低效问题,具体步骤和代码如下:

import pandas as pd

# 构造示例数据(实际使用时替换为你的真实数据)
df1 = pd.DataFrame({
    'A': [['a', 'b', 'c'], ['a', 'd', 'e'], ['b', 'c', 'e']],
    'time': [122, 45, 64]
})
df2 = pd.DataFrame({
    'Origin': ['a', 'b', 'b', 'd'],
    'Destination': ['b', 'c', 'e', 'e']
})

# 1. 为df1添加集合列,提升元素查找效率(集合查找是O(1),比列表O(n)快)
df1['A_set'] = df1['A'].apply(set)
# 保留原索引,用于最终结果的索引对齐
df1 = df1.reset_index().rename(columns={'index': 'original_idx'})

# 2. 生成df1和df2的笛卡尔积,得到所有可能的组合
cross_merge = pd.merge(df1, df2, how='cross')

# 3. 过滤出符合条件的行:Origin和Destination都在对应A列的集合中
filtered = cross_merge[cross_merge.apply(lambda row: row['Origin'] in row['A_set'] and row['Destination'] in row['A_set'], axis=1)]

# 4. 整理结果格式,恢复原索引并保留需要的列
result = filtered.set_index('original_idx')[['A', 'time', 'Origin', 'Destination']]

print(result)

输出结果

运行后会得到与需求完全一致的DataFrame:

A  time Origin Destination
original_idx                            
0        [a, b, c]   122      a            b
0        [a, b, c]   122      b            c
1        [a, d, e]    45      d            e
2        [b, c, e]    64      b            c
2        [b, c, e]    64      b            e

补充说明

  • 该方法比for循环高效得多,尤其适合处理大型DataFrame;
  • 转成集合的操作是为了提升元素存在性检查的速度,如果你的A列列表元素较少,也可以直接用原列表进行判断(把row['A_set']换成row['A']即可)。

内容的提问来源于stack exchange,提问作者Joost van Maanen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 01:35:25