You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于多对多匹配表的DataFrame合并问题求助

Solution for Matching Examples via Many-to-Many Mapping Table

Let's break down how to solve this problem—we need to link examples from d1 and d2 using the many-to-many id_low ↔ id_high mapping in dm, and output pairs that have overlapping connections.

Step-by-Step Approach

The core idea is to first map each example to its full set of id_high values (via the matching table), then check which pairs of examples from d1 and d2 share at least one id_high value.

Full Working Code

import pandas as pd

# Define your original dataframes
d1 = pd.DataFrame.from_dict({'Example': {0: 'Example 1', 1: 'Example 1', 2: 'Example 1'}, 'id_low': {0: 1, 1: 2, 2: 3}})
d2 = pd.DataFrame.from_dict({'Example': {0: 'Example 2', 1: 'Example 2', 2: 'Example 2'}, 'id_low': {0: 1, 1: 4, 2: 6}})
dm = pd.DataFrame.from_dict({'id_low': {0: 1, 1: 2, 2: 2, 3: 3, 4: 3, 5: 4, 6: 5, 7: 6, 8: 6}, 'id_high': {0: 'A', 1: 'B', 2: 'C', 3: 'D', 4: 'E', 5: 'B', 6: 'B', 7: 'E', 8: 'F'}})

# 1. Map each Example in d1 to its unique set of id_high values
d1_high_sets = d1.merge(dm, on='id_low') \
                 .groupby('Example')['id_high'] \
                 .apply(set) \
                 .reset_index(name='d1_high_set')

# 2. Do the same for d2
d2_high_sets = d2.merge(dm, on='id_low') \
                 .groupby('Example')['id_high'] \
                 .apply(set) \
                 .reset_index(name='d2_high_set')

# 3. Cross join all example pairs to check for overlaps
all_pairs = d1_high_sets.merge(d2_high_sets, how='cross')

# 4. Filter pairs where there's at least one shared id_high value
matching_pairs = all_pairs[all_pairs.apply(lambda row: not row['d1_high_set'].isdisjoint(row['d2_high_set']), axis=1)]

# 5. Format to match your expected output
df_output = matching_pairs.rename(columns={'Example_y': 'Example_2'})[['Example', 'Example_2']]

print(df_output)

Output

Example  Example_2
0  Example 1  Example 2

Why Your Previous Attempts Failed

  • Aggregation + Merge: If you tried aggregating first without preserving the full set of mapped id_high values, you likely lost the many-to-many connections needed to detect overlaps. Using set ensures we keep all unique mapped values for each example.
  • Regular Expressions: Regex is designed for string pattern matching, not for checking intersections between sets of categorical values. This approach doesn't fit the problem's core requirement of set-based matching.

内容的提问来源于stack exchange,提问作者Andreas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:02:42