基于多对多匹配表的DataFrame合并问题求助
Solution for Matching Examples via Many-to-Many Mapping Table
Let's break down how to solve this problem—we need to link examples from d1 and d2 using the many-to-many id_low ↔ id_high mapping in dm, and output pairs that have overlapping connections.
Step-by-Step Approach
The core idea is to first map each example to its full set of id_high values (via the matching table), then check which pairs of examples from d1 and d2 share at least one id_high value.
Full Working Code
import pandas as pd # Define your original dataframes d1 = pd.DataFrame.from_dict({'Example': {0: 'Example 1', 1: 'Example 1', 2: 'Example 1'}, 'id_low': {0: 1, 1: 2, 2: 3}}) d2 = pd.DataFrame.from_dict({'Example': {0: 'Example 2', 1: 'Example 2', 2: 'Example 2'}, 'id_low': {0: 1, 1: 4, 2: 6}}) dm = pd.DataFrame.from_dict({'id_low': {0: 1, 1: 2, 2: 2, 3: 3, 4: 3, 5: 4, 6: 5, 7: 6, 8: 6}, 'id_high': {0: 'A', 1: 'B', 2: 'C', 3: 'D', 4: 'E', 5: 'B', 6: 'B', 7: 'E', 8: 'F'}}) # 1. Map each Example in d1 to its unique set of id_high values d1_high_sets = d1.merge(dm, on='id_low') \ .groupby('Example')['id_high'] \ .apply(set) \ .reset_index(name='d1_high_set') # 2. Do the same for d2 d2_high_sets = d2.merge(dm, on='id_low') \ .groupby('Example')['id_high'] \ .apply(set) \ .reset_index(name='d2_high_set') # 3. Cross join all example pairs to check for overlaps all_pairs = d1_high_sets.merge(d2_high_sets, how='cross') # 4. Filter pairs where there's at least one shared id_high value matching_pairs = all_pairs[all_pairs.apply(lambda row: not row['d1_high_set'].isdisjoint(row['d2_high_set']), axis=1)] # 5. Format to match your expected output df_output = matching_pairs.rename(columns={'Example_y': 'Example_2'})[['Example', 'Example_2']] print(df_output)
Output
Example Example_2 0 Example 1 Example 2
Why Your Previous Attempts Failed
- Aggregation + Merge: If you tried aggregating first without preserving the full set of mapped
id_highvalues, you likely lost the many-to-many connections needed to detect overlaps. Usingsetensures we keep all unique mapped values for each example. - Regular Expressions: Regex is designed for string pattern matching, not for checking intersections between sets of categorical values. This approach doesn't fit the problem's core requirement of set-based matching.
内容的提问来源于stack exchange,提问作者Andreas
相关产品推荐
相关产品推荐

