R语言:匹配两个数据框的字符串列并生成匹配标记列
匹配两个DataFrame的组合列并生成标记列
问题说明
给定两个DataFrame:
- df1包含
Object和Thing列 - df2包含
Object1和Thing1列
需要检查df1中每一行的Object+Thing组合是否存在于df2的Object1+Thing1组合中,匹配则新增Value列赋值为1,不匹配则赋值为0。实际场景中列名可能不对应,需灵活处理。
示例数据构造
import pandas as pd # 构造df1 df1 = pd.DataFrame({ "Object": ["apple tini", "vodka cran", "tom collins", "arnie palmer"], "Thing": ["drink", "beverage", "alcohol", "cocktail"] }) # 构造df2 df2 = pd.DataFrame({ "Object1": ["apple tini", "vodka cran", "tom collins", "arnie palmer"], "Thing1": ["drink", "bever", "alc", "cocktail"] })
解决方案
方法一:使用Merge左连接标记
适合可以通过重命名对齐列名的场景:
# 重命名df2的匹配列,与df1对应列名一致 df2_aligned = df2.rename(columns={"Object1": "Object", "Thing1": "Thing"}) # 左连接,仅保留匹配标记列 df1 = df1.merge(df2_aligned[["Object", "Thing"]], on=["Object", "Thing"], how="left", indicator=True) # 生成Value列:匹配为1,不匹配为0 df1["Value"] = df1["_merge"].map({"both": 1, "left_only": 0}) # 删除临时标记列 df1.drop("_merge", axis=1, inplace=True)
方法二:使用元组集合匹配(更灵活)
无需修改列名,直接指定对应匹配列即可,适配任意列名场景:
# 提取df2的目标列组合,转为元组集合(查询效率高) df2_combos = set(zip(df2["Object1"], df2["Thing1"])) # 逐行检查df1的组合是否在集合中,转为整数类型 df1["Value"] = df1.apply(lambda row: (row["Object"], row["Thing"]) in df2_combos, axis=1).astype(int)
预期结果
| Object | Thing | Value |
|---|---|---|
| apple tini | drink | 1 |
| vodka cran | beverage | 0 |
| tom collins | alcohol | 0 |
| arnie palmer | cocktail | 1 |
内容的提问来源于stack exchange,提问作者Jacob
相关产品推荐
相关产品推荐

