如何使用Pandas查找File1列中指定字符并提取File2匹配行
Pandas 跨表匹配筛选实现方案
首先导入依赖并读取文件:
import pandas as pd # 可根据实际文件格式替换为 read_excel、read_table 等方法 df1 = pd.read_csv("File1.csv") df2 = pd.read_csv("File2.csv")
步骤1:从File1中提取匹配指定字符的条目
# 请替换下列参数: # col_name:File1中需要检索的列名 # target_str:你指定的需要匹配的字符 match_keywords = df1[df1["col_name"].str.contains("target_str", na=False)]["col_name"].unique().tolist()
说明:na=False参数会自动跳过空值避免匹配报错,unique()用于对匹配结果去重,减少后续计算冗余。
步骤2:用提取到的关键词筛选File2的行
根据你的需求可选两种匹配模式:
模式1:仅在File2的指定列中匹配
# 替换 target_col 为File2中需要匹配的列名 # 完全匹配写法 filter_result = df2[df2["target_col"].isin(match_keywords)] # 模糊匹配写法(只要包含关键词就算匹配) filter_result = df2[df2["target_col"].str.contains("|".join(match_keywords), na=False)]
模式2:在File2的所有列中全局匹配,只要任意一列包含关键词就保留该行
match_rule = "|".join(match_keywords) mask = df2.apply(lambda x: x.astype(str).str.contains(match_rule, na=False)).any(axis=1) filter_result = df2[mask]
最后可以将结果导出:
filter_result.to_csv("匹配结果.csv", index=False)
如果需要开启大小写敏感匹配,给str.contains方法添加case=True参数即可。
内容的提问来源于stack exchange,提问作者LoganLee
相关产品推荐
相关产品推荐

