判断Pandas DataFrame列值是否存在于另一DataFrame列的问题
解决Pandas判断动物名是否存在于文本列的问题
问题根源
你使用的isin()方法是做精确匹配——它只会判断动物名是否完全等于df_content.document_content中的某一条完整字符串,而非检查动物名是否作为子串出现在任意文本内容里。比如"dog"只有当某条document_content恰好是"dog"时才会返回True,像"the dog was hungry"这种包含"dog"的字符串不会被识别,这就是明明存在却显示False的原因。
正确实现方法
先构造测试数据方便你复现验证:
import pandas as pd # 测试用DataFrame df_animals = pd.DataFrame({'animal': ['cat', 'dog', 'bird', 'rat', 'fish']}) df_content = pd.DataFrame({'document_content': [ 'cat in the hat', 'the dog was hungry', 'a bird flew by', 'fish are swimming' ]})
方法一:高效正则提取法(适合大数据量)
先从所有文本中提取出出现过的动物名,再用isin()匹配:
# 生成匹配完整动物单词的正则模式(\b用于匹配单词边界,避免误匹配类似"rat"和"rate"的情况) animal_pattern = r'\b(' + '|'.join(df_animals['animal']) + r')\b' # 从所有文本中提取出现过的动物并去重 found_animals = df_content['document_content'].str.extractall(animal_pattern)[0].unique() # 新增标记列 df_animals['in_document_content'] = df_animals['animal'].isin(found_animals)
方法二:逐动物检查法(直观易懂)
对每个动物,检查是否存在于任意一条文本内容中:
# 逐行检查动物名,any()表示只要有一条文本包含该动物就返回True df_animals['in_document_content'] = df_animals['animal'].apply( lambda animal: df_content['document_content'].str.contains(r'\b' + animal + r'\b').any() )
补充说明
- 正则中的
\b是为了匹配完整单词,如果不需要严格匹配完整单词(比如允许匹配"cat"在"category"中),可以去掉\b。 - 两种方法最终都会得到正确结果:cat、dog、bird、fish标记为True,rat标记为False。
内容的提问来源于stack exchange,提问作者glucosebat
相关产品推荐
相关产品推荐

