You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用str.contains过滤Pandas字符串列的互斥分类匹配问题

问题分析与解决方案

问题描述

现有如下DataFrame,long_category列存储商家的长分类:

import pandas as pd
import numpy as np

df = pd.DataFrame({
 'long_category': {0: 'Doctors, Traditional Chinese Medicine, Naturopathic/Holistic, Acupuncture, Health & Medical, Nutritionists',
  1: 'Shipping Centers, Local Services, Notaries, Mailbox Centers, Printing Services',
  2: 'Department Stores, Shopping, Fashion, Home & Garden, Electronics, Furniture Stores',
  3: 'Restaurants, Food, Bubble Tea, Coffee & Tea, Bakeries',
  4: 'Brewpubs, Breweries, Food',
  5: 'Burgers, Fast Food, Sandwiches, Food, Ice Cream & Frozen Yogurt, Restaurants',
  6: 'Sporting Goods, Fashion, Shoe Stores, Shopping, Sports Wear, Accessories',
  7: 'Synagogues, Religious Organizations',
  8: 'Pubs, Restaurants, Italian, Bars, American (Traditional), Nightlife, Greek',
  9: 'Ice Cream & Frozen Yogurt, Fast Food, Burgers, Restaurants, Food'}})

目标是将长分类映射为短分类,匹配规则基于known_categories中的关键词,要求短分类互斥,且可通过调整known_categories的顺序设置优先级:

known_categories = ['restaurant', 'beauty & spas', 'hotels', 'health & medical', 'shopping', 'coffee & tea','automotive', 'pets|veterinian', 'services', 'stores',  'grocery', 'ice cream']

原代码未正常工作:第9行长分类包含restaurant但未匹配到任何短分类,且无法实现优先级匹配。原代码如下:

df['short_category'] = np.nan
for cat in known_categories:
    excluded_cats = [x for x in known_categories if x!= cat]
    df['short_category'] [ ~(df.long_category.str.contains('|'.join(excluded_cats), regex = True, case = False, na = False)) & (df.long_category.str.contains(cat, regex = True, case = False, na = False))] = cat

问题根源

原代码的逻辑是仅当当前行完全不包含其他任何已知分类,同时包含当前分类时,才会赋值。这种逻辑会导致:

  • 对于包含多个匹配关键词的行(比如第9行同时包含restaurant和ice cream),循环到restaurant时,因为行内包含ice cream(属于排除列表),所以~(包含排除分类)的条件不成立,无法赋值;
  • 循环到ice cream时,行内又包含restaurant(属于排除列表),同样无法满足条件,最终该行的short_category保持为NaN;
  • 完全不符合优先级匹配的需求,因为它要求行只能匹配一个分类,且不能通过顺序控制优先匹配哪一个。

解决方案:优先级匹配逻辑

正确的逻辑应该是按known_categories的顺序遍历,对尚未匹配分类的行,只要包含当前关键词就赋值,赋值后不再参与后续匹配,这样既保证互斥,又能通过列表顺序控制优先级。

修正后的代码:

df['short_category'] = np.nan

for cat in known_categories:
    # 只对还未匹配分类的行进行判断
    mask = df['short_category'].isna() & df['long_category'].str.contains(cat, regex=True, case=False, na=False)
    df.loc[mask, 'short_category'] = cat

效果验证

  • 当前known_categories顺序下,第9行会优先匹配restaurant(因为它在ice cream前面);
  • 如果将ice cream移到restaurant之前,第9行会匹配ice cream;
  • 第3行同时包含restaurant和coffee & tea,会优先匹配restaurant;
  • 第8行只包含restaurant,正常匹配该分类。

优化建议(可选)

如果需要避免部分单词匹配(比如避免把restaur误识别为restaurant),可以给每个关键词加上正则单词边界\b,修改known_categories为:

known_categories = [r'\brestaurant\b', r'\bbeauty & spas\b', r'\bhotels\b', r'\bhealth & medical\b', r'\bshopping\b', r'\bcoffee & tea\b',r'\bautomotive\b', r'\b(pets|veterinian)\b', r'\bservices\b', r'\bstores\b',  r'\bgrocery\b', r'\bice cream\b']

这样只会匹配完整的单词,提高分类准确性。


内容的提问来源于stack exchange,提问作者Saeed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 17:31:00