使用str.contains过滤Pandas字符串列的互斥分类匹配问题
问题分析与解决方案
问题描述
现有如下DataFrame,long_category列存储商家的长分类:
import pandas as pd import numpy as np df = pd.DataFrame({ 'long_category': {0: 'Doctors, Traditional Chinese Medicine, Naturopathic/Holistic, Acupuncture, Health & Medical, Nutritionists', 1: 'Shipping Centers, Local Services, Notaries, Mailbox Centers, Printing Services', 2: 'Department Stores, Shopping, Fashion, Home & Garden, Electronics, Furniture Stores', 3: 'Restaurants, Food, Bubble Tea, Coffee & Tea, Bakeries', 4: 'Brewpubs, Breweries, Food', 5: 'Burgers, Fast Food, Sandwiches, Food, Ice Cream & Frozen Yogurt, Restaurants', 6: 'Sporting Goods, Fashion, Shoe Stores, Shopping, Sports Wear, Accessories', 7: 'Synagogues, Religious Organizations', 8: 'Pubs, Restaurants, Italian, Bars, American (Traditional), Nightlife, Greek', 9: 'Ice Cream & Frozen Yogurt, Fast Food, Burgers, Restaurants, Food'}})
目标是将长分类映射为短分类,匹配规则基于known_categories中的关键词,要求短分类互斥,且可通过调整known_categories的顺序设置优先级:
known_categories = ['restaurant', 'beauty & spas', 'hotels', 'health & medical', 'shopping', 'coffee & tea','automotive', 'pets|veterinian', 'services', 'stores', 'grocery', 'ice cream']
原代码未正常工作:第9行长分类包含restaurant但未匹配到任何短分类,且无法实现优先级匹配。原代码如下:
df['short_category'] = np.nan for cat in known_categories: excluded_cats = [x for x in known_categories if x!= cat] df['short_category'] [ ~(df.long_category.str.contains('|'.join(excluded_cats), regex = True, case = False, na = False)) & (df.long_category.str.contains(cat, regex = True, case = False, na = False))] = cat
问题根源
原代码的逻辑是仅当当前行完全不包含其他任何已知分类,同时包含当前分类时,才会赋值。这种逻辑会导致:
- 对于包含多个匹配关键词的行(比如第9行同时包含
restaurant和ice cream),循环到restaurant时,因为行内包含ice cream(属于排除列表),所以~(包含排除分类)的条件不成立,无法赋值; - 循环到
ice cream时,行内又包含restaurant(属于排除列表),同样无法满足条件,最终该行的short_category保持为NaN; - 完全不符合优先级匹配的需求,因为它要求行只能匹配一个分类,且不能通过顺序控制优先匹配哪一个。
解决方案:优先级匹配逻辑
正确的逻辑应该是按known_categories的顺序遍历,对尚未匹配分类的行,只要包含当前关键词就赋值,赋值后不再参与后续匹配,这样既保证互斥,又能通过列表顺序控制优先级。
修正后的代码:
df['short_category'] = np.nan for cat in known_categories: # 只对还未匹配分类的行进行判断 mask = df['short_category'].isna() & df['long_category'].str.contains(cat, regex=True, case=False, na=False) df.loc[mask, 'short_category'] = cat
效果验证
- 当前
known_categories顺序下,第9行会优先匹配restaurant(因为它在ice cream前面); - 如果将
ice cream移到restaurant之前,第9行会匹配ice cream; - 第3行同时包含
restaurant和coffee & tea,会优先匹配restaurant; - 第8行只包含
restaurant,正常匹配该分类。
优化建议(可选)
如果需要避免部分单词匹配(比如避免把restaur误识别为restaurant),可以给每个关键词加上正则单词边界\b,修改known_categories为:
known_categories = [r'\brestaurant\b', r'\bbeauty & spas\b', r'\bhotels\b', r'\bhealth & medical\b', r'\bshopping\b', r'\bcoffee & tea\b',r'\bautomotive\b', r'\b(pets|veterinian)\b', r'\bservices\b', r'\bstores\b', r'\bgrocery\b', r'\bice cream\b']
这样只会匹配完整的单词,提高分类准确性。
内容的提问来源于stack exchange,提问作者Saeed
相关产品推荐
相关产品推荐

