如何用Python循环/列表推导式结合正则为DataFrame新增布尔/分类列
解决Pandas中基于多关键词匹配生成布尔/分类列的问题
问题背景
我有一个包含多列的Pandas DataFrame,其中一列是由空格分隔、包含大小写词汇的房产标题字符串。需要基于指定的搜索词列表,通过Python的re模块结合循环/列表推导式遍历该列,生成布尔值(True/False)列或匹配到的分类名称列,并添加到DataFrame中。
示例DataFrame
import pandas as pd data = {'id': [748, 896, 5268], 'name' : ['Bright, Modern Garden Unit - 1BR/1BTH', 'Renovated Alamo Square Victorian', 'Mission Sunny, near Park'], 'price': [209, 255, 180]} df = pd.DataFrame(data)
输出:
id name price 0 748 Bright, Modern Garden Unit - 1BR/1BTH 209 1 896 Renovated Alamo Square Victorian 255 2 5268 Mission Sunny, near Park 180
遇到的问题
单独搜索单个关键词可以实现,但尝试直接将搜索词列表传入str.contains时,会触发unhashable type: list错误,不清楚如何正确整合正则语法实现需求。
解决方案
1. 生成布尔匹配列(amenities_bool)
核心是将搜索词列表转换为正则OR模式,通过re.compile编译并忽略大小写,再用str.contains完成批量匹配:
import re # 定义目标搜索词 search_terms = ['bright', 'renovated', 'near'] # 编译正则模式:用|连接所有关键词,忽略大小写 pattern = re.compile('|'.join(search_terms), re.IGNORECASE) # 生成布尔列:匹配到任意关键词则为True df['amenities_bool'] = df['name'].str.contains(pattern) print(df)
输出结果:
id name price amenities_bool 0 748 Bright, Modern Garden Unit - 1BR/1BTH 209 True 1 896 Renovated Alamo Square Victorian 255 True 2 5268 Mission Sunny, near Park 180 True
说明:str.contains不接受列表作为参数,必须将列表转为单个正则表达式字符串(用|表示逻辑或),编译时添加re.IGNORECASE可自动匹配大小写不同的词汇。
2. 生成分类描述列(amenities_descp)
如果需要提取具体匹配到的关键词,用列表推导式结合re.search提取第一个匹配的词汇(转小写):
# 生成分类列:提取第一个匹配的关键词并转为小写 df['amenities_descp'] = [ re.search(pattern, text).group().lower() if re.search(pattern, text) else None for text in df['name'] ] print(df)
输出结果:
id name price amenities_bool amenities_descp 0 748 Bright, Modern Garden Unit - 1BR/1BTH 209 True bright 1 896 Renovated Alamo Square Victorian 255 True renovated 2 5268 Mission Sunny, near Park 180 True near
扩展:如果需要提取所有匹配的关键词(用逗号分隔),可以改用re.finditer:
df['amenities_descp'] = [ ', '.join([match.group().lower() for match in re.finditer(pattern, text)]) if re.search(pattern, text) else None for text in df['name'] ]
内容的提问来源于stack exchange,提问作者new_to_code
相关产品推荐
相关产品推荐

