如何用Pandas从DataFrame列提取关键词生成新数据框
问题:用Pandas简洁实现从Type列提取关键词分类
嘿,我来帮你搞定这个Pandas数据清洗的问题!你之前用循环的方式确实有点繁琐,咱们用Pandas的向量化操作就能轻松搞定,而且代码更简洁高效。先理清楚你的需求和现有数据:
背景数据
现有Pandas DataFrame df 数据如下:
| id | Type | agent_id | created_at |
|---|---|---|---|
| 0 | 44525 Stunning 6 bedroom villa in New Delhi | 184 | 2018-03-09 |
| 1 | 44859 Villa for sale in Amritsar | 182 | 2017-02-19 |
| 2 | 45465 House in Faridabad | 154 | 2017-04-17 |
| 3 | 50685 5 Hectre land near New Delhi | 113 | 2017-09-01 |
| 4 | 130728 Duplex in Mumbai | 157 | 2017-02-07 |
| 5 | 130856 Large plot with fantastic views in Mumbai | 137 | 2018-01-16 |
| 6 | 130857 Modern Design Penthouse in Bangalore | 199 | 2017-03-24 |
分类规则
你定义的分类列表:
- Apartment = ['apartment', 'penthouse', 'duplex']
- House = ['house', 'villa', 'country estate']
- Plot = ['plot', 'land']
- Location = ['New Delhi','Mumbai','Bangalore','Amritsar']
现有尝试代码
import pandas as pd df = pd.read_csv('test_data.csv') # 我可以用循环逐个提取关键词,但如何用Pandas以最少代码实现? for index, values in df.type.iteritems(): for i in Apartment: if i in values: print(i) df_new = pd.DataFrame(df['id'])
解决方案
推荐用向量化操作替代循环,既简洁又高效,尤其适合大数据量场景。这里给你两种可行的实现方式:
方式一:numpy.select() + 正则匹配(最推荐)
这种方法利用Pandas的字符串方法和numpy的多条件选择,完全避免循环,代码量极少:
import pandas as pd import numpy as np # 读取数据 df = pd.read_csv('test_data.csv') # 1. 生成Property_Type分类列 # 构造匹配条件和对应分类 conditions = [ df['Type'].str.lower().str.contains('|'.join(['apartment', 'penthouse', 'duplex'])), df['Type'].str.lower().str.contains('|'.join(['house', 'villa', 'country estate'])), df['Type'].str.lower().str.contains('|'.join(['plot', 'land'])) ] choices = ['Apartment', 'House', 'Plot'] df['Property_Type'] = np.select(conditions, choices, default='Other') # 2. 提取Location列 location_pattern = '|'.join(['New Delhi','Mumbai','Bangalore','Amritsar']) # 用正则提取匹配到的城市 df['Location'] = df['Type'].str.extract(f'({location_pattern})', expand=False) # 生成最终新数据框 df_new = df[['id', 'Property_Type', 'Location', 'agent_id', 'created_at']]
方式二:自定义函数 + apply()(逻辑更直观)
如果需要更灵活的匹配逻辑,比如要处理特殊情况,可以用apply()配合自定义函数:
import pandas as pd df = pd.read_csv('test_data.csv') # 定义分类函数 def get_property_type(type_str): type_lower = type_str.lower() if any(word in type_lower for word in ['apartment', 'penthouse', 'duplex']): return 'Apartment' elif any(word in type_lower for word in ['house', 'villa', 'country estate']): return 'House' elif any(word in type_lower for word in ['plot', 'land']): return 'Plot' return 'Other' # 定义提取城市的函数 def get_location(type_str): locations = ['New Delhi','Mumbai','Bangalore','Amritsar'] for loc in locations: if loc in type_str: return loc return None # 生成新列 df['Property_Type'] = df['Type'].apply(get_property_type) df['Location'] = df['Type'].apply(get_location) # 生成最终数据框 df_new = df[['id', 'Property_Type', 'Location', 'agent_id', 'created_at']]
最终结果示例
运行后df_new的结构如下:
| id | Property_Type | Location | agent_id | created_at |
|---|---|---|---|---|
| 0 | House | New Delhi | 184 | 2018-03-09 |
| 1 | House | Amritsar | 182 | 2017-02-19 |
| 2 | House | NaN | 154 | 2017-04-17 |
| 3 | Plot | New Delhi | 113 | 2017-09-01 |
| 4 | Apartment | Mumbai | 157 | 2017-02-07 |
| 5 | Plot | Mumbai | 137 | 2018-01-16 |
| 6 | Apartment | Bangalore | 199 | 2017-03-24 |
内容的提问来源于stack exchange,提问作者astroluv
相关产品推荐
相关产品推荐

