You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas从DataFrame列提取关键词生成新数据框

问题:用Pandas简洁实现从Type列提取关键词分类

嘿,我来帮你搞定这个Pandas数据清洗的问题!你之前用循环的方式确实有点繁琐,咱们用Pandas的向量化操作就能轻松搞定,而且代码更简洁高效。先理清楚你的需求和现有数据:

背景数据

现有Pandas DataFrame df 数据如下:

idTypeagent_idcreated_at
044525 Stunning 6 bedroom villa in New Delhi1842018-03-09
144859 Villa for sale in Amritsar1822017-02-19
245465 House in Faridabad1542017-04-17
350685 5 Hectre land near New Delhi1132017-09-01
4130728 Duplex in Mumbai1572017-02-07
5130856 Large plot with fantastic views in Mumbai1372018-01-16
6130857 Modern Design Penthouse in Bangalore1992017-03-24

分类规则

你定义的分类列表:

  • Apartment = ['apartment', 'penthouse', 'duplex']
  • House = ['house', 'villa', 'country estate']
  • Plot = ['plot', 'land']
  • Location = ['New Delhi','Mumbai','Bangalore','Amritsar']

现有尝试代码

import pandas as pd
df = pd.read_csv('test_data.csv')
# 我可以用循环逐个提取关键词,但如何用Pandas以最少代码实现?
for index, values in df.type.iteritems():
    for i in Apartment:
        if i in values:
            print(i)
df_new = pd.DataFrame(df['id'])

解决方案

推荐用向量化操作替代循环,既简洁又高效,尤其适合大数据量场景。这里给你两种可行的实现方式:

方式一:numpy.select() + 正则匹配(最推荐)

这种方法利用Pandas的字符串方法和numpy的多条件选择,完全避免循环,代码量极少:

import pandas as pd
import numpy as np

# 读取数据
df = pd.read_csv('test_data.csv')

# 1. 生成Property_Type分类列
# 构造匹配条件和对应分类
conditions = [
    df['Type'].str.lower().str.contains('|'.join(['apartment', 'penthouse', 'duplex'])),
    df['Type'].str.lower().str.contains('|'.join(['house', 'villa', 'country estate'])),
    df['Type'].str.lower().str.contains('|'.join(['plot', 'land']))
]
choices = ['Apartment', 'House', 'Plot']

df['Property_Type'] = np.select(conditions, choices, default='Other')

# 2. 提取Location列
location_pattern = '|'.join(['New Delhi','Mumbai','Bangalore','Amritsar'])
# 用正则提取匹配到的城市
df['Location'] = df['Type'].str.extract(f'({location_pattern})', expand=False)

# 生成最终新数据框
df_new = df[['id', 'Property_Type', 'Location', 'agent_id', 'created_at']]

方式二:自定义函数 + apply()(逻辑更直观)

如果需要更灵活的匹配逻辑,比如要处理特殊情况,可以用apply()配合自定义函数:

import pandas as pd

df = pd.read_csv('test_data.csv')

# 定义分类函数
def get_property_type(type_str):
    type_lower = type_str.lower()
    if any(word in type_lower for word in ['apartment', 'penthouse', 'duplex']):
        return 'Apartment'
    elif any(word in type_lower for word in ['house', 'villa', 'country estate']):
        return 'House'
    elif any(word in type_lower for word in ['plot', 'land']):
        return 'Plot'
    return 'Other'

# 定义提取城市的函数
def get_location(type_str):
    locations = ['New Delhi','Mumbai','Bangalore','Amritsar']
    for loc in locations:
        if loc in type_str:
            return loc
    return None

# 生成新列
df['Property_Type'] = df['Type'].apply(get_property_type)
df['Location'] = df['Type'].apply(get_location)

# 生成最终数据框
df_new = df[['id', 'Property_Type', 'Location', 'agent_id', 'created_at']]

最终结果示例

运行后df_new的结构如下:

idProperty_TypeLocationagent_idcreated_at
0HouseNew Delhi1842018-03-09
1HouseAmritsar1822017-02-19
2HouseNaN1542017-04-17
3PlotNew Delhi1132017-09-01
4ApartmentMumbai1572017-02-07
5PlotMumbai1372018-01-16
6ApartmentBangalore1992017-03-24

内容的提问来源于stack exchange,提问作者astroluv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:17:05