You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于字符串匹配从DataFrame生成多列?动物列扩展需求

解决方案

核心思路

通过pandas的apply函数提取每行中所有匹配的目标字符串,再将结果拆分到多列中,同时处理无匹配的情况。


1. 准备数据与搜索列表

假设你的目标搜索列表是动物关键词,先定义示例数据和搜索列表:

import pandas as pd

# 定义要搜索的动物列表
search_list = ['dog', 'cat']

# 构造示例DataFrame
data = {
    'weight': [70, 10, 65, 1, 30],
    'String1': [
        'Labrador is a dog',
        'Abyssinian is a cat',
        'German Shepard is a dog',
        'pigeon is a bird',
        'I have a dog and a cat'
    ]
}
df = pd.DataFrame(data)

2. 提取每行匹配的动物

遍历String1列,收集所有命中的关键词;如果无匹配,返回['other']作为默认值:

# 提取所有匹配的动物,无匹配则返回['other']
matched_animals = df['String1'].apply(
    lambda s: [animal for animal in search_list if animal in s] or ['other']
)

3. 拆分结果为多列

将提取到的列表转换为DataFrame,指定列名后与原数据合并:

# 拆分列表为animal_1、animal_2列,不足的位置填充NaN
animals_df = matched_animals.apply(pd.Series).rename(columns={0: 'animal_1', 1: 'animal_2'})

# 合并到原DataFrame
df = pd.concat([df, animals_df], axis=1)

4. 可选:处理空值

如果需要将NaN替换为空字符串或其他默认值,添加以下代码:

df = df.fillna('')

最终输出

处理后的DataFrame如下:

weight                  String1 animal_1 animal_2
0      70        Labrador is a dog      dog         
1      10      Abyssinian is a cat      cat         
2      65  German Shepard is a dog      dog         
3       1         pigeon is a bird    other         
4      30   I have a dog and a cat      dog      cat

内容的提问来源于stack exchange,提问作者Daniel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:42:02