You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中基于另一DataFrame行的列表值创建新DataFrame

问题描述

我有一个示例DataFrame:

| ID | SampleColumn1| SampleColumn2 | SampleColumn3 |
|:-- |:------------:| ------------ :| ------------  |
| 1  |sample Apple  | sample Cherry |sample Lime    |
| 2  |sample Cherry | sample lemon  | sample Grape  |

希望基于这个初始DataFrame创建一个新的DataFrame:如果列表[Apple, Lime, Cherry]中的任一值出现在某一行的任意列中,则新DataFrame对应列标记为1,否则为0。预期输出如下:

| ID | Apple | Lime | Cherry |
| 1  |  1    |  1   |    1   |
| 2  |  0    |  0   |    1   |

我尝试过使用字符串find函数,将每行的Series转为字符串后通过if条件判断是否匹配新DataFrame列名,但出现了逻辑错误。

解决方案

用Pandas的向量化操作可以高效实现需求,避免逐行循环的逻辑问题:

完整代码示例

import pandas as pd

# 初始化原始DataFrame
df = pd.DataFrame({
    'ID': [1, 2],
    'SampleColumn1': ['sample Apple', 'sample Cherry'],
    'SampleColumn2': ['sample Cherry', 'sample lemon'],
    'SampleColumn3': ['sample Lime', 'sample Grape']
})

targets = ['Apple', 'Lime', 'Cherry']

# 生成结果DataFrame,先保留ID列
result = df[['ID']].copy()

# 遍历每个目标关键词,生成对应标记列
for target in targets:
    # 检查每行的所有Sample列是否包含当前关键词,匹配则标记1,否则0
    result[target] = df.filter(like='SampleColumn').apply(
        lambda row: any(target in str(cell) for cell in row), axis=1
    ).astype(int)

print(result)

代码说明

  • df.filter(like='SampleColumn'):自动筛选所有以SampleColumn开头的列,不用手动逐个指定列名
  • apply(..., axis=1):按行处理数据,检查当前行的每个单元格
  • any(target in str(cell) for cell in row):只要行内有一个单元格包含目标关键词,就返回True
  • astype(int):将布尔值True/False转换为1/0的整数格式

运行代码后就能得到你需要的预期输出格式。

内容的提问来源于stack exchange,提问作者plebman952

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:01:00