You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Category列值生成verbatim新列的Pandas报错求助

问题解决:根据Category列生成verbatim新列

需求说明

需要为数据集新增verbatim列,规则如下:

  • 当Category为POSITIVE时,取pos列的值
  • 当Category为BETTER THAN COMP时,取better_than_comp列的值
  • 当Category为LESS WELL THAN COMP时,取less_well_than_comp列的值

原始数据

ID  pos          neg  better_than_comp  less_well_than_comp Category            code
1   good service      quick response    price and range     POSITIVE            Satsfied
2   good service      quick response    price and range     BETTER THAN COMP    Speed
3   good service      quick response    price and range     LESS WELL THAN COMP Cost
4   good service      quick response    price and range     LESS WELL THAN COMP Choice

期望输出

ID  pos          neg    better_than_comp    less_well_than_comp Category            code     verbatim
1   good service        quick response      price and range     POSITIVE            Satsfied  good service
2   good service        quick response      price and range     BETTER THAN COMP    Speed    quick response
3   good service        quick response      price and range     LESS WELL THAN COMP Cost     price and range
4   good service        quick response      price and range     LESS WELL THAN COMP Choice   price and range

报错原因分析

你写的代码用了df['Category'].apply(),这里lambda参数x是Category列的单个字符串值(比如'POSITIVE'),不是整行数据对象。所以x['better_than_comp']相当于试图用字符串索引字符串,自然触发TypeError: string indices must be integers。

正确解法

解法1:行级别的apply(直观易懂)

使用df.apply()并指定axis=1,此时lambda参数是整行数据(Series对象),可以直接访问各列:

df['verbatim'] = df.apply(
    lambda row: row['pos'] if row['Category'] == 'POSITIVE'
    else row['better_than_comp'] if row['Category'] == 'BETTER THAN COMP'
    else row['less_well_than_comp'] if row['Category'] == 'LESS WELL THAN COMP'
    else '',  # 处理未匹配到的Category情况
    axis=1
)

解法2:np.select(高效适合大数据集)

用numpy的select方法批量处理,性能比apply更好:

import numpy as np

# 定义匹配条件
conditions = [
    df['Category'] == 'POSITIVE',
    df['Category'] == 'BETTER THAN COMP',
    df['Category'] == 'LESS WELL THAN COMP'
]
# 对应条件的取值列
choices = [
    df['pos'],
    df['better_than_comp'],
    df['less_well_than_comp']
]

df['verbatim'] = np.select(conditions, choices, default='')

解法3:映射列名+lookup(简洁高效)

先建立Category到目标列名的映射,再用lookup提取对应值:

# 建立Category与目标列的映射关系
cat_col_map = {
    'POSITIVE': 'pos',
    'BETTER THAN COMP': 'better_than_comp',
    'LESS WELL THAN COMP': 'less_well_than_comp'
}

# 生成每行要提取的列名
df['target_col'] = df['Category'].map(cat_col_map)
# 根据行索引和列名提取对应值
df['verbatim'] = df.lookup(df.index, df['target_col'])
# 可选:删除中间生成的target_col列
df.drop('target_col', axis=1, inplace=True)

内容的提问来源于stack exchange,提问作者lala345

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 07:15:42