基于Category列值生成verbatim新列的Pandas报错求助
问题解决:根据Category列生成verbatim新列
需求说明
需要为数据集新增verbatim列,规则如下:
- 当
Category为POSITIVE时,取pos列的值 - 当
Category为BETTER THAN COMP时,取better_than_comp列的值 - 当
Category为LESS WELL THAN COMP时,取less_well_than_comp列的值
原始数据
ID pos neg better_than_comp less_well_than_comp Category code 1 good service quick response price and range POSITIVE Satsfied 2 good service quick response price and range BETTER THAN COMP Speed 3 good service quick response price and range LESS WELL THAN COMP Cost 4 good service quick response price and range LESS WELL THAN COMP Choice
期望输出
ID pos neg better_than_comp less_well_than_comp Category code verbatim 1 good service quick response price and range POSITIVE Satsfied good service 2 good service quick response price and range BETTER THAN COMP Speed quick response 3 good service quick response price and range LESS WELL THAN COMP Cost price and range 4 good service quick response price and range LESS WELL THAN COMP Choice price and range
报错原因分析
你写的代码用了df['Category'].apply(),这里lambda参数x是Category列的单个字符串值(比如'POSITIVE'),不是整行数据对象。所以x['better_than_comp']相当于试图用字符串索引字符串,自然触发TypeError: string indices must be integers。
正确解法
解法1:行级别的apply(直观易懂)
使用df.apply()并指定axis=1,此时lambda参数是整行数据(Series对象),可以直接访问各列:
df['verbatim'] = df.apply( lambda row: row['pos'] if row['Category'] == 'POSITIVE' else row['better_than_comp'] if row['Category'] == 'BETTER THAN COMP' else row['less_well_than_comp'] if row['Category'] == 'LESS WELL THAN COMP' else '', # 处理未匹配到的Category情况 axis=1 )
解法2:np.select(高效适合大数据集)
用numpy的select方法批量处理,性能比apply更好:
import numpy as np # 定义匹配条件 conditions = [ df['Category'] == 'POSITIVE', df['Category'] == 'BETTER THAN COMP', df['Category'] == 'LESS WELL THAN COMP' ] # 对应条件的取值列 choices = [ df['pos'], df['better_than_comp'], df['less_well_than_comp'] ] df['verbatim'] = np.select(conditions, choices, default='')
解法3:映射列名+lookup(简洁高效)
先建立Category到目标列名的映射,再用lookup提取对应值:
# 建立Category与目标列的映射关系 cat_col_map = { 'POSITIVE': 'pos', 'BETTER THAN COMP': 'better_than_comp', 'LESS WELL THAN COMP': 'less_well_than_comp' } # 生成每行要提取的列名 df['target_col'] = df['Category'].map(cat_col_map) # 根据行索引和列名提取对应值 df['verbatim'] = df.lookup(df.index, df['target_col']) # 可选:删除中间生成的target_col列 df.drop('target_col', axis=1, inplace=True)
内容的提问来源于stack exchange,提问作者lala345
相关产品推荐
相关产品推荐

