如何根据DataFrame的FEATURE列中+/-符号新增RESULTS结果列
解决方案
问题原因说明
你之前用np.where失败大概率是没有处理+的正则转义问题:+属于正则特殊字符,直接用str.contains('+')会触发正则语法报错,加regex=False匹配字面量符号即可解决。
完整实现代码
import pandas as pd import numpy as np # 构造样例DataFrame,你可以直接跳过这步用自己读取的df data = { 'FEATURE': { 0: '[NA] this is not a feature.', 1: 'entertainment value[+8]', 2: 'image quality[+6]', 3: 'extras [-2]', 4: 'a list cast[+2],oscar winners[+1]' }, 'SENTENCE': { 0: pd.NA, 1: 'starwars trilogy', 2: ' good for an old dvd', 3: 'dvd no had extra features', 4: ' good cast' } } df = pd.DataFrame(data) # 核心逻辑:先判断两类符号的存在情况 has_plus = df['FEATURE'].str.contains('+', regex=False) has_minus = df['FEATURE'].str.contains('-', regex=False) # 按规则生成RESULTS列 df['RESULTS'] = np.where( has_plus & has_minus, 'neutral', np.where( has_plus, 'positive', np.where( has_minus, 'negative', 'unknown' # 两种符号都不存在的情况可自定义取值 ) ) )
运行结果说明
你的样例数据运行后RESULTS列取值如下:
| 行号 | RESULTS |
|---|---|
| 0 | unknown |
| 1 | positive |
| 2 | positive |
| 3 | negative |
| 4 | positive |
如果不需要保留两种符号都不存在的分类,可自行修改最后一个np.where的默认返回值。
内容的提问来源于stack exchange,提问作者chip
相关产品推荐
相关产品推荐

