如何用Pandas依据列中Action/Drama字符串修改列值?
Pandas 电影类型列值修改实现方案
核心逻辑回顾
- 包含"Action"(无论是否同时存在Drama)→ 设为0
- 不包含"Action"但包含"Drama" → 设为1
- 两者都不包含的情况可自定义默认值(示例中设为
NaN)
方法一:使用numpy.where嵌套判断(高效简洁)
利用np.where的优先级判断特性,一次性完成列映射:
import pandas as pd import numpy as np # 构造示例数据(替换为你的实际数据集) df = pd.DataFrame({ 'Genres': [ 'Action Comedy Adventure', 'Drama Crime Horror', 'Action Drama Adventure' ] }) # 执行映射 df['Genres_Code'] = np.where( df['Genres'].str.contains('Action'), 0, np.where(df['Genres'].str.contains('Drama'), 1, np.nan) )
方法二:自定义函数+apply(逻辑直观)
如果需要更灵活的扩展逻辑,可自定义映射函数后用apply处理:
import pandas as pd import numpy as np def genre_to_code(genre_str): if 'Action' in genre_str: return 0 elif 'Drama' in genre_str: return 1 return np.nan # 无匹配时的默认值,可修改为其他值(如2) df = pd.DataFrame({ 'Genres': [ 'Action Comedy Adventure', 'Drama Crime Horror', 'Action Drama Adventure' ] }) df['Genres_Code'] = df['Genres'].apply(genre_to_code)
验证结果
执行上述代码后,输出的df如下:
Genres Genres_Code 0 Action Comedy Adventure 0.0 1 Drama Crime Horror 1.0 2 Action Drama Adventure 0.0
完全符合需求。
内容的提问来源于stack exchange,提问作者Nathan Pared
相关产品推荐
相关产品推荐

