You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按Type列特定词汇先于Scence出现的规则为pandas DataFrame新增列赋值?

pandas DataFrame按规则新增列实现方法

原有代码问题

  • 判断条件逻辑完全错误,df['Type'][i] == "Action" < df['Type'][i] == "Scene"是Python链式比较,实际执行逻辑为(df['Type'][i] == "Action") and ("Action" < df['Type'][i]) and (df['Type'][i] == "Scene"),永远返回False
  • 循环变量i是整数类型,没有append方法,需要直接给DataFrame的新列赋值
  • 没有考虑“出现在之前”是指行的先后顺序,不是字符串值的大小比较

正确实现思路

将相邻两个Scene(注意保证和数据中拼写完全一致,你需求里有时写为Scence,需统一)之间的行划分为同一组,按组规则赋值:

  • 每组的第一个元素是Scene行,赋值为3
  • 同一组内Scene之后的所有行,根据组内第一个非Scene元素的类型赋值:如果是Action统一赋值1,如果是Dialogue统一赋值2

完整实现代码

import pandas as pd

# 测试用例,可替换为你的实际数据
df = pd.DataFrame({
    'Type': ['Scene', 'Action', 'Dialogue', 'Scene', 'Dialogue', 'Action', 'Scene']
})

# 步骤1:为每个相邻Scene区间分配组ID,每遇到一个Scene组ID+1
df['group_id'] = (df['Type'] == 'Scene').cumsum()

# 步骤2:按组赋值
def group_assign(group):
    res = []
    # 第一个元素是Scene,赋值3
    res.append(3)
    # 取组内Scene之后的所有元素
    other_rows = group[1:]
    if len(other_rows) == 0:
        return pd.Series(res, index=group.index)
    # 取组内第一个非Scene的类型
    first_non_scene = other_rows.iloc[0]
    assign_val = 1 if first_non_scene == 'Action' else 2
    # 组内剩余行统一赋值
    res += [assign_val] * len(other_rows)
    return pd.Series(res, index=group.index)

df['new_col'] = df.groupby('group_id')['Type'].apply(group_assign).droplevel(0)

# 可选:删除辅助用的group_id列
df = df.drop('group_id', axis=1)

测试结果

上述测试用例运行后,最终df的结果如下:

Typenew_col
Scene3
Action1
Dialogue1
Scene3
Dialogue2
Action2
Scene3

内容的提问来源于stack exchange,提问作者Ahmed Abousari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 19:15:00