You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件拆分Pandas DataFrame列的技术实现问题

处理Pandas DataFrame中标签列的提取问题

问题场景

原始DataFrame如下:

VideoIdTags1Tags2
1234NoneNone
3456Movie:RockyGenre:Drama
5678Genre:RomanceMovie:Love Me
7890Movie:ScreamGenre:Horror
9012Movie:ShrekNone

需求:

  • 从Tags1和Tags2列中,将Movie标题提取到新的Movie列,Genre类型提取到新的Genre列
  • 兼容两列中Movie和Genre位置互换的情况
  • 遇到None值时保留None,不触发错误

原有代码存在两个问题:

Movies = df['Tags1'].apply(lambda x: x.split(':')[1] if 'Movie' in x else x) 
Genres = df['Tags2'].apply(lambda x: x.split(':')[1] if 'Genre' in x else x) 
  1. 处理None值时会触发索引越界错误(None无split方法)
  2. 无法处理Tags1为Genre、Tags2为Movie的互换场景

期望结果DataFrame:

VideoIdMovieGenre
1234NoneNone
3456RockyDrama
5678Love MeRomance
7890ScreamHorror
9012ShrekNone

解决方案

通过自定义函数遍历每行标签列,实现兼容None值和位置互换的提取逻辑:

import pandas as pd

# 构造原始数据
data = {
    'VideoId': [1234, 3456, 5678, 7890, 9012],
    'Tags1': [None, 'Movie:Rocky', 'Genre:Romance', 'Movie:Scream', 'Movie:Shrek'],
    'Tags2': [None, 'Genre:Drama', 'Movie:Love Me', 'Genre:Horror', None]
}
df = pd.DataFrame(data)

def extract_tags(row):
    movie = None
    genre = None
    # 遍历当前行的两个标签列
    for tag in [row['Tags1'], row['Tags2']]:
        if tag is None:
            continue
        if 'Movie:' in tag:
            movie = tag.split(':')[1]
        elif 'Genre:' in tag:
            genre = tag.split(':')[1]
    return pd.Series([movie, genre], index=['Movie', 'Genre'])

# 合并提取结果到原DataFrame
result_df = df[['VideoId']].join(df.apply(extract_tags, axis=1))
print(result_df)

逻辑说明

  • 函数extract_tags初始化movie和genre为None,遍历每行的两个标签值
  • 遇到None直接跳过,避免报错
  • 根据标签前缀识别类型,提取对应内容,不管标签在Tags1还是Tags2都能正确匹配
  • 最后通过join将提取出的列与原DataFrame的VideoId列合并,得到目标结果

内容的提问来源于stack exchange,提问作者Aleksei Wolff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:33:16