You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于topic与label条件创建哑列并填充对应value值?

解决方案

首先修正可复现数据,补充缺失的label列:

import pandas as pd

raw_df = pd.DataFrame({
    'reviewId': ['01', '02', '03', '04', '05'],
    'topic': [2, 2, 0, 5, 1],
    'value': [-4, 9, -7, -1, 38],  # 转为数值类型方便后续处理
    'label': ['negative', 'positive', 'negative', 'negative', 'positive']
})

方法一:批量创建列并条件赋值

  1. 生成所有目标列名(覆盖topic1-6与正负标签的组合):
topic_labels = [f"t{t}{l[0]}" for t in range(1,7) for l in ['positive', 'negative']]
  1. 初始化所有目标列为0:
for col in topic_labels:
    raw_df[col] = 0
  1. 遍历每行匹配条件并填充值:
for idx, row in raw_df.iterrows():
    if row['topic'] == 0:
        continue  # 跳过未分配主题的行
    target_col = f"t{row['topic']}{row['label'][0]}"
    raw_df.loc[idx, target_col] = row['value']

方法二:透视表合并(大数据量更高效)

# 生成临时透视表,仅处理已分配主题的行
pivot_df = raw_df[raw_df['topic'] != 0].assign(
    col_name=lambda x: 't' + x['topic'].astype(str) + x['label'].str[0]
).pivot(
    index='reviewId',
    columns='col_name',
    values='value'
).fillna(0).astype(int)

# 合并回原表,补全缺失列并填充0
result_df = raw_df.merge(pivot_df, on='reviewId', how='left')
for col in topic_labels:
    if col not in result_df.columns:
        result_df[col] = 0

# 调整列顺序匹配目标结构
result_df = result_df[['reviewId', 'topic', 'value', 'label'] + topic_labels]

最终结果

处理后的数据结构与目标表完全一致:

reviewIdtopicvaluelabelt1pt1nt2pt2nt3pt3nt4pt4nt5pt5nt6pt6n
012-4negative000-400000000
0229positive009000000000
030-7negative000000000000
045-1negative000000000-100
05138positive3800000000000

内容的提问来源于stack exchange,提问作者Dewani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 09:15:45