You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于sentence列分组,根据entity列值新增Way/Purpose列的Pandas需求

按Sentence分组生成Way/Purpose列解决方案

原始数据

你的原始DataFrame定义如下:

import pandas as pd

df1 = pd.DataFrame({
    'sentence': ['A', "A", "A", "A", 'A', 'B', "B", 'B'],
    'entity': ['Stay home', "Stay home", "WAY", "WAY", "Stay home", 'Go outside', "Go outside", "purpose"],
    'token' : ['Severe weather', "raining", "smt", "SMT0", "Windy", 'Sunny', "Good weather", "smt"]
})

需求说明

按sentence列分组,处理每个分组内的数据:

  • 将非Way/Purpose的实体作为主entity,对应token合并为字符串
  • 若分组内存在Way,合并其对应的token放入Way列,无则填充NaN
  • 若分组内存在Purpose,合并其对应的token放入Purpose列,无则填充NaN

实现代码

# 统一entity的大小写,避免大小写不匹配问题
df1['entity'] = df1['entity'].str.title()

# 定义分组聚合逻辑
def group_agg(group):
    # 提取主实体(排除Way/Purpose的第一个实体值)
    main_entity = group[~group['entity'].isin(['Way', 'Purpose'])]['entity'].iloc[0]
    # 合并主实体对应的token
    main_token = ', '.join(group[~group['entity'].isin(['Way', 'Purpose'])]['token'])
    # 处理Way列的token合并
    way_tokens = group[group['entity'] == 'Way']['token']
    way_col = ', '.join(way_tokens) if not way_tokens.empty else pd.NA
    # 处理Purpose列的token合并
    purpose_tokens = group[group['entity'] == 'Purpose']['token']
    purpose_col = ', '.join(purpose_tokens) if not purpose_tokens.empty else pd.NA
    
    return pd.Series(
        [main_entity, main_token, way_col, purpose_col],
        index=['entity', 'token', 'Way', 'Purpose']
    )

# 执行分组聚合并重置索引
result_df = df1.groupby('sentence').apply(group_agg).reset_index()

输出结果

运行代码后得到的结果:

sentence       entity                          token       Way Purpose
0        A    Stay home  Severe weather, raining, Windy  smt, SMT0    <NA>
1        B  Go outside            Sunny, Good weather      <NA>     smt

内容的提问来源于stack exchange,提问作者xavi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:30:50