You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按case_number分组并生成指定格式的合并文本列?

按Case分组合并文本的实现方案

问题描述

现有如下Pandas DataFrame:

import pandas as pd
new_df = pd.DataFrame([('231', '2122', '1', 'some_text', 'agent_text_1', 'cust text_1'),
                   ('231', '2682', '2', 'some_text', 'agent_text_2', 'cust text_2'),
                   ('232', '3982', '1', 'some_text', 'agent_text_1', 'cust text_1'),
                   ('233', '1503', '1', 'some_text', 'agent_text_1', 'cust text_1'),
                   ('233', '1692', '2', 'some_text', 'agent_text_2', 'cust text_2'),
                   ('233', '4113', '3', 'some_text', 'agent_text_3', 'cust text_3')
                   ],
                  columns=['case_number', 'call_id', 'call_order', 'description', 'agent_text', 'cust_text'])

需求:

  • 按case_number对行分组
  • 创建all_texts_combined列,格式为:summary: {description}. [每个call的文本片段]
  • 同一case_number下的description仅显示一次
  • 每个call的文本片段格式:Call order: {call_order}. Agent said: {agent_text}. Customer said: {cust_text}.

尝试的代码存在问题:丢失call_order信息,无法保留description并放在文本开头:

grouped_df = new_df.groupby(['case_number', 'call_order']).agg({'agent_text': ' '.join, 'cust_text': ' '.join}).reset_index()
grouped_df = grouped_df .groupby(['case_number']).agg({'agent_text': ' '.join, 'cust_text': ' '.join}).reset_index()

正确实现方法

步骤1:生成单个call的文本片段

先为每行生成对应call的标准化文本:

new_df['call_segment'] = new_df.apply(
    lambda row: f"Call order: {row['call_order']}. Agent said: {row['agent_text']}. Customer said: {row['cust_text']}.",
    axis=1
)

步骤2:分组聚合并拼接完整文本

按case_number分组,提取唯一的description,再拼接所有call片段:

# 分组聚合,获取每个case的description和所有call片段
grouped_data = new_df.groupby('case_number').agg(
    first_description=('description', 'first'),
    combined_calls=('call_segment', ' '.join)
).reset_index()

# 拼接成最终的all_texts_combined列
grouped_data['all_texts_combined'] = grouped_data.apply(
    lambda row: f"summary: {row['first_description']}. {row['combined_calls']}",
    axis=1
)

# 保留需要的列
grouped_df = grouped_data[['case_number', 'all_texts_combined']]

最终结果示例

grouped_df的输出如下:

case_numberall_texts_combined
231summary: some_text. Call order: 1. Agent said: agent_text_1. Customer said: cust text_1. Call order: 2. Agent said: agent_text_2. Customer said: cust text_2.
232summary: some_text. Call order: 1. Agent said: agent_text_1. Customer said: cust text_1.
233summary: some_text. Call order: 1. Agent said: agent_text_1. Customer said: cust text_1. Call order: 2. Agent said: agent_text_2. Customer said: cust text_2. Call order: 3. Agent said: agent_text_3. Customer said: cust text_3.

内容的提问来源于stack exchange,提问作者Bilal Sedef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 03:53:16