You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按Tasks列分组拼接Instructions列,满足分句与字符长度限制

pandas按任务分组拼接指令文本(满足长度与句号截断要求)

我有个带Tasks、Instructions、len(存储Instructions字符长度)三列的DataFrame,需要实现以下需求:

  • 按Tasks列分组,对同组的Instructions内容进行拼接
  • 拼接时需截断到下一个句号的位置
  • 拼接后的每行内容字符数最多不超过100
  • 同步更新len列,改为拼接后文本的字符长度

示例输入数据

import pandas as pd

d = {
    'Tasks': ["TaskOne","TaskOne","TaskTwo","TaskTwo"], 
    'Instructions': [
        "longTextofInstructionsforTask1", 
        "longTextofInstructionsforTask1",
        "longTextofInstructionsforTask2 additional text.", 
        "longTextofInstructionsforTask2 additional text."
    ]
}
df = pd.DataFrame(data=d)
df['len'] = df['Instructions'].str.len()

当前DataFrame内容

Tasks                  Instructions                    len
0   TaskOne longTextofInstructionsforTask1                  30
1   TaskOne longTextofInstructionsforTask1                  30

2   TaskTwo longTextofInstructionsforTask2 additional text.  47
3   TaskTwo longTextofInstructionsforTask2 additional text.  47

期望输出结果

Tasks                  Instructions                                                                 len
0   TaskOne longTextofInstructionsforTask1 longTextofInstructionsforTask1                                60
1   TaskTwo longTextofInstructionsforTask2 additional text. longTextofInstructionsforTask2 additional text. 94

实现代码

1. 定义拼接处理函数

这个函数负责完成同组指令的拼接、句号截断和长度控制:

def process_group(instructions):
    full_text = ' '.join(instructions)
    segments = []
    current_segment = ''
    
    for char in full_text:
        current_segment += char
        # 当当前内容接近100字符且遇到句号时,截断保存
        if len(current_segment) >= 90 and char == '.':
            segments.append(current_segment.strip())
            current_segment = ''
    # 处理剩余未截断的内容
    if current_segment.strip():
        segments.append(current_segment.strip())
    
    # 返回分段后的文本及对应长度
    return pd.Series({'Instructions': segments, 'len': [len(seg) for seg in segments]})

2. 分组执行并整理结果

# 按Tasks分组处理
result_df = df.groupby('Tasks')['Instructions'].apply(process_group).reset_index()
# 将分段结果展开为多行
result_df = result_df.explode(['Instructions', 'len']).reset_index(drop=True)

# 打印结果
print(result_df)

补充说明

  • 设置len(current_segment) >=90是为了预留缓冲空间,确保截断到句号后总长度不超过100,可根据实际需求调整该阈值
  • 如果文本中没有句号,函数会自动保留不超过100字符的内容;若需要严格仅按句号截断(即使超过100字符),可修改判断逻辑

内容的提问来源于stack exchange,提问作者Michail Koumpas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 05:25:30