按Tasks列分组拼接Instructions列,满足分句与字符长度限制
pandas按任务分组拼接指令文本(满足长度与句号截断要求)
我有个带Tasks、Instructions、len(存储Instructions字符长度)三列的DataFrame,需要实现以下需求:
- 按
Tasks列分组,对同组的Instructions内容进行拼接 - 拼接时需截断到下一个句号的位置
- 拼接后的每行内容字符数最多不超过100
- 同步更新
len列,改为拼接后文本的字符长度
示例输入数据
import pandas as pd d = { 'Tasks': ["TaskOne","TaskOne","TaskTwo","TaskTwo"], 'Instructions': [ "longTextofInstructionsforTask1", "longTextofInstructionsforTask1", "longTextofInstructionsforTask2 additional text.", "longTextofInstructionsforTask2 additional text." ] } df = pd.DataFrame(data=d) df['len'] = df['Instructions'].str.len()
当前DataFrame内容
Tasks Instructions len 0 TaskOne longTextofInstructionsforTask1 30 1 TaskOne longTextofInstructionsforTask1 30 2 TaskTwo longTextofInstructionsforTask2 additional text. 47 3 TaskTwo longTextofInstructionsforTask2 additional text. 47
期望输出结果
Tasks Instructions len 0 TaskOne longTextofInstructionsforTask1 longTextofInstructionsforTask1 60 1 TaskTwo longTextofInstructionsforTask2 additional text. longTextofInstructionsforTask2 additional text. 94
实现代码
1. 定义拼接处理函数
这个函数负责完成同组指令的拼接、句号截断和长度控制:
def process_group(instructions): full_text = ' '.join(instructions) segments = [] current_segment = '' for char in full_text: current_segment += char # 当当前内容接近100字符且遇到句号时,截断保存 if len(current_segment) >= 90 and char == '.': segments.append(current_segment.strip()) current_segment = '' # 处理剩余未截断的内容 if current_segment.strip(): segments.append(current_segment.strip()) # 返回分段后的文本及对应长度 return pd.Series({'Instructions': segments, 'len': [len(seg) for seg in segments]})
2. 分组执行并整理结果
# 按Tasks分组处理 result_df = df.groupby('Tasks')['Instructions'].apply(process_group).reset_index() # 将分段结果展开为多行 result_df = result_df.explode(['Instructions', 'len']).reset_index(drop=True) # 打印结果 print(result_df)
补充说明
- 设置
len(current_segment) >=90是为了预留缓冲空间,确保截断到句号后总长度不超过100,可根据实际需求调整该阈值 - 如果文本中没有句号,函数会自动保留不超过100字符的内容;若需要严格仅按句号截断(即使超过100字符),可修改判断逻辑
内容的提问来源于stack exchange,提问作者Michail Koumpas
相关产品推荐
相关产品推荐

