Pandas合并连续行实现每行单词数≥10且不拆分原句子的实现方案
实现方案
核心逻辑是为连续行生成专属分组标签:遍历每行累计单词数,一旦累计值≥设定的最低阈值(此处为10)就闭合当前分组,后续行进入新分组重新累计,最终按分组聚合即可实现需求,全程不会拆分原有完整句子。
完整实现代码
import pandas as pd # 构造示例DataFrame df = pd.DataFrame(data=[['Tyger Tyger burning.', 3], ['bright.', 1], ['In the forests of the night.', 6], ['What immortal.', 2], ['hand or eye could frame thy fearful symmetry.', 8], ['In what distant deeps or skies.', 6], ['Burnt the fire of thine eyes.', 6] ],columns=['SENTENCE','NO_WORDS']) # 生成分组标签 current_word_sum = 0 group_id = 0 group_list = [] min_word_threshold = 10 # 可按需调整最低单词数阈值 for word_cnt in df['NO_WORDS']: current_word_sum += word_cnt group_list.append(group_id) # 累计单词数达标,开启新分组 if current_word_sum >= min_word_threshold: group_id += 1 current_word_sum = 0 # 按分组聚合得到结果 result_df = df.groupby(group_list).agg( SENTENCE = ('SENTENCE', ' '.join), NO_WORDS = ('NO_WORDS', 'sum') ).reset_index(drop=True) print(result_df)
输出结果
SENTENCE NO_WORDS 0 Tyger Tyger burning. bright. In the forests of... 10 1 What immortal. hand or eye could frame thy fea... 10 2 In what distant deeps or skies. Burnt the fire... 12
内容的提问来源于stack exchange,提问作者rumpet
相关产品推荐
相关产品推荐

