You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按分隔符和长度(整句拆分)分割Pandas文本列

解决DataFrame超长文本按「长度限制+完整句子」拆分的问题

实现思路

要同时满足「单块不超300字符」和「以完整句子结尾」,得先把文本拆成独立句子,再按长度规则拼接成块:

  • 先按句子结束符(.!?)拆分文本,保留标点符号保证句子完整性
  • 逐个累加句子,直到加上下一句会超过300字符时,就把当前累加内容作为一个文本块
  • 剩余句子继续按相同规则生成后续块

代码实现

import pandas as pd
import re

def split_text_into_blocks(text, max_len=300):
    # 按句子结束符拆分,保留标点
    sentences = re.split(r'(?<=[.!?])', text)
    # 过滤空句子和纯空格内容
    sentences = [s.strip() for s in sentences if s.strip()]
    
    blocks = []
    current_block = ""
    
    for sent in sentences:
        # 检查累加当前句子后的长度是否符合要求
        if current_block and len(current_block + " " + sent) <= max_len:
            current_block += " " + sent
        elif len(sent) <= max_len:
            # 当前句子本身不超长度,作为新块起始
            if current_block:
                blocks.append(current_block)
            current_block = sent
        else:
            # 极端情况:单个句子超长,直接截断到限制长度(可按需调整逻辑)
            if current_block:
                blocks.append(current_block)
            blocks.append(sent[:max_len])
            current_block = ""
    # 加入最后剩余的文本块
    if current_block:
        blocks.append(current_block)
    
    # 返回text1、text2...格式的字典
    return {f'text{i+1}': block for i, block in enumerate(blocks)}

# 模拟用户提供的DataFrame结构
data = {
    'doc': ['doc_1', 'doc_1', 'doc_2'],
    'name': ['Texas', 'Texas', 'Georgia'],
    'text': [
        "I have a dream that one day this nation will rise up and live out the true meaning of its creed: 'We hold these truths to be self-evident, that all men are created equal.' I have a dream that one day on the red hills of Georgia, the sons of former slaves and the sons of former slave owners will be able to sit down together at the table of brotherhood.",
        "This dream is a deeply rooted dream in the American dream. I have a dream that my four little children will one day live in a nation where they will not be judged by the color of their skin but by the content of their character. I have a dream today!",
        "I love eating all kinds of food. Chick fil A is my favorite fast food chain, their chicken sandwiches are always fresh and delicious. World of Coke is a great place to visit in Atlanta, you can taste different Coke products from around the world."
    ],
    'len': [3110, 3111, 3325]
}

df = pd.DataFrame(data)

# 应用拆分函数并展开为多列
split_cols = df['text'].apply(lambda x: pd.Series(split_text_into_blocks(x)))

# 合并原DataFrame与拆分后的列
result_df = pd.concat([df.drop('text', axis=1), split_cols], axis=1)

print(result_df)

关键细节说明

  • 正则表达式(?<=[.!?])采用正向肯定后顾,拆分时会保留句子结尾的标点,确保每个拆分出的句子语义完整
  • 处理了「单个句子超长」的极端场景,默认逻辑是直接截断,你可以根据需求改成按词拆分等其他规则
  • 生成的text1、text2等列会自动适配最长文本的块数,不足的位置会填充NaN,可通过fillna('')统一处理空值

内容的提问来源于stack exchange,提问作者Simeon Simeonov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 14:40:00