You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则提取指定模式句子:长原始句子截断优化方法问询

简洁实现长句子优化方案

当然有简洁的实现方式!核心思路是先定位到匹配模式在原始句子中的词位置,然后精准截取模式前后各30个词,用省略号替换超出范围的部分。下面是基于Python的完整实现,适配你的需求:

步骤说明

  1. 通过正则找到目标模式在句子中的字符位置,再映射到词列表的索引;
  2. 判断句子长度是否超过200词,若超过则截取模式前后30词范围;
  3. 用省略号补充被截断的前后部分,拼接成优化后的句子。

代码实现

import re

def optimize_long_sentences(sentences, my_words_markers, max_words=200, keep_context=30):
    # 构建你的匹配模式
    pattern = r'\b' + ' (?:\w+ )?(?:\w+ )?'.join(my_words_markers) + r'\b'
    optimized = []
    
    for sent in sentences:
        # 拆分句子为词列表(用空格拆分,若需更精准分词可替换为nltk.word_tokenize等)
        word_list = sent.split()
        total_words = len(word_list)
        
        # 直接保留短句子
        if total_words <= max_words:
            optimized.append(sent)
            continue
        
        # 找到模式匹配的位置(理论上所有传入句子都能匹配到)
        match = re.search(pattern, sent)
        if not match:
            optimized.append(sent)
            continue
        
        # 计算每个词的字符位置范围,用于映射词索引
        word_positions = []
        current_pos = 0
        for word in word_list:
            word_end = current_pos + len(word)
            word_positions.append((current_pos, word_end))
            current_pos = word_end + 1  # 加上空格的长度
        
        # 定位匹配模式对应的起始和结束词索引
        start_word_idx = None
        end_word_idx = None
        for idx, (pos_start, pos_end) in enumerate(word_positions):
            if pos_end >= match.start() and start_word_idx is None:
                start_word_idx = idx
            if pos_start <= match.end():
                end_word_idx = idx
        
        # 计算要保留的词范围,避免索引越界
        left_start = max(0, start_word_idx - keep_context)
        right_end = min(total_words, end_word_idx + keep_context)
        
        # 构建优化后的句子片段
        parts = []
        if left_start > 0:
            parts.append("...")
        parts.extend(word_list[left_start:right_end+1])
        if right_end < total_words - 1:
            parts.append("...")
        
        optimized_sentence = ' '.join(parts)
        optimized.append(optimized_sentence)
    
    return optimized

关键细节说明

  • 分词灵活调整:示例用split()简单拆分,如果你需要区分标点和单词,可以替换为nltk.word_tokenize()或spaCy的分词器,只需修改word_list的生成逻辑;
  • 精准定位:通过计算每个词的字符位置,把正则匹配的字符范围映射到词索引,确保模式所在的词位置不会偏移;
  • 边界处理:用max(0, ...)和min(total_words, ...)避免索引越界,同时通过判断截断情况决定是否添加省略号,保证输出格式自然。

你可以直接把原始句子列表、my_words_markers传入这个函数,就能得到符合要求的优化结果啦!

内容的提问来源于stack exchange,提问作者Fed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:15:23