You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对字符串数组执行re.sub()并保留分割点(格式)

问题描述

我有一组字符串数组,每个元素对应带格式的段落文本片段(类似HTML <span> 标签)。需要对整段文本执行类似re.sub()的替换操作,但要保留原有分割点(即保留格式),不使用re.sub()的方案也可接受。

无格式场景的实现示例

不考虑格式时,直接对完整字符串处理的代码如下:

import re

def repl(match):
    ix = next(i for i, val in enumerate(match.groups()) if val is not None)
    return f'{match.group(0)} ({keyword_annotate_map[ix]["annotation"]})'

before = 'keyword1 asdafljd asdanfnfg keyword2 snbsbsdbns'

keyword_annotate_map = [
    { 'regex': 'keyword1', 'annotation': 'annotation1' },
    { 'regex': 'keyword2', 'annotation': 'annotation2' }
]

# 修正原代码正则错误:用|分隔关键词而非逗号
after = re.sub(rf'({"|".join(re.escape(val["regex"]) for val in keyword_annotate_map)})', repl, before, flags=re.IGNORECASE)
print(after) # 输出: keyword1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) snbsbsdbns

带格式场景的输入与预期输出

当文本被分割为数组时,需要保留分割结构,仅替换对应关键词:

# ''.join(before) 得到无格式原始字符串
before = ['key', 'word1 asdafljd asdanfnfg keyword2 ', 'snbsbsdbns']

# 预期输出
print(after) # ['key', 'word1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) ', 'snbsbsdbns']
解决方案

基础方案:逐片段处理(无跨片段关键词)

核心思路是对数组中的每个文本片段单独执行替换,既完成关键词标注,又保留原分割结构。

import re

def repl(match):
    ix = next(i for i, val in enumerate(match.groups()) if val is not None)
    return f'{match.group(0)} ({keyword_annotate_map[ix]["annotation"]})'

before = ['key', 'word1 asdafljd asdanfnfg keyword2 ', 'snbsbsdbns']

keyword_annotate_map = [
    { 'regex': 'keyword1', 'annotation': 'annotation1' },
    { 'regex': 'keyword2', 'annotation': 'annotation2' }
]

# 构建正则模式:用|分隔转义后的关键词
pattern = re.compile(
    rf'({"|".join(re.escape(val["regex"]) for val in keyword_annotate_map)})',
    flags=re.IGNORECASE
)

# 对每个片段单独执行替换
after = [pattern.sub(repl, segment) for segment in before]

print(after)
# 输出: ['key', 'word1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) ', 'snbsbsdbns']

进阶方案:处理跨片段关键词

如果存在关键词被分割在两个片段中的情况(比如['key', 'word1']),可以先拼接完整文本执行替换,再根据原片段长度重新分割:

import re

def repl(match):
    ix = next(i for i, val in enumerate(match.groups()) if val is not None)
    return f'{match.group(0)} ({keyword_annotate_map[ix]["annotation"]})'

before = ['key', 'word1 asdafljd asdanfnfg keyword2 ', 'snbsbsdbns']
keyword_annotate_map = [
    { 'regex': 'keyword1', 'annotation': 'annotation1' },
    { 'regex': 'keyword2', 'annotation': 'annotation2' }
]

# 记录原片段长度,用于后续分割
segment_lengths = [len(seg) for seg in before]
# 拼接为完整文本
full_text = ''.join(before)

# 执行替换
pattern = re.compile(rf'({"|".join(re.escape(val["regex"]) for val in keyword_annotate_map)})', flags=re.IGNORECASE)
full_annotated_text = pattern.sub(repl, full_text)

# 根据原长度重新分割数组
after = []
current_pos = 0
for length in segment_lengths:
    after.append(full_annotated_text[current_pos:current_pos+length])
    current_pos += length

print(after)
# 输出: ['key', 'word1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) ', 'snbsbsdbns']

注意:该方案会因替换后文本长度变化导致分割位置偏移,仅在确有跨片段关键词需求时使用。


内容的提问来源于stack exchange,提问作者tearfur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 11:45:43