如何对字符串数组执行re.sub()并保留分割点(格式)
问题描述
我有一组字符串数组,每个元素对应带格式的段落文本片段(类似HTML <span> 标签)。需要对整段文本执行类似re.sub()的替换操作,但要保留原有分割点(即保留格式),不使用re.sub()的方案也可接受。
无格式场景的实现示例
不考虑格式时,直接对完整字符串处理的代码如下:
import re def repl(match): ix = next(i for i, val in enumerate(match.groups()) if val is not None) return f'{match.group(0)} ({keyword_annotate_map[ix]["annotation"]})' before = 'keyword1 asdafljd asdanfnfg keyword2 snbsbsdbns' keyword_annotate_map = [ { 'regex': 'keyword1', 'annotation': 'annotation1' }, { 'regex': 'keyword2', 'annotation': 'annotation2' } ] # 修正原代码正则错误:用|分隔关键词而非逗号 after = re.sub(rf'({"|".join(re.escape(val["regex"]) for val in keyword_annotate_map)})', repl, before, flags=re.IGNORECASE) print(after) # 输出: keyword1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) snbsbsdbns
带格式场景的输入与预期输出
当文本被分割为数组时,需要保留分割结构,仅替换对应关键词:
# ''.join(before) 得到无格式原始字符串 before = ['key', 'word1 asdafljd asdanfnfg keyword2 ', 'snbsbsdbns'] # 预期输出 print(after) # ['key', 'word1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) ', 'snbsbsdbns']
解决方案
基础方案:逐片段处理(无跨片段关键词)
核心思路是对数组中的每个文本片段单独执行替换,既完成关键词标注,又保留原分割结构。
import re def repl(match): ix = next(i for i, val in enumerate(match.groups()) if val is not None) return f'{match.group(0)} ({keyword_annotate_map[ix]["annotation"]})' before = ['key', 'word1 asdafljd asdanfnfg keyword2 ', 'snbsbsdbns'] keyword_annotate_map = [ { 'regex': 'keyword1', 'annotation': 'annotation1' }, { 'regex': 'keyword2', 'annotation': 'annotation2' } ] # 构建正则模式:用|分隔转义后的关键词 pattern = re.compile( rf'({"|".join(re.escape(val["regex"]) for val in keyword_annotate_map)})', flags=re.IGNORECASE ) # 对每个片段单独执行替换 after = [pattern.sub(repl, segment) for segment in before] print(after) # 输出: ['key', 'word1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) ', 'snbsbsdbns']
进阶方案:处理跨片段关键词
如果存在关键词被分割在两个片段中的情况(比如['key', 'word1']),可以先拼接完整文本执行替换,再根据原片段长度重新分割:
import re def repl(match): ix = next(i for i, val in enumerate(match.groups()) if val is not None) return f'{match.group(0)} ({keyword_annotate_map[ix]["annotation"]})' before = ['key', 'word1 asdafljd asdanfnfg keyword2 ', 'snbsbsdbns'] keyword_annotate_map = [ { 'regex': 'keyword1', 'annotation': 'annotation1' }, { 'regex': 'keyword2', 'annotation': 'annotation2' } ] # 记录原片段长度,用于后续分割 segment_lengths = [len(seg) for seg in before] # 拼接为完整文本 full_text = ''.join(before) # 执行替换 pattern = re.compile(rf'({"|".join(re.escape(val["regex"]) for val in keyword_annotate_map)})', flags=re.IGNORECASE) full_annotated_text = pattern.sub(repl, full_text) # 根据原长度重新分割数组 after = [] current_pos = 0 for length in segment_lengths: after.append(full_annotated_text[current_pos:current_pos+length]) current_pos += length print(after) # 输出: ['key', 'word1 (annotation1) asdafljd asdanfnfg keyword2 (annotation2) ', 'snbsbsdbns']
注意:该方案会因替换后文本长度变化导致分割位置偏移,仅在确有跨片段关键词需求时使用。
内容的提问来源于stack exchange,提问作者tearfur
相关产品推荐
相关产品推荐

