如何对OCR提取的字符串列表分句并保留原元素组成映射
分句结果与原OCR文本元素的映射方案
当你用Azure Form Recognizer得到带位置坐标的OCR文本行列表后,直接拼接分句会丢失原行的关联信息。以下是两种可行的解决方案,能帮你建立分句结果到原OCR行索引的映射,从而保留位置坐标。
方法一:基于字符位置的精准匹配
核心思路是先记录每一行OCR文本在拼接后总字符串中的字符位置区间,再结合NLTK/Spacy返回的句子字符位置,匹配出对应原行。
实现步骤
- 遍历原OCR行列表,记录每行的起始/结束字符位置及索引;
- 拼接所有行得到完整文本,用分句工具处理并获取每个句子的字符位置;
- 对比句子位置与原行位置,找到所有覆盖句子的原行索引。
代码示例
import spacy # 示例OCR行列表,包含原位置对应的索引(这里用列表索引+1表示行号) ocr_lines = [ "Much of Singapore's infrastructure had been destroyed during the war", "including those needed to supply utilities.", "A shortage of food led to malnutrition, disease,", " and rampant crime and violence.", "A series of strikes in 1947 caused massive,", "stoppages in public transport and other services." ] # 构建原行的字符位置映射 line_positions = [] current_char_pos = 0 for line_idx, line in enumerate(ocr_lines): line_length = len(line) line_positions.append({ "line_number": line_idx + 1, "start": current_char_pos, "end": current_char_pos + line_length }) current_char_pos += line_length # 若原行之间有空格/换行,需同步调整这里的偏移量 # 拼接文本并分句 full_text = "".join(ocr_lines) nlp = spacy.load("en_core_web_sm") doc = nlp(full_text) # 建立句子到原行的映射 sentence_to_lines = {} for sent_idx, sent in enumerate(doc.sents, 1): sent_start = sent.start_char sent_end = sent.end_char # 筛选所有与当前句子位置重叠的原行 related_lines = [lp["line_number"] for lp in line_positions if not (lp["end"] <= sent_start or lp["start"] >= sent_end)] sentence_to_lines[f"句子{sent_idx}"] = tuple(related_lines) print(sentence_to_lines) # 输出结果: # {'句子1': (1, 2), '句子2': (3, 4), '句子3': (5, 6)}
注意事项
- 拼接文本时需保证与分句工具处理的文本完全一致,若原OCR行之间存在换行/空格,需在拼接时同步添加,否则字符位置会错位;
- 该方法适用于所有分句场景,匹配精度高,是优先推荐的方案。
方法二:基于句子片段的逐步补全
核心思路是维护一个句子缓冲区,逐步拼接原OCR行,直到缓冲区内容被分句工具判定为完整句子,此时记录参与拼接的原行索引。
实现步骤
- 初始化缓冲区和当前行索引列表;
- 逐行添加到缓冲区,用分句工具检查缓冲区内容是否为完整句子;
- 若判定为完整句子,建立映射并重置缓冲区;
- 处理循环结束后剩余的未完成句子片段。
代码示例
import spacy ocr_lines = [ "Much of Singapore's infrastructure had been destroyed during the war", "including those needed to supply utilities.", "A shortage of food led to malnutrition, disease,", " and rampant crime and violence.", "A series of strikes in 1947 caused massive,", "stoppages in public transport and other services." ] nlp = spacy.load("en_core_web_sm") sentence_buffer = [] current_line_numbers = [] sentence_to_lines = {} sentence_idx = 0 for line_idx, line in enumerate(ocr_lines): current_line_numbers.append(line_idx + 1) sentence_buffer.append(line) combined_text = "".join(sentence_buffer) # 用Spacy判断缓冲区是否为完整句子 doc = nlp(combined_text) sents = list(doc.sents) if len(sents) == 1 and sents[0].end_char == len(combined_text): sentence_idx += 1 sentence_to_lines[f"句子{sentence_idx}"] = tuple(current_line_numbers) # 重置缓冲区 sentence_buffer = [] current_line_numbers = [] # 处理剩余的未完成句子 if sentence_buffer: sentence_idx += 1 sentence_to_lines[f"句子{sentence_idx}"] = tuple(current_line_numbers) print(sentence_to_lines)
注意事项
- 该方法适合处理原OCR行拆分不规则的场景;
- 对于包含缩写(如
Mr.)的句子,分句工具可能误判,需根据实际场景调整判断逻辑。
内容的提问来源于stack exchange,提问作者newbie101
相关产品推荐
相关产品推荐

