You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对OCR提取的字符串列表分句并保留原元素组成映射

分句结果与原OCR文本元素的映射方案

当你用Azure Form Recognizer得到带位置坐标的OCR文本行列表后,直接拼接分句会丢失原行的关联信息。以下是两种可行的解决方案,能帮你建立分句结果到原OCR行索引的映射,从而保留位置坐标。

方法一:基于字符位置的精准匹配

核心思路是先记录每一行OCR文本在拼接后总字符串中的字符位置区间,再结合NLTK/Spacy返回的句子字符位置,匹配出对应原行。

实现步骤

  1. 遍历原OCR行列表,记录每行的起始/结束字符位置及索引;
  2. 拼接所有行得到完整文本,用分句工具处理并获取每个句子的字符位置;
  3. 对比句子位置与原行位置,找到所有覆盖句子的原行索引。

代码示例

import spacy

# 示例OCR行列表,包含原位置对应的索引(这里用列表索引+1表示行号)
ocr_lines = [
    "Much of Singapore's infrastructure had been destroyed during the war",
    "including those needed to supply utilities.",
    "A shortage of food led to malnutrition, disease,",
    " and rampant crime and violence.",
    "A series of strikes in 1947 caused massive,",
    "stoppages in public transport and other services."
]

# 构建原行的字符位置映射
line_positions = []
current_char_pos = 0
for line_idx, line in enumerate(ocr_lines):
    line_length = len(line)
    line_positions.append({
        "line_number": line_idx + 1,
        "start": current_char_pos,
        "end": current_char_pos + line_length
    })
    current_char_pos += line_length  # 若原行之间有空格/换行,需同步调整这里的偏移量

# 拼接文本并分句
full_text = "".join(ocr_lines)
nlp = spacy.load("en_core_web_sm")
doc = nlp(full_text)

# 建立句子到原行的映射
sentence_to_lines = {}
for sent_idx, sent in enumerate(doc.sents, 1):
    sent_start = sent.start_char
    sent_end = sent.end_char
    # 筛选所有与当前句子位置重叠的原行
    related_lines = [lp["line_number"] for lp in line_positions if not (lp["end"] <= sent_start or lp["start"] >= sent_end)]
    sentence_to_lines[f"句子{sent_idx}"] = tuple(related_lines)

print(sentence_to_lines)
# 输出结果:
# {'句子1': (1, 2), '句子2': (3, 4), '句子3': (5, 6)}

注意事项

  • 拼接文本时需保证与分句工具处理的文本完全一致,若原OCR行之间存在换行/空格,需在拼接时同步添加,否则字符位置会错位;
  • 该方法适用于所有分句场景,匹配精度高,是优先推荐的方案。

方法二:基于句子片段的逐步补全

核心思路是维护一个句子缓冲区,逐步拼接原OCR行,直到缓冲区内容被分句工具判定为完整句子,此时记录参与拼接的原行索引。

实现步骤

  1. 初始化缓冲区和当前行索引列表;
  2. 逐行添加到缓冲区,用分句工具检查缓冲区内容是否为完整句子;
  3. 若判定为完整句子,建立映射并重置缓冲区;
  4. 处理循环结束后剩余的未完成句子片段。

代码示例

import spacy

ocr_lines = [
    "Much of Singapore's infrastructure had been destroyed during the war",
    "including those needed to supply utilities.",
    "A shortage of food led to malnutrition, disease,",
    " and rampant crime and violence.",
    "A series of strikes in 1947 caused massive,",
    "stoppages in public transport and other services."
]

nlp = spacy.load("en_core_web_sm")
sentence_buffer = []
current_line_numbers = []
sentence_to_lines = {}
sentence_idx = 0

for line_idx, line in enumerate(ocr_lines):
    current_line_numbers.append(line_idx + 1)
    sentence_buffer.append(line)
    combined_text = "".join(sentence_buffer)
    
    # 用Spacy判断缓冲区是否为完整句子
    doc = nlp(combined_text)
    sents = list(doc.sents)
    if len(sents) == 1 and sents[0].end_char == len(combined_text):
        sentence_idx += 1
        sentence_to_lines[f"句子{sentence_idx}"] = tuple(current_line_numbers)
        # 重置缓冲区
        sentence_buffer = []
        current_line_numbers = []

# 处理剩余的未完成句子
if sentence_buffer:
    sentence_idx += 1
    sentence_to_lines[f"句子{sentence_idx}"] = tuple(current_line_numbers)

print(sentence_to_lines)

注意事项

  • 该方法适合处理原OCR行拆分不规则的场景;
  • 对于包含缩写(如Mr.)的句子,分句工具可能误判,需根据实际场景调整判断逻辑。

内容的提问来源于stack exchange,提问作者newbie101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 22:05:12