You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言OCR文本清洗:合并断行并生成句子级向量/列表

OCR文本清洗与句子拆分优化(R语言)

问题分析

当前OCR识别结果存在两个核心问题:

  • 完整句子被不合理拆分为多行
  • 文本中存在多余空格干扰

以下是针对性的优化方案:

优化步骤

1. 增强图片预处理(提升OCR准确率)

在现有预处理流程中加入阈值处理,进一步强化文本与背景的对比度,减少识别误差:

image_greek <- image_greek %>% 
  image_scale("600") %>% 
  image_crop("600x400+220+150") %>% 
  image_convert(type = 'Grayscale') %>% 
  image_contrast(sharpen = 1) %>% 
  image_threshold(type = "white", threshold = "40%") %>% # 新增阈值处理,突出文本
  image_write(format="jpg")

2. OCR后文本清洗与句子拆分

先合并所有换行内容,清理多余空格,再按句子标点(句号、问号、感叹号)拆分,得到每个元素对应一个完整句子的向量:

heraclitus_sentences <- magick::image_read(image_greek) %>% 
  ocr() %>% 
  str_replace_all("\n", " ") %>% # 将换行符替换为空格,合并断行内容
  str_squish() %>% # 清理所有多余空格(合并连续空格、去除首尾空格)
  str_split("(?<=[.!?])\\s+") # 按句子结束符后的空格拆分,保留标点在句尾

完整优化代码

heraclitus <- "greek.png"
library(tidyverse)
library(tesseract)
library(magick)

# 图片预处理
image_greek <- image_read(heraclitus) %>% 
  image_scale("600") %>% 
  image_crop("600x400+220+150") %>% 
  image_convert(type = 'Grayscale') %>% 
  image_contrast(sharpen = 1) %>% 
  image_threshold(type = "white", threshold = "40%") %>% 
  image_write(format="jpg")

# OCR识别+文本清洗+句子拆分
heraclitus_sentences <- magick::image_read(image_greek) %>% 
  ocr() %>% 
  str_replace_all("\n", " ") %>% 
  str_squish() %>% 
  str_split("(?<=[.!?])\\s+") %>% 
  pluck(1) # 将拆分后的列表转为向量,方便后续处理

关键说明

  • str_squish() 自动处理所有冗余空格,包括单词间的连续空格、文本首尾空格
  • str_split("(?<=[.!?])\\s+") 使用正向零宽断言,确保拆分后每个句子保留结尾标点,且仅在句子结束后的空格处拆分
  • pluck(1) 去除外层列表结构,直接得到每个元素为完整句子的向量

内容的提问来源于stack exchange,提问作者zachi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 03:43:13