You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Tesseract.js返回空字符串/Null而非噪声OCR结果?

解决Tesseract.js识别普通照片返回噪声文本的问题

我需要处理PDF和普通照片,判断文件是否包含OCR可识别的文本,无文本时用LLM分析图像。但当前用Tesseract.js识别普通照片时会返回大量噪声文本(示例如下),希望返回空字符串或null替代噪声。

a i § V5
i LIS Ia ~ as # L

go “ry ‘

2 4 ray .
3 = say
Fo a Fe. FTE
A pS 4
a 2

  • 3 = — ¢
    eT TY
    Ey
    @ ir
    . -
    —-
    d
    |
    pv 4
    = TT)
    EFTA00000003

现有代码如上,且部分文档是图片格式、部分照片为PDF,无法单纯按扩展名过滤。


解决方案

1. 增强图像预处理

普通照片的噪声多来自复杂背景,通过更严格的图像预处理可以减少Tesseract误识别:

  • 阈值处理:将图像转为黑白,强化文本与背景边界
  • 中值滤波:消除微小干扰点
  • 对比度提升:让文本更清晰

修改图像预处理逻辑,替换原有sharp处理代码:

const processedBuffer = await sharp(buffer)
  .grayscale()
  .normalize()
  .threshold(128) // 黑白阈值,可根据实际调整
  .median(2) // 中值滤波去噪
  .linear(1.5, 0) // 提升对比度
  .toBuffer();

2. 基于Tesseract置信度过滤噪声

Tesseract会为每个识别出的单词返回置信度(confidence),过滤低置信度内容可直接剔除大部分噪声:

修改ocrImage函数,利用置信度筛选有效文本:

async function ocrImage(buffer) {
  const worker = await createWorker("eng");
  const { data } = await worker.recognize(buffer);
  await worker.terminate();

  const confidenceThreshold = 60; // 置信度阈值,可调整
  const validWords = data.words.filter(word => word.confidence >= confidenceThreshold);
  const text = validWords.map(word => word.text).join(" ").trim();

  return text.length ? text : "";
}

3. 添加文本质量校验

即使经过上述处理,仍可能残留少量噪声,可通过以下规则进一步过滤:

  • 统计有效字符(字母、数字)比例,低于阈值判定为噪声
  • 检查是否包含长度≥3的有效单词

添加文本校验辅助函数:

function isTextValid(text) {
  if (!text) return false;
  
  // 计算有效字符比例
  const validChars = text.match(/[a-zA-Z0-9]/g) || [];
  const validRatio = validChars.length / text.length;
  if (validRatio < 0.5) return false;
  
  // 检查是否有有效长单词
  const hasValidWord = text.split(/\s+/).some(word => /[a-zA-Z0-9]{3,}/.test(word));
  return hasValidWord;
}

在ocrImage函数末尾调用校验逻辑:

async function ocrImage(buffer) {
  // ... 原有置信度过滤逻辑 ...
  
  const text = validWords.map(word => word.text).join(" ").trim();
  
  return isTextValid(text) ? text : "";
}

整合后的完整代码

const fs = require("fs");
const fsp = fs.promises;
const path = require("path");
const sharp = require("sharp");
const { createWorker } = require("tesseract.js");
const { fromPath } = require("pdf2pic");
const os = require("os");

function isTextValid(text) {
  if (!text) return false;
  
  const validChars = text.match(/[a-zA-Z0-9]/g) || [];
  const validRatio = validChars.length / text.length;
  
  if (validRatio < 0.5) return false;
  
  const hasValidWord = text.split(/\s+/).some(word => /[a-zA-Z0-9]{3,}/.test(word));
  return hasValidWord;
}

async function ocrImage(buffer) {
  const worker = await createWorker("eng");
  const { data } = await worker.recognize(buffer);
  await worker.terminate();

  const confidenceThreshold = 60;
  const validWords = data.words.filter(word => word.confidence >= confidenceThreshold);
  const text = validWords.map(word => word.text).join(" ").trim();

  return isTextValid(text) ? text : "";
}

async function extractText(filePath) {
  const ext = path.extname(filePath).toLowerCase();
  console.log("filepath: ", filePath);

  if (ext === ".pdf") {
    const converter = fromPath(filePath, {
      density: 200,
      format: "png",
      width: 1654,
      height: 2339,
      popplerPath: "/usr/local/bin"
    });

    const pages = await converter.bulk(-1, { responseType: "base64" });

    let fullText = "";
    for (const page of pages) {
      console.log("page: ", page);
      const imgBuffer = Buffer.from(page.base64, "base64");

      const processedBuffer = await sharp(imgBuffer)
        .grayscale()
        .normalize()
        .threshold(128)
        .median(2)
        .linear(1.5, 0)
        .toBuffer();

      const text = await ocrImage(processedBuffer);
      if (text) fullText += text + "\n";
    }
    return fullText.trim();
  }

  const supported = new Set([
    ".png", ".jpg", ".jpeg", ".bmp", ".tiff", ".webp", ".gif"
  ]);

  if (!supported.has(ext)) {
    return ""
  }

  const processedBuffer = await sharp(filePath)
    .grayscale()
    .normalize()
    .threshold(128)
    .median(2)
    .linear(1.5, 0)
    .toBuffer();
  return await ocrImage(processedBuffer);
}

module.exports = { extractText }

注意事项

  • 所有阈值(置信度、图像阈值、有效字符比例)需根据实际场景调整,建议测试不同类型照片后优化参数
  • PDF中的照片页面会自动转为图片处理,上述逻辑同样适用

内容的提问来源于stack exchange,提问作者Logan Besecker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 14:04:52