You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

希伯来语PDF解析乱码求助:未知编码致符号输出

解决希伯来语PDF解析乱码问题

问题根源

用「打印另存为」去除PDF密码的操作,大概率破坏了原文档的字符编码映射——尤其是希伯来语这种从右到左的语言,pdf2json这类工具没法正确识别被篡改后的内嵌字体或编码规则,才会出现乱码。

Node.js 解决方案

1. 优化pdf2json的编码处理

直接修改解析逻辑,尝试多种希伯来语常用编码解码:

const PDFParser = require("pdf2json");
const pdfParser = new PDFParser(this, 1);

pdfParser.on("pdfParser_dataError", errData => console.error(errData.parserError));
pdfParser.on("pdfParser_dataReady", pdfData => {
  const rawText = pdfData.Pages.reduce((acc, page) => {
    return acc + page.Texts.map(text => {
      // 先解码默认的URI转义
      let decoded = decodeURIComponent(text.R[0].T);
      // 乱码的话,试试希伯来语常用的Windows-1255编码
      try {
        decoded = new TextDecoder("windows-1255").decode(Buffer.from(text.R[0].T, 'hex'));
      } catch(e) {}
      return decoded;
    }).join(' ');
  }, '');
  console.log(rawText);
});

pdfParser.loadPDF("你的解密后PDF文件.pdf");

2. 换用pdf-parse配合iconv-lite

如果pdf2json不好使,试试这个组合,支持更多编码转换:

const fs = require('fs');
const pdfParse = require('pdf-parse');
const iconv = require('iconv-lite');

fs.readFile('你的解密后PDF文件.pdf', (err, data) => {
  pdfParse(data).then(result => {
    // 挨个试希伯来语常用编码
    const possibleEncodings = ['utf8', 'windows-1255', 'iso-8859-8'];
    possibleEncodings.forEach(enc => {
      try {
        const decoded = iconv.decode(Buffer.from(result.text), enc);
        console.log(`编码${enc}解析结果:\n`, decoded);
      } catch(e) {}
    });
  });
});

Python 解决方案(更适配复杂编码)

Python的编码处理工具更成熟,推荐用pdfplumber配合chardet自动检测编码:

import pdfplumber
import chardet

with pdfplumber.open("你的解密后PDF文件.pdf") as pdf:
    for page in pdf.pages:
        raw_text = page.extract_text()
        # 自动检测文本编码
        encoding_result = chardet.detect(raw_text.encode('latin1'))
        detected_enc = encoding_result['encoding']
        
        if detected_enc:
            fixed_text = raw_text.encode('latin1').decode(detected_enc)
            print(fixed_text)
        else:
            # 检测失败就手动试希伯来语编码
            for enc in ['windows-1255', 'iso-8859-8']:
                try:
                    fixed_text = raw_text.encode('latin1').decode(enc)
                    print(fixed_text)
                    break
                except UnicodeDecodeError:
                    continue

极端情况:字体完全损坏

如果打印后的PDF已经把文本转成了图像,直接用OCR工具处理:

# 先安装OCRmyPDF
pip install ocrmypdf
# 对PDF做希伯来语OCR,生成可解析的新PDF
ocrmypdf 你的解密后PDF文件.pdf ocr处理后的文件.pdf -l heb

之后再用上面的工具解析生成的新PDF即可。

重要提示

别再用「打印另存为」解密PDF了,用专业工具qpdf更稳妥,不会破坏原文档结构:

qpdf --password=你的PDF密码 --decrypt 加密的原文件.pdf 解密后的文件.pdf

内容的提问来源于stack exchange,提问作者Yoni Kohn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:43:39