You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用iText提取PDF文本时欧元符号丢失的解决方法求助

解决iText提取PDF文本丢失欧元符号的问题

以下是几个可行的解决方案,按优先级尝试:

1. 启用完整字体提取模式

PDF中的欧元符号常因字体嵌入不完整、编码映射缺失导致默认提取逻辑无法识别,开启完整字体提取模式可让iText尝试解析更多字体字符映射信息。

修改代码,在创建PdfReader时添加字体提取配置:

public void TextExtraction()
{
    StringBuilder allTextBuilder = new StringBuilder();
    // 添加字体提取配置
    var readerProperties = new PdfReaderProperties();
    readerProperties.SetFontExtractionMode(PdfReader.FontExtractionMode.FULL);
    
    using (PdfReader pdfReader = new PdfReader(SourceFileName, readerProperties))
    using (PdfDocument pdfDocument = new PdfDocument(pdfReader))
    {
        for (int page = 1; page <= pdfDocument.GetNumberOfPages(); page++)
        {
            ITextExtractionStrategy strategy = new SimpleTextExtractionStrategy();
            string currentPageText = PdfTextExtractor.GetTextFromPage(pdfDocument.GetPage(page), strategy);
            allTextBuilder.Append(currentPageText); // 无需使用AppendFormat,直接追加文本即可
        }
    }

    File.WriteAllText(DestinationFileName, allTextBuilder.ToString(), Encoding.Unicode);
}

2. 替换文本提取策略

如果上述方法无效,可尝试使用LocationTextExtractionStrategy,它在处理复杂布局、特殊字符的文本时表现更稳定:

// 替换策略初始化代码
ITextExtractionStrategy strategy = new LocationTextExtractionStrategy();

3. 调整输出编码

虽然你已使用Encoding.Unicode,可尝试换成Encoding.UTF8写入文件,部分文本查看工具对UTF-8的特殊字符支持更友好:

File.WriteAllText(DestinationFileName, allTextBuilder.ToString(), Encoding.UTF8);

内容的提问来源于stack exchange,提问作者Ehsan Sharifi Esfahani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 20:09:10