You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用C#/Spire.PDF NuGet包转换PDF到TXT时移除左右边距问题

解决方案

问题出在你设置的RectangleF参数错误:new RectangleF(45, 0, 0, 0)的宽度和高度都是0,导致提取区域被压缩成一条线,自然会截断底部文本并丢失行。要通过Spire.Pdf直接移除开头空白,需要正确定义提取区域的范围,覆盖页面除左边距外的全部有效内容区域。

标准A4页面的默认尺寸(以96DPI计算)是794×1123像素,你可以根据实际边距调整提取区域的坐标,以下是修正后的代码:

PdfDocument doc = new PdfDocument();
doc.LoadFromFile(@"path");

var content = new List<string>();

foreach (PdfPageBase page in doc.Pages)
{
    // 获取当前页面的实际尺寸
    float pageWidth = page.Size.Width;
    float pageHeight = page.Size.Height;
    // 定义提取区域:左边距45,上下无限制,宽度为页面宽度减去左边距
    RectangleF extractArea = new RectangleF(45, 0, pageWidth - 45, pageHeight);
    
    PdfTextExtractOptions options = new() 
    { 
        IsExtractAllText = true, 
        IsShowHiddenText = true, 
        ExtractArea = extractArea 
    };

    PdfTextExtractor textExtractor = new(page);
    string extractedText = textExtractor.ExtractText(options);
    content.Add(extractedText);
}

using (FileStream fs = new FileStream(@"outputFile.txt", FileMode.Create))
using (StreamWriter sw = new StreamWriter(fs))
{
    sw.Write(string.Join("\n", content));
}

关键说明:

  • RectangleF的参数顺序是(x, y, width, height):x是左偏移,y是上偏移,width是区域宽度,height是区域高度。
  • 如果你不确定左边距的具体数值,可以先提取完整页面文本,查看空白字符的长度,再对应调整x的值;或者尝试使用page.Margins.Left获取PDF定义的左边距(部分PDF可能未设置该属性,需测试)。
  • 代码中添加了using语句自动释放文件流资源,避免资源泄漏。

内容的提问来源于stack exchange,提问作者mv1999

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 14:29:09