使用C#/Spire.PDF NuGet包转换PDF到TXT时移除左右边距问题
解决方案
问题出在你设置的RectangleF参数错误:new RectangleF(45, 0, 0, 0)的宽度和高度都是0,导致提取区域被压缩成一条线,自然会截断底部文本并丢失行。要通过Spire.Pdf直接移除开头空白,需要正确定义提取区域的范围,覆盖页面除左边距外的全部有效内容区域。
标准A4页面的默认尺寸(以96DPI计算)是794×1123像素,你可以根据实际边距调整提取区域的坐标,以下是修正后的代码:
PdfDocument doc = new PdfDocument(); doc.LoadFromFile(@"path"); var content = new List<string>(); foreach (PdfPageBase page in doc.Pages) { // 获取当前页面的实际尺寸 float pageWidth = page.Size.Width; float pageHeight = page.Size.Height; // 定义提取区域:左边距45,上下无限制,宽度为页面宽度减去左边距 RectangleF extractArea = new RectangleF(45, 0, pageWidth - 45, pageHeight); PdfTextExtractOptions options = new() { IsExtractAllText = true, IsShowHiddenText = true, ExtractArea = extractArea }; PdfTextExtractor textExtractor = new(page); string extractedText = textExtractor.ExtractText(options); content.Add(extractedText); } using (FileStream fs = new FileStream(@"outputFile.txt", FileMode.Create)) using (StreamWriter sw = new StreamWriter(fs)) { sw.Write(string.Join("\n", content)); }
关键说明:
RectangleF的参数顺序是(x, y, width, height):x是左偏移,y是上偏移,width是区域宽度,height是区域高度。- 如果你不确定左边距的具体数值,可以先提取完整页面文本,查看空白字符的长度,再对应调整
x的值;或者尝试使用page.Margins.Left获取PDF定义的左边距(部分PDF可能未设置该属性,需测试)。 - 代码中添加了
using语句自动释放文件流资源,避免资源泄漏。
内容的提问来源于stack exchange,提问作者mv1999
相关产品推荐
相关产品推荐

