使用iText7和内存流在C#提取PDF表单字段值遇‘PDF头未找到’错误
解决iText7加载内存流PDF时“PDF header not found”错误
错误原因
你通过memoryStream.ReadAsync将流内容读取到buffer后,内存流的位置指针已经移动到了流的末尾。此时直接将该流传入PdfReader,它会从当前指针位置开始读取,无法获取到任何PDF数据,因此触发“PDF header not found”错误。
解决方案
方案一:重置内存流指针到起始位置
在将内存流传给PdfReader前,重置流的位置指针到开头:
using (var memoryStream = await oneDrive.DriveItemDownloadAsync(SharedDriveID, Path.Combine(fileDirectory, "testPDFForm.pdf"), "path")) { byte[] buffer = new byte[memoryStream.Length]; await memoryStream.ReadAsync(buffer, 0, (int)memoryStream.Length); document.Base64 = Convert.ToBase64String(buffer); document.DownloadLink = driveItem.AdditionalData["@microsoft.graph.downloadUrl"].ToString(); try { memoryStream.Position = 0; // 重置流指针到起始位置 using (PdfReader reader = new PdfReader(memoryStream)) using (PdfDocument pdfDocument = new PdfDocument(reader)) { PdfAcroForm acroForm = PdfAcroForm.GetAcroForm(pdfDocument, false); var allFormFields = acroForm.GetAllFormFieldsAndAnnotations(); // 处理表单字段逻辑 } } catch (Exception ex) { Ok(ex); } }
方案二:用已读取的byte数组创建新内存流
利用已经读取到buffer中的数据,创建一个新的内存流传入PdfReader,无需处理指针问题:
using (var memoryStream = await oneDrive.DriveItemDownloadAsync(SharedDriveID, Path.Combine(fileDirectory, "testPDFForm.pdf"), "path")) { byte[] buffer = new byte[memoryStream.Length]; await memoryStream.ReadAsync(buffer, 0, (int)memoryStream.Length); document.Base64 = Convert.ToBase64String(buffer); document.DownloadLink = driveItem.AdditionalData["@microsoft.graph.downloadUrl"].ToString(); try { using (var newStream = new MemoryStream(buffer)) using (PdfReader reader = new PdfReader(newStream)) using (PdfDocument pdfDocument = new PdfDocument(reader)) { PdfAcroForm acroForm = PdfAcroForm.GetAcroForm(pdfDocument, false); var allFormFields = acroForm.GetAllFormFieldsAndAnnotations(); // 处理表单字段逻辑 } } catch (Exception ex) { Ok(ex); } }
额外注意事项
- 始终用
using包裹PdfReader和PdfDocument,确保资源被正确释放,避免内存泄漏 - 若问题仍存在,可验证
buffer转换的Base64内容是否能正常解析为PDF,排查是否是OneDrive下载的文件本身不完整或损坏
内容的提问来源于stack exchange,提问作者Qiuzman
相关产品推荐
相关产品推荐

