You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure函数中如何将Blob存储的PDF流转为可提取文本的PDF?

当然有可行的实现方法!你的PdfSharp扩展方法已经搭好了文本提取的核心逻辑,接下来只需要把Azure Blob里的PDF流转成PdfSharp能识别的PdfDocument对象就行,我给你一步步拆解具体实现:

实现步骤

1. 准备依赖包

首先确保你的Azure函数项目里安装了必要的NuGet包:

  • PdfSharpCore(如果是.NET Core/.NET 5+环境,优先用这个,是PdfSharp的.NET Core移植版,兼容性更好)
  • Azure.Storage.Blobs(用于操作Azure Blob存储)

可以通过NuGet包管理器或者命令行安装:

Install-Package PdfSharpCore
Install-Package Azure.Storage.Blobs

2. 处理Blob流并转换为PdfDocument

Azure Blob返回的流通常是不可查找的流(比如BlobDownloadInfo.Content),而PdfSharp加载PDF需要可查找的流,所以我们需要先把Blob流复制到MemoryStream里,再用来加载PDF。

下面是Blob触发的Azure函数示例代码:

using Azure.Storage.Blobs;
using PdfSharpCore.Pdf;
using PdfSharpCore.Pdf.Content;
using PdfSharpCore.Pdf.Content.Objects;
using System.IO;
using System.Linq;

public static class PdfTextExtractFunction
{
    [FunctionName("ExtractPdfTextFromBlob")]
    public static void Run(
        [BlobTrigger("pdf-container/{name}.pdf", Connection = "AzureWebJobsStorage")] Stream blobStream,
        string name,
        ILogger log)
    {
        log.LogInformation($"Processing PDF file: {name}.pdf");

        try
        {
            // 将Blob流复制到可查找的MemoryStream
            using var memoryStream = new MemoryStream();
            blobStream.CopyTo(memoryStream);
            memoryStream.Position = 0; // 重置流位置到开头

            // 加载PDF文档
            using var document = PdfDocument.Load(memoryStream);

            // 遍历所有页面提取文本
            var allText = document.Pages
                .SelectMany(page => page.ExtractText())
                .Aggregate((current, next) => current + " " + next);

            log.LogInformation($"Extracted text from {name}.pdf:\n{allText}");

            // 这里可以添加后续逻辑,比如把提取的文本保存到数据库、另一个Blob等
        }
        catch (Exception ex)
        {
            log.LogError($"Error processing {name}.pdf: {ex.Message}");
        }
    }
}

// 你的PdfSharp扩展方法(适配PdfSharpCore的命名空间)
public static class PdfSharpExtensions
{
    public static IEnumerable<string> ExtractText(this PdfPage page)
    {
        var content = ContentReader.ReadContent(page);
        var text = content.ExtractText();
        return text;
    }

    public static IEnumerable<string> ExtractText(this CObject cObject)
    {
        if (cObject is COperator cOperator)
        {
            if (cOperator.OpCode.Name == OpCodeName.Tj.ToString() || 
                cOperator.OpCode.Name == OpCodeName.TJ.ToString())
            {
                foreach (var cOperand in cOperator.Operands)
                    foreach (var txt in ExtractText(cOperand))
                        yield return txt;
            }
        }
        else if (cObject is CSequence cSequence)
        {
            foreach (var element in cSequence)
                foreach (var txt in ExtractText(element))
                    yield return txt;
        }
        else if (cObject is CString cString)
        {
            yield return cString.Value;
        }
    }
}

3. 关键注意事项

  • 流的可查找性:一定要把Blob流复制到MemoryStream,因为PdfSharp在加载PDF时需要对流进行多次读取/定位,原生Blob流不支持这些操作。
  • 依赖版本:如果你的Azure函数是基于.NET Core/.NET 6+,别用旧版的PdfSharp,优先选PdfSharpCore避免兼容性问题。
  • 异常处理:添加适当的异常捕获,处理PDF损坏、Blob权限不足等情况,保证函数的稳定性。
  • 资源释放:所有流和PdfDocument对象都要用using语句包裹,确保资源被正确释放,避免内存泄漏。

这样就能顺利把Azure Blob里的PDF流转成可处理的PdfSharp对象,并用你的扩展方法完成文本提取啦!

内容的提问来源于stack exchange,提问作者Xavi Andreu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:11:47