You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Saxon扩展的Apache FOP IF转独立HTML:图片/PDF转Base64问题

解决Apache FOP IF转HTML时的图片/PDF嵌入及Saxon扩展函数报错问题

一、修复Saxon扩展函数找不到的错误

1. 正确实现Saxon HE扩展函数

Saxon 12 HE需通过ExtensionFunctionDefinition接口实现自定义扩展函数,以下是适配需求的完整实现:

package com.example.saxon.ext;

import net.sf.saxon.expr.XPathContext;
import net.sf.saxon.lib.ExtensionFunctionCall;
import net.sf.saxon.lib.ExtensionFunctionDefinition;
import net.sf.saxon.om.Sequence;
import net.sf.saxon.om.StructuredQName;
import net.sf.saxon.trans.XPathException;
import net.sf.saxon.value.StringValue;

import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Base64;

public class ImageToBase64Extension extends ExtensionFunctionDefinition {

    // 对应XSLT中使用的命名空间与函数名
    private static final StructuredQName FUNCTION_NAME =
            new StructuredQName("ext", "http://example.com/saxon-extension", "imageToBase64");

    @Override
    public StructuredQName getFunctionQName() {
        return FUNCTION_NAME;
    }

    @Override
    public int getMinimumNumberOfArguments() {
        return 1;
    }

    @Override
    public int getMaximumNumberOfArguments() {
        return 1;
    }

    @Override
    public ExtensionFunctionCall makeCallExpression() {
        return new ExtensionFunctionCall() {
            @Override
            public Sequence call(XPathContext context, Sequence[] arguments) throws XPathException {
                try {
                    // 获取传入的文件路径参数
                    String filePath = arguments[0].head().getStringValue();
                    // 读取文件字节流
                    byte[] fileBytes = Files.readAllBytes(Paths.get(filePath));
                    // 编码为Base64字符串
                    String base64 = Base64.getEncoder().encodeToString(fileBytes);
                    // 生成带MIME类型的Data URI
                    String mimeType = resolveMimeType(filePath);
                    return StringValue.makeStringValue("data:" + mimeType + ";base64," + base64);
                } catch (Exception e) {
                    throw new XPathException("文件转Base64失败: " + e.getMessage(), e);
                }
            }
        };
    }

    // 解析文件MIME类型,可按需扩展支持更多格式
    private String resolveMimeType(String filePath) {
        String lowerPath = filePath.toLowerCase();
        if (lowerPath.endsWith(".png")) return "image/png";
        if (lowerPath.endsWith(".jpg") || lowerPath.endsWith(".jpeg")) return "image/jpeg";
        if (lowerPath.endsWith(".pdf")) return "application/pdf";
        return "application/octet-stream";
    }
}

2. 在XSLT中声明并调用扩展函数

确保XSLT根元素绑定扩展命名空间,并针对IF文件的<image>元素编写转换逻辑:

<xsl:stylesheet version="3.0"
                xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
                xmlns:ext="http://example.com/saxon-extension"
                exclude-result-prefixes="ext">

    <!-- 处理IF文件中的<image>节点 -->
    <xsl:template match="image">
        <xsl:variable name="sourcePath" select="@src"/>
        <xsl:variable name="embeddedData" select="ext:imageToBase64($sourcePath)"/>
        
        <!-- 区分图片与PDF,生成对应HTML标签 -->
        <xsl:choose>
            <xsl:when test="lower-case(substring-after($sourcePath, '.')) = 'pdf'">
                <embed src="{$embeddedData}" type="application/pdf" width="100%" height="600px"/>
            </xsl:when>
            <xsl:otherwise>
                <img src="{$embeddedData}" alt="Embedded media"/>
            </xsl:otherwise>
        </xsl:choose>
    </xsl:template>

</xsl:stylesheet>

3. 在转换代码中注册扩展函数

使用Saxon Processor完成扩展函数注册与转换流程:

import net.sf.saxon.Processor;
import net.sf.saxon.s9api.XsltCompiler;
import net.sf.saxon.s9api.XsltExecutable;
import net.sf.saxon.s9api.XsltTransformer;
import com.example.saxon.ext.ImageToBase64Extension;

import java.io.File;
import java.io.FileOutputStream;

public class IfToHtmlConverter {
    public static void main(String[] args) throws Exception {
        Processor processor = new Processor(false);
        // 注册自定义扩展函数
        processor.getUnderlyingConfiguration().registerExtensionFunction(new ImageToBase64Extension());
        
        XsltCompiler compiler = processor.newXsltCompiler();
        XsltExecutable executable = compiler.compile(new File("your-stylesheet.xsl"));
        
        XsltTransformer transformer = executable.load();
        transformer.setSource(new File("input.if"));
        transformer.setDestination(new FileOutputStream("output.html"));
        
        transformer.transform();
    }
}

报错排查要点

  • 命名空间一致性:XSLT中声明的命名空间http://example.com/saxon-extension必须与Java类中StructuredQName的命名空间完全匹配。
  • 类路径配置:确保包含扩展类的JAR或class文件在运行时类路径中。
  • 参数数量匹配:XSLT调用函数时传入的参数个数,需与Java实现中getMinimumNumberOfArguments、getMaximumNumberOfArguments定义一致。

二、实现图片/PDF的Base64嵌入

上述扩展函数已完成核心逻辑:

  1. 读取外部图片/PDF文件的字节流。
  2. 将字节流编码为Base64字符串。
  3. 生成带MIME类型的Data URI,直接嵌入HTML标签,无需依赖外部文件,实现HTML独立运行。

优化建议

  • 增强MIME类型识别:可替换resolveMimeType方法为Files.probeContentType(Paths.get(filePath)),实现更准确的类型判断(需依赖系统文件类型支持)。
  • 异常容错处理:添加文件不存在、权限不足等场景的分支,返回默认占位资源的Base64,避免转换流程中断。

内容的提问来源于stack exchange,提问作者CrazyEight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 07:38:22