使用XSL提取外部HTML生成XHTML、XSLT转换生成HTML的技术问询
Alright, let's walk through how to use XSLT to pull content from external HTML files and generate either XHTML or HTML from your DITA topic. I've broken this down into actionable steps with code examples to make it concrete:
First, we'll set up the core XSLT file that targets your DITA topic, handles namespaces, and defines the output format. This ensures we're working with the right context from the start.
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:ditaarch="http://dita.oasis-open.org/architecture/2005/" xmlns:r="http://www.Corecms.com/Core/ns/metadata" exclude-result-prefixes="ditaarch r"> <!-- Configure output: switch to method="html" if you don't need strict XHTML --> <xsl:output method="xhtml" version="1.0" encoding="UTF-8" indent="yes" doctype-public="-//W3C//DTD XHTML 1.0 Transitional//EN" doctype-system="http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd"/>
Next, we'll create a template that matches the root <topic> element of your DITA file. Here's where we'll load the external HTML using the document() function and inject its content into our output.
<!-- Match the root DITA topic --> <xsl:template match="/topic"> <html xmlns="http://www.w3.org/1999/xhtml"> <head> <title><xsl:value-of select="title"/></title> </head> <body> <!-- First, include content from your DITA topic --> <h1><xsl:value-of select="title"/></h1> <xsl:apply-templates select="body"/> <!-- Now load and extract content from external HTML --> <xsl:variable name="external-content" select="document('path/to/your/external.html')"/> <section class="external-html-content"> <h2>External Content</h2> <!-- Example: Extract all paragraphs from a specific div in the external HTML --> <xsl:copy-of select="$external-content//div[@class='target-section']/p"/> </section> </body> </html> </xsl:template>
To ensure all elements in your DITA topic's body are properly copied to the output, add these default templates. They'll preserve the structure of your original DITA content while we inject the external HTML.
<!-- Default template to copy DITA body elements --> <xsl:template match="body//*"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> <!-- Copy all attributes from DITA elements --> <xsl:template match="@*"> <xsl:copy/> </xsl:template> </xsl:stylesheet>
- Path to External HTML: Make sure the path in
document('path/to/your/external.html')is correct—use a relative path if the file is in the same directory as your XSLT, or an absolute path if it's elsewhere. - Well-Formed HTML: XSLT requires XML-compatible input, so your external HTML must be well-formed (properly closed tags, no unquoted attributes). If it's not, use a tool like HTML Tidy to convert it to valid XHTML first.
- XSLT Version: If you're using XSLT 1.0, the
document()function behaves slightly differently—double-check your processor's documentation for path handling. - Customize Extraction: Modify the XPath in
<xsl:copy-of select="..."/>to target exactly the content you need from the external HTML (e.g.,//h3,//ul[@id='nav'], etc.).
内容的提问来源于stack exchange,提问作者Aakash Kumar

