基于父属性匹配将节点移至前序父节点的XSL问题求助
Hey there! Let’s work through your XSLT challenges with that PDF-converted XML file—those auto-generated XMLs can be tricky with all their layout-specific nodes, so let’s break down the two issues you’re facing:
1. Fixing the "can’t copy desired nodes" problem
PDF-to-XML outputs (like from Acrobat or tools like pdf2xml) usually have a nested structure—think Page → TextBlock → TextLine → Word with attributes like bbox (bounding box) for positioning. Chances are your current XSL isn’t targeting the right nodes correctly. Here’s how to fix this:
- Start with an identity template to preserve the base XML structure, then add specific templates to copy only what you need:
<!-- Identity template: copies all nodes/attributes by default --> <xsl:template match="node()|@*"> <xsl:copy> <xsl:apply-templates select="node()|@*"/> </xsl:copy> </xsl:template> <!-- Example: Copy only <Word> nodes nested under <TextLine> into a custom element --> <xsl:template match="TextLine"> <CustomTextLine> <xsl:apply-templates select="Word"/> </CustomTextLine> </xsl:template> - If you need to target nodes with specific properties (like a certain
bboxor text content), add conditions to yourmatchorselectstatements:<!-- Copy only <Word> nodes with text containing "important" --> <xsl:template match="Word[contains(text(), 'important')]"> <xsl:copy-of select="."/> </xsl:template>
2. Excluding specific nodes
Excluding nodes is straightforward once you know their structure—you just create a template that matches the node and outputs nothing:
- To exclude an entire node type (e.g.,
<Metadata>or<Annotation>):<!-- This template matches <Metadata> and does nothing, effectively removing it --> <xsl:template match="Metadata"/> - To exclude child nodes under a specific parent (e.g., remove
<Footer>only inside<Page>):<xsl:template match="Page/Footer"/> - For duplicate nodes, use
<xsl:for-each-group>to group by a unique identifier (like@textor@bbox) and keep only the first instance:<!-- Group <Word> nodes by their text content, keep only the first in each group --> <xsl:for-each-group select="Word" group-by="@text"> <xsl:copy-of select="current-group()[1]"/> </xsl:for-each-group>
Quick tips for your geometric grouping attempt
Since you tried using text geometry (like bounding boxes) to group/move nodes, make sure you’re targeting the right attributes. Most PDF XMLs use bbox in the format x1 y1 x2 y2—you can group by the y-coordinate (second value) to cluster text from the same line:
<xsl:for-each-group select="Word" group-by="substring-before(substring-after(@bbox, ' '), ' ')"> <TextLineGroup> <xsl:apply-templates select="current-group()"/> </TextLineGroup> </xsl:for-each-group>
If you can share a small snippet of your XML input and your current XSL code, we can pinpoint exactly where things are going wrong—but these steps should get you past the two main blockers.
内容的提问来源于stack exchange,提问作者Phyllis

