You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于父属性匹配将节点移至前序父节点的XSL问题求助

Hey there! Let’s work through your XSLT challenges with that PDF-converted XML file—those auto-generated XMLs can be tricky with all their layout-specific nodes, so let’s break down the two issues you’re facing:

1. Fixing the "can’t copy desired nodes" problem

PDF-to-XML outputs (like from Acrobat or tools like pdf2xml) usually have a nested structure—think Page → TextBlock → TextLine → Word with attributes like bbox (bounding box) for positioning. Chances are your current XSL isn’t targeting the right nodes correctly. Here’s how to fix this:

  • Start with an identity template to preserve the base XML structure, then add specific templates to copy only what you need:
    <!-- Identity template: copies all nodes/attributes by default -->
    <xsl:template match="node()|@*">
      <xsl:copy>
        <xsl:apply-templates select="node()|@*"/>
      </xsl:copy>
    </xsl:template>
    
    <!-- Example: Copy only <Word> nodes nested under <TextLine> into a custom element -->
    <xsl:template match="TextLine">
      <CustomTextLine>
        <xsl:apply-templates select="Word"/>
      </CustomTextLine>
    </xsl:template>
    
  • If you need to target nodes with specific properties (like a certain bbox or text content), add conditions to your match or select statements:
    <!-- Copy only <Word> nodes with text containing "important" -->
    <xsl:template match="Word[contains(text(), 'important')]">
      <xsl:copy-of select="."/>
    </xsl:template>
    

2. Excluding specific nodes

Excluding nodes is straightforward once you know their structure—you just create a template that matches the node and outputs nothing:

  • To exclude an entire node type (e.g., <Metadata> or <Annotation>):
    <!-- This template matches <Metadata> and does nothing, effectively removing it -->
    <xsl:template match="Metadata"/>
    
  • To exclude child nodes under a specific parent (e.g., remove <Footer> only inside <Page>):
    <xsl:template match="Page/Footer"/>
    
  • For duplicate nodes, use <xsl:for-each-group> to group by a unique identifier (like @text or @bbox) and keep only the first instance:
    <!-- Group <Word> nodes by their text content, keep only the first in each group -->
    <xsl:for-each-group select="Word" group-by="@text">
      <xsl:copy-of select="current-group()[1]"/>
    </xsl:for-each-group>
    

Quick tips for your geometric grouping attempt

Since you tried using text geometry (like bounding boxes) to group/move nodes, make sure you’re targeting the right attributes. Most PDF XMLs use bbox in the format x1 y1 x2 y2—you can group by the y-coordinate (second value) to cluster text from the same line:

<xsl:for-each-group select="Word" group-by="substring-before(substring-after(@bbox, ' '), ' ')">
  <TextLineGroup>
    <xsl:apply-templates select="current-group()"/>
  </TextLineGroup>
</xsl:for-each-group>

If you can share a small snippet of your XML input and your current XSL code, we can pinpoint exactly where things are going wrong—but these steps should get you past the two main blockers.

内容的提问来源于stack exchange,提问作者Phyllis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:20:48