You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TEI书籍评注词条提取:如何用XSLT获取anchor与note间文本

Solution: Extracting Commentary Entries from TEI with XSLT

To solve your problem of extracting the text between <anchor type='commentary'> and its corresponding <note type='commentary'> and formatting it into the desired entry format, here's a targeted XSLT solution that handles special content like math formulas:

Step-by-Step Explanation & Code

1. Core XSLT Stylesheet

This stylesheet assumes your TEI uses a predictable relationship between anchor and note IDs (e.g., A1 → N1). It also retrieves page numbers from preceding <pb> (page break) elements, which is standard in TEI.

<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <!-- Output plain text entries with UTF-8 encoding -->
  <xsl:output method="text" encoding="UTF-8"/>

  <!-- Default template: skip all nodes unless explicitly matched -->
  <xsl:template match="node()|@*">
    <xsl:apply-templates select="node()|@*"/>
  </xsl:template>

  <!-- Process commentary anchors to build entries -->
  <xsl:template match="anchor[@type='commentary']">
    <!-- Link anchor to its corresponding note using ID pattern -->
    <xsl:variable name="anchor-id" select="@xml:id"/>
    <xsl:variable name="note-id" select="concat('N', substring-after($anchor-id, 'A'))"/>
    <xsl:variable name="target-note" select="//note[@type='commentary' and @xml:id=$note-id]"/>

    <!-- Get page number from the most recent preceding page break -->
    <xsl:variable name="page-num" select="preceding::pb[1]/@n"/>

    <!-- Capture and process content between anchor and note -->
    <xsl:variable name="entry-content">
      <xsl:apply-templates select="following-sibling::node()[. << $target-note]" mode="content-handler"/>
    </xsl:variable>

    <!-- Format the final entry -->
    <xsl:text>p. </xsl:text>
    <xsl:value-of select="$page-num"/>
    <xsl:text> </xsl:text>
    <xsl:value-of select="$entry-content"/>
    <xsl:text>) </xsl:text>
    <xsl:value-of select="$target-note"/>
    <xsl:text>&#10;</xsl:text> <!-- New line for each entry -->
  </xsl:template>

  <!-- Skip commentary notes since we process them via the anchor -->
  <xsl:template match="note[@type='commentary']"/>

  <!-- Mode for processing content between anchor and note -->
  <xsl:template match="text()" mode="content-handler">
    <xsl:value-of select="normalize-space(.)"/> <!-- Clean up whitespace -->
  </xsl:template>

  <!-- Handle math elements by converting to plain text -->
  <xsl:template match="math" mode="content-handler">
    <xsl:apply-templates select="node()" mode="content-handler"/>
  </xsl:template>
  <xsl:template match="mi|mo|mn|mtext" mode="content-handler">
    <xsl:value-of select="."/> <!-- Output math components as text -->
  </xsl:template>

  <!-- Handle other inline elements (e.g., <emph>, <quote>) by processing their content -->
  <xsl:template match="*" mode="content-handler">
    <xsl:apply-templates select="node()" mode="content-handler"/>
  </xsl:template>
</xsl:stylesheet>

2. Key Features

  • Anchor-Note Linking: Uses the ID pattern (A→N) to pair anchors with their notes. If your TEI uses @target on anchors instead, replace the $note-id variable with substring-after(@target, '#').
  • Page Number Retrieval: Pulls the page number from the nearest preceding <pb n="X"> element. Adjust this if your TEI uses a different page marking system.
  • Special Content Handling: Converts math elements (like <math>, <mi>, <mo>) to plain text, and processes other inline elements by extracting their content.
  • Whitespace Cleanup: Uses normalize-space() to remove extra spaces from the extracted text.

3. Example Output

For your sample TEI snippet:

<p>
  <pb n="1"/>
  <anchor type='commentary' xml:id='A1'/>Romeo and Juliet<note type='commentary' xml:id='N1'>Some blabla</note>
  <pb n="2"/>
  <anchor type='commentary' xml:id='A2'/>The formula <math><mi>x</mi><mo>=</mo><mn>2</mn></math> is important<note type='commentary' xml:id='N2'>Explains the formula</note>
</p>

The stylesheet will generate:

p. 1 Romeo and Juliet) Some blabla
p. 2 The formula x=2 is important) Explains the formula

4. Adjustments for Your TEI

  • If your anchor-note pairing uses a different ID scheme, modify the $note-id calculation.
  • If you need to preserve formatting for special elements (e.g., keep math as structured XML), adjust the content-handler templates to copy elements instead of extracting text.

内容的提问来源于stack exchange,提问作者martinanton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:49:21