TEI书籍评注词条提取:如何用XSLT获取anchor与note间文本
Solution: Extracting Commentary Entries from TEI with XSLT
To solve your problem of extracting the text between <anchor type='commentary'> and its corresponding <note type='commentary'> and formatting it into the desired entry format, here's a targeted XSLT solution that handles special content like math formulas:
Step-by-Step Explanation & Code
1. Core XSLT Stylesheet
This stylesheet assumes your TEI uses a predictable relationship between anchor and note IDs (e.g., A1 → N1). It also retrieves page numbers from preceding <pb> (page break) elements, which is standard in TEI.
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <!-- Output plain text entries with UTF-8 encoding --> <xsl:output method="text" encoding="UTF-8"/> <!-- Default template: skip all nodes unless explicitly matched --> <xsl:template match="node()|@*"> <xsl:apply-templates select="node()|@*"/> </xsl:template> <!-- Process commentary anchors to build entries --> <xsl:template match="anchor[@type='commentary']"> <!-- Link anchor to its corresponding note using ID pattern --> <xsl:variable name="anchor-id" select="@xml:id"/> <xsl:variable name="note-id" select="concat('N', substring-after($anchor-id, 'A'))"/> <xsl:variable name="target-note" select="//note[@type='commentary' and @xml:id=$note-id]"/> <!-- Get page number from the most recent preceding page break --> <xsl:variable name="page-num" select="preceding::pb[1]/@n"/> <!-- Capture and process content between anchor and note --> <xsl:variable name="entry-content"> <xsl:apply-templates select="following-sibling::node()[. << $target-note]" mode="content-handler"/> </xsl:variable> <!-- Format the final entry --> <xsl:text>p. </xsl:text> <xsl:value-of select="$page-num"/> <xsl:text> </xsl:text> <xsl:value-of select="$entry-content"/> <xsl:text>) </xsl:text> <xsl:value-of select="$target-note"/> <xsl:text> </xsl:text> <!-- New line for each entry --> </xsl:template> <!-- Skip commentary notes since we process them via the anchor --> <xsl:template match="note[@type='commentary']"/> <!-- Mode for processing content between anchor and note --> <xsl:template match="text()" mode="content-handler"> <xsl:value-of select="normalize-space(.)"/> <!-- Clean up whitespace --> </xsl:template> <!-- Handle math elements by converting to plain text --> <xsl:template match="math" mode="content-handler"> <xsl:apply-templates select="node()" mode="content-handler"/> </xsl:template> <xsl:template match="mi|mo|mn|mtext" mode="content-handler"> <xsl:value-of select="."/> <!-- Output math components as text --> </xsl:template> <!-- Handle other inline elements (e.g., <emph>, <quote>) by processing their content --> <xsl:template match="*" mode="content-handler"> <xsl:apply-templates select="node()" mode="content-handler"/> </xsl:template> </xsl:stylesheet>
2. Key Features
- Anchor-Note Linking: Uses the ID pattern (A→N) to pair anchors with their notes. If your TEI uses
@targeton anchors instead, replace the$note-idvariable withsubstring-after(@target, '#'). - Page Number Retrieval: Pulls the page number from the nearest preceding
<pb n="X">element. Adjust this if your TEI uses a different page marking system. - Special Content Handling: Converts math elements (like
<math>,<mi>,<mo>) to plain text, and processes other inline elements by extracting their content. - Whitespace Cleanup: Uses
normalize-space()to remove extra spaces from the extracted text.
3. Example Output
For your sample TEI snippet:
<p> <pb n="1"/> <anchor type='commentary' xml:id='A1'/>Romeo and Juliet<note type='commentary' xml:id='N1'>Some blabla</note> <pb n="2"/> <anchor type='commentary' xml:id='A2'/>The formula <math><mi>x</mi><mo>=</mo><mn>2</mn></math> is important<note type='commentary' xml:id='N2'>Explains the formula</note> </p>
The stylesheet will generate:
p. 1 Romeo and Juliet) Some blabla p. 2 The formula x=2 is important) Explains the formula
4. Adjustments for Your TEI
- If your anchor-note pairing uses a different ID scheme, modify the
$note-idcalculation. - If you need to preserve formatting for special elements (e.g., keep math as structured XML), adjust the
content-handlertemplates to copy elements instead of extracting text.
内容的提问来源于stack exchange,提问作者martinanton
相关产品推荐
相关产品推荐

