请求优化XSLT实现文本到XML的结构化转换
Alright, let's refine your XSLT 2.0 code to hit all three of your requirements perfectly. Here's the optimized solution with breakdowns of how each part works:
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:xs="http://www.w3.org/2001/XMLSchema" exclude-result-prefixes="xs"> <xsl:output indent="yes" method="xml"/> <xsl:strip-space elements="*"/> <xsl:param name="txt-encoding" as="xs:string" select="'iso-8859-1'"/> <xsl:param name="txt-uri" as="xs:string" select="'linktofile'"/> <!-- 字段映射:源字段名 → 目标XML元素名(按需扩展) --> <xsl:variable name="field-mappings" as="element()*"> <map source="REF FOU" target="REF_VEN"/> <map source="GENCOD" target="EAN"/> </xsl:variable> <xsl:template match="/" name="text2xml"> <xsl:variable name="txt" select="unparsed-text($txt-uri, $txt-encoding)"/> <!-- 拆分整个文本为单个订单块(匹配[ENTETE]到[FIN]的完整订单内容) --> <xsl:analyze-string select="$txt" regex="\[ENTETE\](.*?)(\[LIGNE\].*?)*\[FIN\]" flags="s"> <xsl:matching-substring> <ORDER> <!-- 处理ENTETE部分的键值对,生成顶层XML元素 --> <xsl:analyze-string select="regex-group(1)" regex="(\S+(\s\S+)*)\s*=\s*([^=]+?)(?=\s+\S+\s*=|\s*$)"> <xsl:matching-substring> <xsl:variable name="field-name" select="normalize-space(regex-group(1))"/> <!-- 把字段名里的空格替换成下划线 --> <xsl:variable name="element-name" select="replace($field-name, ' ', '_')"/> <xsl:element name="{$element-name}"> <xsl:value-of select="normalize-space(regex-group(3))"/> </xsl:element> </xsl:matching-substring> </xsl:analyze-string> <!-- 处理所有LIGNE段落,转换为<LINE>子元素 --> <xsl:analyze-string select="." regex="\[LIGNE\](.*?)(?=\[LIGNE\]|\[FIN\])" flags="s"> <xsl:matching-substring> <LINE> <xsl:analyze-string select="regex-group(1)" regex="(\S+(\s\S+)*)\s*=\s*([^=]+?)(?=\s+\S+\s*=|\s*$)"> <xsl:matching-substring> <xsl:variable name="source-field" select="normalize-space(regex-group(1))"/> <!-- 优先使用字段映射,没有映射则用空格转下划线的名称 --> <xsl:variable name="target-element" select="$field-mappings[@source = $source-field]/@target"/> <xsl:variable name="element-name" select="if ($target-element) then $target-element else replace($source-field, ' ', '_')"/> <xsl:element name="{$element-name}"> <!-- 清理REF CLI字段的引号和多余空格 --> <xsl:value-of select="normalize-space(replace(regex-group(3), '"', ''))"/> </xsl:element> </xsl:matching-substring> </xsl:analyze-string> </LINE> </xsl:matching-substring> </xsl:analyze-string> </ORDER> </xsl:matching-substring> </xsl:analyze-string> </xsl:template> </xsl:stylesheet>
How each requirement is fulfilled:
Replace spaces in header fields with underscores
- We use
replace($field-name, ' ', '_')to swap out spaces for underscores in field names (e.g.,CODE VENDbecomesCODE_VEND). - For LIGNE-specific fields that need a full rename (like
REF FOU→REF_VEN), we added afield-mappingsvariable to handle those custom mappings cleanly.
- We use
Group header data by NO_CDE
- First, we split the entire input text into individual order blocks using the regex
\[ENTETE\](.*?)(\[LIGNE\].*?)*\[FIN\]. Each block represents one order, naturally grouped by its uniqueNO_CDEvalue. - Each order block gets its own
<ORDER>element, containing all the matching header fields and LINE items.
- First, we split the entire input text into individual order blocks using the regex
Convert multiple [LIGNE] sections to
child elements - We use another regex
\[LIGNE\](.*?)(?=\[LIGNE\]|\[FIN\])to capture each standalone LIGNE block. - Each captured block is converted into a
<LINE>element, with its internal fields processed the same way as header fields (plus cleaning up quotes fromREF CLIvalues).
- We use another regex
Extra tweaks:
- Improved regex matching to avoid cross-field value capture, making the key-value parsing more reliable.
- Added quote removal for
REF CLIto match your target XML format exactly. - The field mapping variable is easy to extend if you need to rename more fields later.
内容的提问来源于stack exchange,提问作者Komagain
相关产品推荐
相关产品推荐

