You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求优化XSLT实现文本到XML的结构化转换

Alright, let's refine your XSLT 2.0 code to hit all three of your requirements perfectly. Here's the optimized solution with breakdowns of how each part works:

<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:xs="http://www.w3.org/2001/XMLSchema" exclude-result-prefixes="xs">
    <xsl:output indent="yes" method="xml"/>
    <xsl:strip-space elements="*"/>

    <xsl:param name="txt-encoding" as="xs:string" select="'iso-8859-1'"/>
    <xsl:param name="txt-uri" as="xs:string" select="'linktofile'"/>

    <!-- 字段映射:源字段名 → 目标XML元素名(按需扩展) -->
    <xsl:variable name="field-mappings" as="element()*">
        <map source="REF FOU" target="REF_VEN"/>
        <map source="GENCOD" target="EAN"/>
    </xsl:variable>

    <xsl:template match="/" name="text2xml">
        <xsl:variable name="txt" select="unparsed-text($txt-uri, $txt-encoding)"/>
        
        <!-- 拆分整个文本为单个订单块(匹配[ENTETE]到[FIN]的完整订单内容) -->
        <xsl:analyze-string select="$txt" regex="\[ENTETE\](.*?)(\[LIGNE\].*?)*\[FIN\]" flags="s">
            <xsl:matching-substring>
                <ORDER>
                    <!-- 处理ENTETE部分的键值对,生成顶层XML元素 -->
                    <xsl:analyze-string select="regex-group(1)" regex="(\S+(\s\S+)*)\s*=\s*([^=]+?)(?=\s+\S+\s*=|\s*$)">
                        <xsl:matching-substring>
                            <xsl:variable name="field-name" select="normalize-space(regex-group(1))"/>
                            <!-- 把字段名里的空格替换成下划线 -->
                            <xsl:variable name="element-name" select="replace($field-name, ' ', '_')"/>
                            <xsl:element name="{$element-name}">
                                <xsl:value-of select="normalize-space(regex-group(3))"/>
                            </xsl:element>
                        </xsl:matching-substring>
                    </xsl:analyze-string>

                    <!-- 处理所有LIGNE段落,转换为<LINE>子元素 -->
                    <xsl:analyze-string select="." regex="\[LIGNE\](.*?)(?=\[LIGNE\]|\[FIN\])" flags="s">
                        <xsl:matching-substring>
                            <LINE>
                                <xsl:analyze-string select="regex-group(1)" regex="(\S+(\s\S+)*)\s*=\s*([^=]+?)(?=\s+\S+\s*=|\s*$)">
                                    <xsl:matching-substring>
                                        <xsl:variable name="source-field" select="normalize-space(regex-group(1))"/>
                                        <!-- 优先使用字段映射,没有映射则用空格转下划线的名称 -->
                                        <xsl:variable name="target-element" select="$field-mappings[@source = $source-field]/@target"/>
                                        <xsl:variable name="element-name" select="if ($target-element) then $target-element else replace($source-field, ' ', '_')"/>
                                        <xsl:element name="{$element-name}">
                                            <!-- 清理REF CLI字段的引号和多余空格 -->
                                            <xsl:value-of select="normalize-space(replace(regex-group(3), '&quot;', ''))"/>
                                        </xsl:element>
                                    </xsl:matching-substring>
                                </xsl:analyze-string>
                            </LINE>
                        </xsl:matching-substring>
                    </xsl:analyze-string>
                </ORDER>
            </xsl:matching-substring>
        </xsl:analyze-string>
    </xsl:template>
</xsl:stylesheet>

How each requirement is fulfilled:

  1. Replace spaces in header fields with underscores

    • We use replace($field-name, ' ', '_') to swap out spaces for underscores in field names (e.g., CODE VEND becomes CODE_VEND).
    • For LIGNE-specific fields that need a full rename (like REF FOU → REF_VEN), we added a field-mappings variable to handle those custom mappings cleanly.
  2. Group header data by NO_CDE

    • First, we split the entire input text into individual order blocks using the regex \[ENTETE\](.*?)(\[LIGNE\].*?)*\[FIN\]. Each block represents one order, naturally grouped by its unique NO_CDE value.
    • Each order block gets its own <ORDER> element, containing all the matching header fields and LINE items.
  3. Convert multiple [LIGNE] sections to child elements

    • We use another regex \[LIGNE\](.*?)(?=\[LIGNE\]|\[FIN\]) to capture each standalone LIGNE block.
    • Each captured block is converted into a <LINE> element, with its internal fields processed the same way as header fields (plus cleaning up quotes from REF CLI values).

Extra tweaks:

  • Improved regex matching to avoid cross-field value capture, making the key-value parsing more reliable.
  • Added quote removal for REF CLI to match your target XML format exactly.
  • The field mapping variable is easy to extend if you need to rename more fields later.

内容的提问来源于stack exchange,提问作者Komagain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 20:59:04