如何使用XSLT同时拆分多字段字符串
关联字段拆分映射的解决思路
针对你这种需要同时拆分两个关联字段并一一映射的场景,这里提供两种实用方案:
方案一:使用XSLT 2.0+(适合XML原生处理场景)
利用XSLT的tokenize()函数分别拆分两个字段的文本,再通过索引对应关联,最终输出JSON:
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="text" encoding="UTF-8"/> <xsl:template match="/"> <xsl:variable name="codes" select="tokenize(somefield/code, ',')"/> <xsl:variable name="descriptions" select="tokenize(somefield/description, ';')"/> <xsl:text>{"somefield": [</xsl:text> <xsl:for-each select="$codes"> <xsl:variable name="pos" select="position()"/> <xsl:text>{"code": "</xsl:text> <xsl:value-of select="normalize-space(.)"/> <xsl:text>", "description": "</xsl:text> <xsl:value-of select="normalize-space($descriptions[$pos])"/> <xsl:text>"}</xsl:text> <xsl:if test="$pos != last()">,</xsl:if> </xsl:for-each> <xsl:text>]}</xsl:text> </xsl:template> </xsl:stylesheet>
- 核心逻辑:用
tokenize()拆分字符串得到两个序列,通过position()获取当前索引,匹配对应位置的code和description; normalize-space()用来清理字符串前后的多余空格(比如description拆分后可能带的前置空格)。
方案二:使用Python(适合脚本化处理场景)
通过解析XML提取字段,拆分后用zip()配对生成目标结构:
import xml.etree.ElementTree as ET import json # 解析XML xml_content = ''' <somefield> <code>A,B,C</code> <description>Description of A; Description of B; Description of C</description> </somefield> ''' root = ET.fromstring(xml_content) # 提取并拆分字段 codes = root.find('code').text.strip().split(',') # 拆分description时处理分号+空格的分隔符,同时清理每个项的空格 descriptions = [desc.strip() for desc in root.find('description').text.strip().split(';')] # 配对生成列表 somefield_list = [{"code": code, "description": desc} for code, desc in zip(codes, descriptions)] # 生成最终JSON result = {"somefield": somefield_list} print(json.dumps(result, indent=2))
- 核心逻辑:拆分两个字段后,用
zip()将对应位置的元素配对,推导式生成字典列表; - 用列表推导式配合
strip()清理每个description项的前后空格,保证输出格式整洁。
内容的提问来源于stack exchange,提问作者DR - Idemia
相关产品推荐
相关产品推荐

