正则表达式问题:获取文本行所属的最近XML标签内容
问题解决:获取文本行所属的最近XML标签完整内容
需求说明
给定如下XML内容:
<Artificial name="Artifical name"> <Machine> <MachineEnvironment uri="environment" /> </Machine> <Mobile>taken phone, test when r1 100m SUV then FireFly is High end when r2 Order of the Phonenix then Magic is High end </Mobile> </Artificial>
需要编写TypeScript函数getLineContent,接收一行字符串(line)和XML内容字符串(content),返回该行所属的最近标签的完整内容。例如传入行FireFly is High时,应返回<Mobile>及其内部所有内容到</Mobile>的完整片段。
当前实现问题
当前处理纯文本行的正则逻辑存在问题,代码如下:
if (isPlainTextLine) { const regex = new RegExp(`(<[^>]*>)([\\s\\S]*?${trimmedLine.split(' ')[0].substr(1)}[\\s\\S]*?</[a-zA-Z]+>)`) const match = content.match(regex) console.log('isPlainTextLine', match) if (match && match[1] && match[2]) { return match[2] } }
传入FireFly is High时,返回结果错误包含了<Machine>标签内容,原因是该正则会匹配文本开头的第一个标签,没有考虑标签的层级和最近匹配的逻辑。
正确实现方案
换用位置扫描+标签计数的方式,准确找到目标行所属的最近标签:
function getLineContent(line: string, content: string): string | null { // 定位目标行在XML内容中的起始位置 const normalizedLine = line.trimEnd(); const lineIndex = content.indexOf(normalizedLine); if (lineIndex === -1) return null; // 向前扫描找到最近的非自闭合开始标签 let startTagInfo: { index: number; name: string; length: number } | null = null; const startTagRegex = /<([a-zA-Z]+)(?:\s[^>]*?)?>/g; let match: RegExpMatchArray | null; // 遍历所有起始标签,筛选出目标行之前的最近有效标签 while ((match = startTagRegex.exec(content)) !== null) { if (match.index < lineIndex && !match[0].endsWith('/>')) { startTagInfo = { index: match.index, name: match[1], length: match[0].length }; } } if (!startTagInfo) return null; const { index: startIndex, name: tagName, length: tagLength } = startTagInfo; let scanPos = startIndex + tagLength; let tagCount = 1; const endTagRegex = new RegExp(`(</${tagName}>|<${tagName}(?:\\s[^>]*?)?>|<${tagName}(?:\\s[^>]*?)?/>)`, 'g'); // 向后扫描匹配对应的结束标签,处理嵌套标签的计数 while ((match = endTagRegex.exec(content)) !== null) { if (match.index < scanPos) continue; if (match[0].startsWith(`</${tagName}>`)) { tagCount--; if (tagCount === 0) { // 返回从开始标签到结束标签的完整内容 return content.slice(startIndex, match.index + match[0].length); } } else if (match[0].startsWith(`<${tagName}`) && !match[0].endsWith('/>')) { tagCount++; } scanPos = match.index + match[0].length; } return null; }
实现说明
- 定位目标行:先找到目标行在XML内容中的位置,确保能准确锁定范围。
- 寻找最近开始标签:从目标行位置向前遍历所有非自闭合的起始标签,记录最后一个(最近的)标签的位置和名称。
- 匹配对应结束标签:从开始标签之后的位置向后扫描,通过标签计数处理嵌套情况(遇到同名起始标签计数+1,遇到结束标签计数-1,计数归0时找到对应闭合标签)。
- 返回结果:截取从开始标签到对应结束标签的完整字符串。
测试传入FireFly is High和给定XML内容,将正确返回<Mobile>标签的完整内容。
内容的提问来源于stack exchange,提问作者CraZyDroiD
相关产品推荐
相关产品推荐

