You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式问题:获取文本行所属的最近XML标签内容

问题解决:获取文本行所属的最近XML标签完整内容

需求说明

给定如下XML内容:

<Artificial name="Artifical name">
    <Machine>
        <MachineEnvironment uri="environment" />
    </Machine>
    <Mobile>taken phone, test

when r1
    100m SUV
then
    FireFly is High
end


when r2
    Order of the Phonenix 
    
then
    Magic is High
end


</Mobile>
</Artificial>

需要编写TypeScript函数getLineContent,接收一行字符串(line)和XML内容字符串(content),返回该行所属的最近标签的完整内容。例如传入行FireFly is High时,应返回<Mobile>及其内部所有内容到</Mobile>的完整片段。

当前实现问题

当前处理纯文本行的正则逻辑存在问题,代码如下:

if (isPlainTextLine) {
    const regex = new RegExp(`(&lt;[^&gt;]*&gt;)([\\s\\S]*?${trimmedLine.split(' ')[0].substr(1)}[\\s\\S]*?&lt;/[a-zA-Z]+&gt;)`)
    const match = content.match(regex)
    console.log('isPlainTextLine', match)
    if (match && match[1] && match[2]) {
        return match[2]
    }
}

传入FireFly is High时,返回结果错误包含了<Machine>标签内容,原因是该正则会匹配文本开头的第一个标签,没有考虑标签的层级和最近匹配的逻辑。

正确实现方案

换用位置扫描+标签计数的方式,准确找到目标行所属的最近标签:

function getLineContent(line: string, content: string): string | null {
    // 定位目标行在XML内容中的起始位置
    const normalizedLine = line.trimEnd();
    const lineIndex = content.indexOf(normalizedLine);
    if (lineIndex === -1) return null;

    // 向前扫描找到最近的非自闭合开始标签
    let startTagInfo: { index: number; name: string; length: number } | null = null;
    const startTagRegex = /<([a-zA-Z]+)(?:\s[^>]*?)?>/g;
    let match: RegExpMatchArray | null;
    
    // 遍历所有起始标签,筛选出目标行之前的最近有效标签
    while ((match = startTagRegex.exec(content)) !== null) {
        if (match.index < lineIndex && !match[0].endsWith('/>')) {
            startTagInfo = {
                index: match.index,
                name: match[1],
                length: match[0].length
            };
        }
    }
    if (!startTagInfo) return null;

    const { index: startIndex, name: tagName, length: tagLength } = startTagInfo;
    let scanPos = startIndex + tagLength;
    let tagCount = 1;
    const endTagRegex = new RegExp(`(</${tagName}>|<${tagName}(?:\\s[^>]*?)?>|<${tagName}(?:\\s[^>]*?)?/>)`, 'g');

    // 向后扫描匹配对应的结束标签,处理嵌套标签的计数
    while ((match = endTagRegex.exec(content)) !== null) {
        if (match.index < scanPos) continue;
        
        if (match[0].startsWith(`</${tagName}>`)) {
            tagCount--;
            if (tagCount === 0) {
                // 返回从开始标签到结束标签的完整内容
                return content.slice(startIndex, match.index + match[0].length);
            }
        } else if (match[0].startsWith(`<${tagName}`) && !match[0].endsWith('/>')) {
            tagCount++;
        }
        scanPos = match.index + match[0].length;
    }

    return null;
}

实现说明

  1. 定位目标行:先找到目标行在XML内容中的位置,确保能准确锁定范围。
  2. 寻找最近开始标签:从目标行位置向前遍历所有非自闭合的起始标签,记录最后一个(最近的)标签的位置和名称。
  3. 匹配对应结束标签:从开始标签之后的位置向后扫描,通过标签计数处理嵌套情况(遇到同名起始标签计数+1,遇到结束标签计数-1,计数归0时找到对应闭合标签)。
  4. 返回结果:截取从开始标签到对应结束标签的完整字符串。

测试传入FireFly is High和给定XML内容,将正确返回<Mobile>标签的完整内容。

内容的提问来源于stack exchange,提问作者CraZyDroiD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:43:24