JS正则表达式未按预期工作:换行符检测及匹配异常问题
Let's break down the two problems in your regex and walk through how to fix them:
Problem 1: Limited Whitespace Handling Between </abc> and <xyz>
Your original regex uses \n* to match characters between </abc> and <xyz>, which only accounts for newlines. If there are spaces, tabs, or a mix of whitespace types in that gap, the regex will fail to match correctly—this is likely the "newline detection issue" you're seeing.
Problem 2: Matching Across Multiple <abc> Elements When <xyz> Is Missing
The [\s\S]*? pattern is non-greedy but doesn't exclude other <abc> tags. When the <xyz> immediately after an <abc> is removed, the regex engine backtracks and keeps matching until it finds an </abc> followed by <xyz>—even if that means including the next <abc> element in the match, causing the unexpected output you mentioned.
The Corrected Regex
Here's the fixed regex that addresses both issues:
var regexapi = /<abc\s*name\s*=\s*"(.*?)"(?:(?!<abc)[\s\S])*?<\/abc>\s*<xyz>/gim;
Key Changes:
- Replace
\n*with\s*:\s*matches any whitespace character (spaces, tabs, newlines, carriage returns, etc.), so the regex works regardless of what whitespace separates</abc>and<xyz>. - Replace
[\s\S]*?with(?:(?!<abc)[\s\S])*?: This uses a negative lookahead(?!<abc)to ensure we don't match any characters that start a new<abc>tag. This prevents the regex from crossing into subsequent<abc>elements when the expected<xyz>is missing.
Corrected Code Example
const data = `<abc name = "foo"> <child>bar</child> </abc> <xyz>1</xyz> <abc name = "foo2"> <child>bar2</child> </abc> <xyz>5</xyz>`; const array1 = []; var regexapi = /<abc\s*name\s*=\s*"(.*?)"(?:(?!<abc)[\s\S])*?<\/abc>\s*<xyz>/gim; let resApi; while ((resApi = regexapi.exec(data))) { array1.push(resApi[0]); } console.log(array1[0]); // Outputs: <abc name = "foo"> <child>bar</child> </abc> <xyz>
Testing the Modified Input (Without <xyz>1</xyz>)
If you remove <xyz>1</xyz> from the input:
const data = `<abc name = "foo"> <child>bar</child> </abc> <abc name = "foo2"> <child>bar2</child> </abc> <xyz>5</xyz>`; // ... same code as above ... console.log(array1[0]); // Outputs: <abc name = "foo2"> <child>bar2</child> </abc> <xyz>
Now the regex correctly skips the first <abc> (since there's no <xyz> after it) and matches only the second <abc> that's followed by <xyz>.
Bonus: Capturing the <xyz> Content
If you want to capture the content inside <xyz> as well, adjust the regex to include an additional capture group:
var regexapi = /<abc\s*name\s*=\s*"(.*?)"(?:(?!<abc)[\s\S])*?<\/abc>\s*<xyz>(.*?)<\/xyz>/gim; // Then in the loop: array1.push({ abcName: resApi[1], xyzContent: resApi[2] });
Content of the question来源于stack exchange,提问作者Rogmier

