正则表达式问题:如何匹配>与<之间的第三个单词
> and < Got it, let's break this down. Your original regex (?<=>)(\w*)(?=<) works great for grabbing every word wrapped between > and <, but adding {2} directly didn't hit the mark because you weren't accounting for the full surrounding structure of the first two matches. Here's how to fix it:
The Working Regex
(?:>\w*<){2}>(\w*)<
Let's Break It Down
(?:>\w*<){2}: This is a non-capturing group ((?:...)) that matches the full>WORD<structure exactly 2 times. The non-capturing part means we don't save these first two words—we just skip over them.>(\w*)<: After skipping the first two matches, this targets the third occurrence. The parentheses around\w*capture the word we actually want.
Example Usage
If your input string is:
Lorem>First<Ipsum>Second<Dolor>Third<Sit
Running the regex will capture Third as the first (and only) capturing group value.
Why Your Previous Attempts Failed
- Adding
{2}to the lookbehind ((?<=>){2}) would look for two consecutive>characters, which isn't your structure. - Adding
{2}to the lookahead ((?=<){2}) would require two consecutive<characters, which also doesn't match your pattern. - You need to match the entire
>WORD<block for the first two instances before targeting the third.
Alternative (If You Prefer Lookarounds)
If you specifically want to stick with lookarounds (though the above method is more straightforward), you can use a positive lookbehind that checks for the first two full blocks:
(?<=(?:>\w*<){2}>)(\w*)(?=<)
This works by verifying the word is preceded by two complete >WORD< blocks plus a >, then followed by <.
内容的提问来源于stack exchange,提问作者Fayaz

