You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VB.NET正则表达式匹配范围过大求助:如何精准匹配目标文本?

Fixing Unintended Early Matches in Your Regex for VB.NET

Let’s break down why your current regex is grabbing content above your target snippet and how to fix it—even when there are stray "On" words or HTML tags in the mix.

The Core Problem

Your regex On[\s\S]+?at[\s\S]+?person[\s\S]+?has written: uses lazy matching, but it doesn’t enforce that the "On" you’re targeting is the last valid starting point before the ending "has written:" marker. If there’s an "On" earlier in the text, the regex will start matching from that first "On" and keep going until it finds the required subsequent terms (at, person, has written:), pulling in all the content in between (including your target snippet’s upper context).

Solution 1: Target the Last Valid "On" (Best for Clean Text)

To ensure you only match the "On" that’s immediately associated with the final "has written:" marker, use a negative lookahead to eliminate any earlier "On" instances that could trigger a false match. Here’s the adjusted regex for VB.NET:

(?s)\bOn\b(?=.*?\bat\b.*?\bperson\b.*?has written:)(?!.*\bOn\b.*?\bat\b.*?\bperson\b.*?has written:).*?has written:.*

Let’s break this down:

  • (?s): Enables single-line mode, so . matches newlines (same as your [\s\S] trick, but more concise).
  • \bOn\b: Matches the word "On" as a whole word (avoids partial matches like "Onward").
  • (?=.*?\bat\b.*?\bperson\b.*?has written:): Positive lookahead confirms this "On" is followed by the full sequence of required terms.
  • (?!.*\bOn\b.*?\bat\b.*?\bperson\b.*?has written:): Negative lookahead ensures there are no other valid "On...has written:" sequences later in the text—so we only pick the last (and correct) one.
  • .*?has written:.*: Matches from the valid "On" through the ending marker and any content below it (per your requirement to allow matching target text below).

Solution 2: HTML-Specific Adjustments

For HTML content, stray tags or characters (like <) can throw off basic matching. Build on the above regex by adding constraints for HTML structure, and use [\s\S] instead of . if you need to explicitly handle tag line breaks:

(?s)\bOn\b(?=.*?\bat\b.*?\bperson\b.*?has written:)(?!.*\bOn\b.*?\bat\b.*?\bperson\b.*?has written:)[\s\S]+?has written:[\s\S]*

If your target snippet is wrapped in specific HTML tags (e.g., <blockquote>), you can add that to the lookahead to make it even more reliable:

(?s)\bOn\b(?=.*?<blockquote>.*?\bat\b.*?\bperson\b.*?has written:)(?!.*\bOn\b.*?<blockquote>.*?\bat\b.*?\bperson\b.*?has written:)[\s\S]+?has written:[\s\S]*

Quick Fix for Line-Specific "On"

If your target "On" always starts a new line (common in email replies), you can combine line anchors with the above logic to narrow things down faster:

(?s)(?<=^|\r\n)\bOn\b(?=.*?\bat\b.*?\bperson\b.*?has written:)(?!.*\bOn\b.*?\bat\b.*?\bperson\b.*?has written:).*?has written:.*

The (?<=^|\r\n) ensures "On" is at the start of a line, eliminating any "On" words buried in sentences above.

Testing in VB.NET

Remember that in VB.NET, you need to escape backslashes in string literals, so your regex string would look like this:

Dim pattern As String = "(?s)\bOn\b(?=.*?\bat\b.*?\bperson\b.*?has written:)(?!.*\bOn\b.*?\bat\b.*?\bperson\b.*?has written:).*?has written:.*"
Dim regex As New Regex(pattern)
Dim match As Match = regex.Match(yourInputText)

This approach guarantees you never pull in content above your target snippet, while still allowing matches of the target and everything below it.

内容的提问来源于stack exchange,提问作者Whitekang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:47:39