You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页中Regex Negative Lookahead应用:匹配不含指定短语的代码

Regex to Match Web Code Blocks & Exclude Those With a Specific Phrase

Got it, let's work through this regex problem together. You’re looking to match all code blocks on a webpage but exclude any that contain a specific phrase—your example uses the word 'is', but you mentioned you’ll make the negative lookahead more specific later. Here’s how to pull this off:

Step 1: Start with a Base Regex for Code Blocks

First, let's target the typical code block structure used on most webpages: wrapped in <pre><code> tags. If your site uses different containers (like standalone <code> tags or custom classed divs), you can adjust this part later.

The base regex to match these blocks is:

<pre><code>([\s\S]*?)<\/code><\/pre>
  • [\s\S]*? matches every character (including newlines) in a non-greedy way—this ensures we stop at the first closing </code></pre> instead of grabbing all code blocks on the page in one go.

Step 2: Add the Negative Lookahead to Exclude Target Phrases

Now we need to modify this to skip any code block that contains your target phrase. We’ll use a negative lookahead ((?!...)) right after the opening tags to check for the forbidden content before matching the rest of the block.

Here’s the adjusted regex for your 'is' example:

<pre><code>(?![\s\S]*\bis\b)([\s\S]*?)<\/code><\/pre>

Let’s break down the key parts:

  • (?![\s\S]*\bis\b): This negative lookahead checks that the word is (wrapped in \b word boundaries to avoid partial matches like this or island) does NOT appear anywhere in the code block. The [\s\S]* ensures we check across all lines in the block, not just the first one.
  • If you want to exclude blocks with the substring is (not just the standalone word), remove the \b boundaries: (?![\s\S]*is).

Step 3: Customization Tips for Your Actual Use Case

  • Adjust code containers: If your site uses different code wrappers (e.g., <code class="snippet">...</code> or <div class="code-block">...</div>), swap out the <pre><code> and </code></pre> parts to match your HTML structure.
  • Escape special characters: If your target phrase has regex-specific characters (like ., *, or ?), escape them with a backslash (\). For example, to exclude blocks with user.name, use \buser\.name\b.
  • Case insensitivity: Add the i flag (depending on your regex engine) if you want to match the phrase regardless of capitalization (e.g., Is, IS).

Important Caveat

Regex isn’t perfect for parsing HTML. If your webpage has nested tags, malformed HTML, or dynamically loaded code blocks, a proper HTML parser (like BeautifulSoup for Python, Cheerio for JavaScript) paired with a simple regex check on the extracted code content will be more reliable. But for well-structured, static code blocks, this regex works great.

内容的提问来源于stack exchange,提问作者Dylan Frost

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:59:37