You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

相同字符正则匹配问题:如何提取Markdown中##间内容

Hey there! Great question—this is a super common gotcha when working with regex delimiters that are identical. Let’s break down why your original pattern falls short and fix it properly.

Why Your Original Regex Fails

Your initial regex ##(.*?)(?=##) has two key issues:

  1. By default, the . wildcard doesn’t match newline characters. Since Markdown content almost always has line breaks after headings, this would only capture text up to the first line break, not the full section.
  2. The lookahead (?=##) doesn’t distinguish between actual section headings (which start on a new line) and any ## that might appear in the content itself (like in code blocks or inline text). This can cause premature truncation of your sections.

The Fix: Target Section Headings Explicitly

We need a regex that:

  • Starts with ## (the heading we want to capture)
  • Matches all content until the next line starting with ## (a new section heading) or the end of the document
  • Includes the starting ## but excludes the next section’s ##

Here’s the pattern, with variations for different regex flavors:

For PCRE (Python, PHP, etc.)

Use the DOTALL flag to make . match newlines, and target new-line-prefixed ## for the next heading:

##.*?(?=\n##|\Z)
  • ##: Matches the start of your target section heading
  • .*?: Non-greedily matches all characters (including newlines, thanks to DOTALL)
  • (?=\n##|\Z): Positive lookahead that stops when it hits either a new line followed by ## (next section) or the end of the string (\Z, for the final section)

In Python, you’d use it like this:

import re

with open("your_large_doc.md", "r", encoding="utf-8") as f:
    markdown_content = f.read()

# Capture all sections
sections = re.findall(r'##.*?(?=\n##|\Z)', markdown_content, re.DOTALL)

# Export each section to a separate file
for idx, section in enumerate(sections, 1):
    with open(f"section_{idx}.md", "w", encoding="utf-8") as out_f:
        out_f.write(section.strip() + "\n")

For JavaScript (Browser or Node.js)

JavaScript doesn’t have a DOTALL flag, so use [\s\S] (matches all whitespace and non-whitespace characters, i.e., everything including newlines) instead of .:

##[\s\S]*?(?=\n##|\Z)

In Node.js, here’s how you’d use it to export sections:

const fs = require('fs');

const markdownContent = fs.readFileSync('your_large_doc.md', 'utf8');
const sectionRegex = /##[\s\S]*?(?=\n##|\Z)/g;
const sections = markdownContent.match(sectionRegex);

sections.forEach((section, idx) => {
    fs.writeFileSync(`section_${idx + 1}.md`, section.trim() + '\n', 'utf8');
});

Testing with Your Example Text

If we run this against your sample content:

RegExr was created by gskinner.com, and is proudly hosted by Media Temple. Edit the Expression & Text to see matches. Roll over matches or the expression for details. PCRE & JavaScript flavors of RegEx are supported.

The side bar includes a Cheatsheet, full Reference, and Help. You can also Save & Share with the Community, and view patterns you create or favorite in My Patterns.

Tools

Explore results with the Tools below. Replace & List output custom results. Details lists capture groups. Explain describes your expression in plain English.

We’ll get three separate sections, each including the starting ## heading and all content up to (but not including) the next ## heading. Perfect for exporting to individual files!

内容的提问来源于stack exchange,提问作者Will Simpson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:57:21