正则表达式生成段落:如何排除空段落及h*标签开头内容?
Great question! To avoid wrapping empty lines or lines starting with heading tags (<h1> to <h6>) in <p> tags, you can combine two simple checks before applying your regex: one for empty/whitespace-only content, and another for heading tag prefixes.
Here's a step-by-step breakdown of how to do this:
1. Check for Empty or Whitespace-Only Paragraphs
A "empty" paragraph might include lines with just spaces, tabs, or line breaks—so we need to account for those too. Use this regex to match such lines:^\s*$
^= start of the line\s*= zero or more whitespace characters (spaces, tabs, newlines)$= end of the line
2. Check for Heading Tag Prefixes
To catch lines starting with any heading tag (even if they have attributes like <h2 class="section">), use this regex:^<h[1-6]\s*>.*$
^<h[1-6]= starts with<h1>to<h6>\s*= allows optional whitespace/attributes after the heading number>.*$= matches the rest of the line after the opening tag
Add the i flag (e.g., /^<h[1-6]\s*>.*$/i) if you need case-insensitive matching (for tags like <H1>).
3. Integrate Checks into Your Workflow
Let’s use JavaScript as an example (adjust syntax for your language of choice):
const rawText = "Your input text here"; const lines = rawText.split('\n'); const formattedText = lines.map(line => { const trimmedLine = line.trim(); // Skip wrapping if line is empty or starts with a heading tag if (!trimmedLine || /^<h[1-6]\s*>/.test(line)) { return line; // Leave the line unchanged } // Wrap valid paragraphs in <p> tags return `<p>${line}</p>`; }).join('\n');
Quick Tips:
- If you’re working with multi-line paragraphs (not just single lines), make sure you split your content into logical paragraph blocks first instead of processing line-by-line.
- Test your regex with edge cases (like headings with extra spaces or attributes) to ensure it catches all valid heading lines.
内容的提问来源于stack exchange,提问作者party34

