PHP移除所有标题标签及内容、提取首个h2内容的实现方法
Hey there! Let's break down solutions for your two PHP text processing needs, and fix the issue with your current code for Task 1.
Task 1: Remove All Heading Tags (h1-h6) & Their Content, Keep Plain Text
Your existing code only handles tab replacement, trimming, and HTML entity escaping—it doesn't target heading tags or their content, which is why it's not working. Here are two reliable approaches:
Approach 1: Use DOMDocument (Recommended for Robust HTML Handling)
This method works even with malformed HTML (like your example's mismatched </h3> tag) and is far more reliable than regex for HTML manipulation:
$html = '<h2>Title of product</h2> </h3>Subtitle of product</h3> Here starts the first paragraph and goes on to the next... <h3>Subtitle of the paragraph</h3> Here goes on the second paragraph.'; $dom = new DOMDocument(); // Suppress warnings for non-standard HTML libxml_use_internal_errors(true); $dom->loadHTML($html); libxml_clear_errors(); $xpath = new DOMXPath($dom); // Target all h1 to h6 elements $headings = $xpath->query('//h1|//h2|//h3|//h4|//h5|//h6'); // Remove each heading from the DOM tree foreach ($headings as $heading) { $heading->parentNode->removeChild($heading); } // Extract remaining text, strip any leftover HTML, and clean up whitespace $plainText = trim(strip_tags($dom->saveHTML())); echo $plainText; // Output: Here starts the first paragraph and goes on to the next... Here goes on the second paragraph.
Approach 2: Regular Expressions (For Simple, Well-Formed HTML)
If you're certain your HTML is always properly structured, regex can work as a quicker alternative:
$html = '<h2>Title of product</h2> </h3>Subtitle of product</h3> Here starts the first paragraph and goes on to the next... <h3>Subtitle of the paragraph</h3> Here goes on the second paragraph.'; // Match any h1-h6 tag and its inner content (case-insensitive, multi-line) $cleanedText = preg_replace('/<h[1-6]>.*?<\/h[1-6]>/si', '', $html); // Strip remaining HTML tags and trim whitespace $plainText = trim(strip_tags($cleanedText)); echo $plainText; // Output: Here starts the first paragraph and goes on to the next... Here goes on the second paragraph.
Task 2: Extract Text from the First h2 Tag
Again, using DOMDocument is the most reliable way to grab this content without edge cases:
$html = '<h2>Title of product</h2> </h3>Subtitle of product</h3> Here starts the first paragraph and goes on to the next... <h3>Subtitle of the paragraph</h3> Here goes on the second paragraph.'; $dom = new DOMDocument(); libxml_use_internal_errors(true); $dom->loadHTML($html); libxml_clear_errors(); $xpath = new DOMXPath($dom); // Fetch the first h2 element in the document $firstH2 = $xpath->query('//h2[1]')->item(0); if ($firstH2) { $h2Content = trim($firstH2->nodeValue); echo $h2Content; // Output: Title of product } else { echo "No h2 tag found in the content."; }
Key Note
Always prefer DOMDocument over regex for HTML processing—HTML is not a regular language, so regex can break easily with unexpected tag structures, nested elements, or malformed markup.
内容的提问来源于stack exchange,提问作者matt

