PHP提取标签内字符串转htmlentities及带<code>标签内容网页展示方案
Got it, let's tackle your two PHP/HTML questions one by one, with practical code examples you can use right away:
1. 从<tag></tag>标签提取字符串并做htmlentities转义
If you're dealing with simple, non-nested <tag> elements (no attributes or complex HTML structure), regular expressions work perfectly here. Just note that regex isn't ideal for full HTML parsing, but it's totally fine for this specific use case.
Here's a reusable function that does both extraction and escaping:
function extractAndEscapeTagContent($inputHtml, $targetTag) { // Regex to match everything between opening and closing tags (including newlines) $regexPattern = "/<$targetTag>(.*?)<\/$targetTag>/s"; preg_match_all($regexPattern, $inputHtml, $matches); // Apply htmlentities to each extracted string to escape special characters $escapedResults = array_map(function($content) { // ENT_QUOTES escapes both single and double quotes; UTF-8 avoids garbled text return htmlentities($content, ENT_QUOTES, 'UTF-8'); }, $matches[1]); return $escapedResults; } // Example usage $sampleHtml = '<title>My Page Title</title><description>包含特殊字符的内容:& < > "</description>'; $titles = extractAndEscapeTagContent($sampleHtml, 'title'); $descriptions = extractAndEscapeTagContent($sampleHtml, 'description'); print_r($titles); // Outputs: Array ( [0] => My Page Title ) echo $descriptions[0]; // Outputs: 包含特殊字符的内容:& < > "
2. 展示带编程语言类名的<code>标签内容
When pulling content from your database that includes <code class="language-xxx"> blocks, the key goals are:
- Make sure special characters in the code (like
<,>,") don't break your HTML - Preserve the
language-xxxclass so you can add syntax highlighting later if you want
Option 1: Preprocess raw database data (Recommended)
If your database stores raw code + language info (instead of pre-wrapped <code> tags), this is the cleanest approach:
// Example data fetched from database $dbRecord = [ 'body_text' => '下面是一段PHP代码示例:', 'code_snippets' => [ ['language' => 'php', 'code' => '<?php echo "Hello World!"; ?>'], ['language' => 'css', 'code' => 'body { background-color: #f0f0f0; }'] ] ]; // First, escape and print the regular text echo htmlentities($dbRecord['body_text'], ENT_QUOTES, 'UTF-8'); // Loop through each code snippet and render properly foreach ($dbRecord['code_snippets'] as $snippet) { // Escape the language class to prevent XSS $safeLang = htmlentities($snippet['language'], ENT_QUOTES, 'UTF-8'); // Escape the code content to display it correctly in HTML $safeCode = htmlentities($snippet['code'], ENT_QUOTES, 'UTF-8'); echo "<br><code class=\"language-{$safeLang}\">{$safeCode}</code>"; }
Option 2: Process existing <code> tags in database content
If your database already returns a string with mixed text and <code> tags (e.g., "来看这段CSS:<code class='language-css'>body{color:red;}</code>"), use this function to escape the code inside the tags:
function escapeCodeInsideTags($inputContent) { // Regex to match <code> tags with language classes and their content $codeTagPattern = "/(<code class=\"language-([a-z]+)\">)(.*?)(<\/code>)/s"; // Use a callback to escape only the code inside the tags return preg_replace_callback($codeTagPattern, function($matches) { $openingTag = $matches[1]; $escapedCode = htmlentities($matches[3], ENT_QUOTES, 'UTF-8'); $closingTag = $matches[4]; return "{$openingTag}{$escapedCode}{$closingTag}"; }, $inputContent); } // Example usage $dbContent = '这里有段PHP代码:<code class="language-php"><?php $name = "Alice"; echo $name; ?></code>'; $safeContent = escapeCodeInsideTags($dbContent); echo $safeContent;
Bonus: Add syntax highlighting
If you want your code blocks to look polished, drop in a lightweight library like Prism.js:
- Include Prism's CSS and JS files in your HTML
<head> - Ensure your
<code>tags have the correctlanguage-xxxclass - Prism will automatically apply syntax highlighting when the page loads
内容的提问来源于stack exchange,提问作者aidron

