PHP编码HTML中{{}}占位符后解码,适配Magento DOMDocument处理
Great question! I've run into exactly this issue with Magento's DOMDocument handling of widget placeholders before, and the key is to isolate the placeholders from DOM's HTML parsing engine entirely. Here's the optimal solution I've implemented for this exact scenario:
Optimal Solution: Placeholder Mapping & Replacement
The core idea is to temporarily replace all {{...}} placeholders with unique, valid string markers before passing the content to DOMDocument. This avoids parsing errors caused by unescaped quotes in placeholders (like your <img src="{{image url="mediadir/someimage.jpg"}}"/> example) and ensures DOMDocument doesn't modify the placeholder content. After processing, we map the markers back to the original placeholders.
Step 1: Encode Placeholders (Before DOM Processing)
Use a regex to capture all {{...}} placeholders, assign each a unique ID, and store a mapping of IDs to original placeholder content. This turns invalid HTML snippets into valid ones DOMDocument can parse safely.
$content = $this->cleanMagentoCode($page->getContent()); // Initialize a map to track unique IDs -> original placeholders $placeholderMap = []; // Replace all {{...}} placeholders with unique markers $encodedContent = preg_replace_callback( '/{{(.*?)}}/', // Non-greedy match to capture individual placeholders function ($matches) use (&$placeholderMap) { $uniqueId = 'MAGENTO_WIDGET_' . uniqid(); // Unique, collision-resistant ID $placeholderMap[$uniqueId] = '{{' . $matches[1] . '}}'; return $uniqueId; }, $content );
Step 2: Process HTML with DOMDocument
Load the encoded content (now valid HTML) into DOMDocument and run your existing link replacement logic as usual:
$dom = new \DOMDocument(); libxml_use_internal_errors(true); $dom->loadHTML($encodedContent, LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD); // Your existing link replacement logic here (unchanged) $elements = $dom->getElementsByTagName("a"); for ($i = $elements->length - 1; $i >= 0; $i--) { $link = $elements->item($i); $found = false; // ... [keep all your existing link validation logic] ... if (!$found) { $url = parse_url($link->getAttribute('href')); if (isset($url['path'])) { $identifier = rtrim(ltrim($url['path'], '/'), '/'); try { $pagelink = $this->pageRepository->getById($identifier); if ($this->fixLinksFlag($input)) { $link_template = stripos($link->getAttribute('class'), "btn") !== FALSE ? "widget/link/link_block.phtml" : "widget/link/link_inline.phtml"; $widgetcode = '{{widget type="Magento\Cms\Block\Widget\Page\Link" anchor_text="' . $link->nodeValue . '" template="' . $link_template . '" page_id="' . $pagelink->getId() . '"}}'; $widget = $dom->createTextNode($widgetcode); $link->parentNode->replaceChild($widget, $link); } } catch (\Exception $e) { // Handle missing pages or errors as needed $missing_pages++; } } } }
Step 3: Decode Placeholders (After DOM Processing)
Replace the unique markers back with the original placeholder content before saving the page:
$processedContent = $dom->saveHTML(); // Restore all original placeholders from the mapping foreach ($placeholderMap as $uniqueId => $originalPlaceholder) { $processedContent = str_replace($uniqueId, $originalPlaceholder, $processedContent); } // Save the final processed content $page->setContent($this->dirtyMagentoCode($processedContent)); $page->save();
Why This Works Better Than Other Approaches
- Full Compatibility: Handles placeholders in both text nodes and HTML attributes (the biggest pain point in your original issue).
- No DOM Tampering: The unique markers are treated as plain text by DOMDocument, so no escaping or structural changes happen to your placeholder content.
- Collision Resistance: The
uniqid()prefix ensures almost zero chance of conflicting with existing page content.
Bonus Optimization
If you're processing a large number of pages, you can pre-generate a longer, more unique prefix (e.g., MAGENTO_STORE_123_WIDGET_) to further reduce collision risk, though this is rarely necessary for most Magento sites.
内容的提问来源于stack exchange,提问作者Wisey

