PHP从字符串中移除XML标签的实现方案
Got it, let's break down how to strip those XML tags from your PHP string that’s made up of multiple separate XML snippets. Here are a few reliable approaches, ranging from quick fixes to robust solutions:
If your XML structure is straightforward (no nested tags with tricky characters, no CDATA sections), a regex can get the job done fast. It’ll strip all XML tags and clean up extra whitespace:
$rawString = '<?xml version=\'1.0\' encoding=\'UTF-8\'?><root available-locales="en_US" default-locale="en_US"><Title language-id="en_US">Batman</Title></root> <?xml version=\'1.0\' encoding=\'UTF-8\'?><root available-locales="en_US" default-locale="en_US"><Title language-id="en_US">Wonder Woman</Title></root>'; // Strip all XML tags $strippedText = preg_replace('/<[^>]+>/', '', $rawString); // Clean up extra spaces/newlines $cleanText = trim(preg_replace('/\s+/', ' ', $strippedText)); echo $cleanText; // Output: Batman Wonder Woman
Note: This works for basic cases, but avoid it if your XML might contain edge cases like tags with attributes containing > characters or CDATA blocks—it can break the regex matching.
For a reliable solution that handles any valid XML structure, use PHP’s built-in DOMDocument. Since your string has multiple separate XML documents, we’ll first split them apart, then parse each one individually to extract text content:
$rawString = '<?xml version=\'1.0\' encoding=\'UTF-8\'?><root available-locales="en_US" default-locale="en_US"><Title language-id="en_US">Batman</Title></root> <?xml version=\'1.0\' encoding=\'UTF-8\'?><root available-locales="en_US" default-locale="en_US"><Title language-id="en_US">Wonder Woman</Title></root>'; // Split the string into individual XML snippets $xmlSnippets = preg_split('/(<\?xml)/', $rawString, -1, PREG_SPLIT_NO_EMPTY | PREG_SPLIT_DELIM_CAPTURE); $extractedTexts = []; // Process each snippet foreach ($xmlSnippets as $snippet) { // Reattach the <?xml prefix if it was split off if (strpos($snippet, '<?xml') !== 0) { $snippet = '<?xml' . $snippet; } // Suppress XML parsing warnings (in case of minor formatting issues) libxml_use_internal_errors(true); $dom = new DOMDocument(); $dom->loadXML($snippet); libxml_clear_errors(); // Extract and trim the text content $text = trim($dom->textContent); if (!empty($text)) { $extractedTexts[] = $text; } } // Combine all extracted text $finalText = implode(' ', $extractedTexts); echo $finalText; // Output: Batman Wonder Woman
Why this is better: It properly handles nested tags, attributes, and special characters that regex would mangle. It’s the go-to solution for production code where XML structure might vary.
If your XML snippets have a consistent, flat structure, SimpleXML offers a more concise syntax than DOMDocument:
$rawString = '<?xml version=\'1.0\' encoding=\'UTF-8\'?><root available-locales="en_US" default-locale="en_US"><Title language-id="en_US">Batman</Title></root> <?xml version=\'1.0\' encoding=\'UTF-8\'?><root available-locales="en_US" default-locale="en_US"><Title language-id="en_US">Wonder Woman</Title></root>'; $xmlSnippets = preg_split('/(<\?xml)/', $rawString, -1, PREG_SPLIT_NO_EMPTY | PREG_SPLIT_DELIM_CAPTURE); $extractedTexts = []; foreach ($xmlSnippets as $snippet) { if (strpos($snippet, '<?xml') !== 0) { $snippet = '<?xml' . $snippet; } libxml_use_internal_errors(true); $sxml = simplexml_load_string($snippet); libxml_clear_errors(); if ($sxml) { $text = trim((string)$sxml); if (!empty($text)) { $extractedTexts[] = $text; } } } $finalText = implode(' ', $extractedTexts); echo $finalText;
Note: This works great when your XML doesn’t have deeply nested elements, and you want code that’s easy to read and maintain.
内容的提问来源于stack exchange,提问作者dukenukem6487

