DOMDocument中title标签的短破折号(en dash)为何会破坏Unicode字符串?
<title> tag break Unicode text extraction with DOMDocument? Problem Reproduction
Here's the code that demonstrates the frustrating issue you ran into:
<?php $html = <<<'HTML' <!DOCTYPE html> <html><head> <title>example.org – example.org - example.org</title> <meta charset="utf-8" /> </head> <body>Trädgård</body> </html> HTML; $domd = new DOMDocument("1.0", "UTF-8"); @$domd->loadHTML($html); $xp = new DOMXPath($domd); $interesting = $domd->getElementsByTagName("body")->item(0)->textContent; var_dump($interesting, bin2hex($interesting));
Broken Output (with en dash in title):
string(14) "Trädgård" string(28) "5472c383c2a46467c383c2a57264"
Correct Output (after removing the en dash from title):
string(10) "Trädgård" string(20) "5472c3a46467c3a57264"
Root Cause
The issue boils down to DOMDocument's finicky automatic encoding detection logic when handling HTML that lacks an explicit, fully qualified encoding declaration.
Here's the step-by-step breakdown:
- The en dash character (
–, U+2013) uses the UTF-8 byte sequenceE2 80 93. - When
loadHTMLparses your document, it scans for clues to guess the encoding. Your existing<meta charset="utf-8">isn't enough to override its auto-detection—DOMDocument prioritizes byte patterns over partial meta tags. That en dash's byte sequence triggers a false positive: it looks like valid Windows-1252 (a common single-byte encoding) to the detector. - DOMDocument then decodes the entire document as Windows-1252, re-encodes it to UTF-8 for internal storage, and this double-encoding mangles your original UTF-8 characters. For example:
- The original
ä(UTF-8:C3 A4) gets split into two separate Windows-1252 characters (Ãand¤), which when re-encoded to UTF-8 becomesC3 83 C2 A4—the exact gibberish you saw.
- The original
Remove the en dash, and there are no ambiguous byte sequences to throw off the detector, so DOMDocument correctly identifies UTF-8 and processes the text normally.
Solutions
You can fix this by forcing DOMDocument to use UTF-8 explicitly, bypassing its error-prone auto-detection:
Add an XML encoding prefix to your HTML string before loading it:
@$domd->loadHTML('<?xml encoding="UTF-8">' . $html);This tells the underlying libxml library exactly which encoding to use, leaving no room for guesswork.
Use a fully qualified Content-Type meta tag:
Replace your existing meta tag with this more explicit version:<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />This gives DOMDocument a clear, unmissable hint about the encoding.
Combine with libxml flags for extra reliability:
Pass these flags toloadHTMLto disable default DTD elements and suppress minor errors, which can help avoid additional encoding edge cases:@$domd->loadHTML($html, LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD | LIBXML_NOERROR);
内容的提问来源于stack exchange,提问作者hanshenrik

