You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DOMDocument中title标签的短破折号(en dash)为何会破坏Unicode字符串?

Why does an en dash in the <title> tag break Unicode text extraction with DOMDocument?

Problem Reproduction

Here's the code that demonstrates the frustrating issue you ran into:

<?php 
$html = <<<'HTML'
<!DOCTYPE html>
<html><head>
<title>example.org – example.org - example.org</title>
<meta charset="utf-8" />
</head>
<body>Trädgård</body>
</html>
HTML;
$domd = new DOMDocument("1.0", "UTF-8");
@$domd->loadHTML($html);
$xp = new DOMXPath($domd);
$interesting = $domd->getElementsByTagName("body")->item(0)->textContent;
var_dump($interesting, bin2hex($interesting));

Broken Output (with en dash in title):

string(14) "Trädgård"
string(28) "5472c383c2a46467c383c2a57264"

Correct Output (after removing the en dash from title):

string(10) "Trädgård"
string(20) "5472c3a46467c3a57264"

Root Cause

The issue boils down to DOMDocument's finicky automatic encoding detection logic when handling HTML that lacks an explicit, fully qualified encoding declaration.

Here's the step-by-step breakdown:

  • The en dash character (–, U+2013) uses the UTF-8 byte sequence E2 80 93.
  • When loadHTML parses your document, it scans for clues to guess the encoding. Your existing <meta charset="utf-8"> isn't enough to override its auto-detection—DOMDocument prioritizes byte patterns over partial meta tags. That en dash's byte sequence triggers a false positive: it looks like valid Windows-1252 (a common single-byte encoding) to the detector.
  • DOMDocument then decodes the entire document as Windows-1252, re-encodes it to UTF-8 for internal storage, and this double-encoding mangles your original UTF-8 characters. For example:
    • The original ä (UTF-8: C3 A4) gets split into two separate Windows-1252 characters (Ã and ¤), which when re-encoded to UTF-8 becomes C3 83 C2 A4—the exact gibberish you saw.

Remove the en dash, and there are no ambiguous byte sequences to throw off the detector, so DOMDocument correctly identifies UTF-8 and processes the text normally.

Solutions

You can fix this by forcing DOMDocument to use UTF-8 explicitly, bypassing its error-prone auto-detection:

  1. Add an XML encoding prefix to your HTML string before loading it:

    @$domd->loadHTML('<?xml encoding="UTF-8">' . $html);
    

    This tells the underlying libxml library exactly which encoding to use, leaving no room for guesswork.

  2. Use a fully qualified Content-Type meta tag:
    Replace your existing meta tag with this more explicit version:

    <meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
    

    This gives DOMDocument a clear, unmissable hint about the encoding.

  3. Combine with libxml flags for extra reliability:
    Pass these flags to loadHTML to disable default DTD elements and suppress minor errors, which can help avoid additional encoding edge cases:

    @$domd->loadHTML($html, LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD | LIBXML_NOERROR);
    

内容的提问来源于stack exchange,提问作者hanshenrik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 19:38:09