PHP Simple HTML DOM爬取开启WordPress缓存的网站返回乱码如何解决
乱码根因分析
- 核心原因是WordPress内置缓存生成静态缓存时,默认会根据请求头的
Accept-Encoding返回gzip/brotli压缩后的内容。你第一次请求时缓存还未生成,返回的是未压缩的实时渲染内容,第二次请求命中已生成的压缩缓存,服务端直接返回压缩流,你未解码就输出就会出现乱码。
解决方案
方案1:修改请求头声明不接受压缩内容
直接在请求头中添加Accept-Encoding: identity,告知服务器不要返回压缩内容,无需修改原有逻辑:
include('/simple_html_dom.php'); $url = "https://www.2311666.com.tw/"; $opts = array( 'http'=> array( 'header'=>"User-Agent:Chrome/94.0.4606.81\r\n"."Accept-Encoding: identity\r\n" ) ); $context = stream_context_create($opts); $html = new simple_html_dom(); $get = file_get_contents($url,false,$context); $html->load($get); echo $html;
方案2:自动解码压缩内容(兼容性更强)
如果要开启压缩提升传输效率,可以先判断响应头的编码类型,对压缩内容主动解码:
include('/simple_html_dom.php'); $url = "https://www.2311666.com.tw/"; $opts = array( 'http'=> array( 'header'=>"User-Agent:Chrome/94.0.4606.81\r\n"."Accept-Encoding: gzip, deflate, br\r\n" ) ); $context = stream_context_create($opts); $get = file_get_contents($url,false,$context); // 解析返回内容的编码类型 $encoding = ''; foreach ($http_response_header as $header) { $headerParts = explode(':', strtolower(trim($header)), 2); if (count($headerParts) == 2 && $headerParts[0] == 'content-encoding') { $encoding = trim($headerParts[1]); break; } } // 按编码类型解码 if ($encoding == 'gzip' || $encoding == 'x-gzip') { $get = gzdecode($get); } elseif ($encoding == 'deflate') { $get = gzinflate(substr($get, 2)); } elseif ($encoding == 'br') { $get = brotli_uncompress($get); } $html = new simple_html_dom(); $html->load($get); echo $html;
补充说明
如果使用cURL请求,可以直接添加CURLOPT_ENCODING => ''参数,让cURL自动处理所有编码类型的解码,不需要手动编写解码逻辑。请确认服务器已开启zlib、brotli扩展,避免解码失败。
内容的提问来源于stack exchange,提问作者Chris Lin
相关产品推荐
相关产品推荐

