使用Curl获取指定网页内容时出现乱码,如何正确获取?
解决curl获取realestate.com.au页面返回乱码的问题
你的代码返回乱码主要有两个原因:
- 网站返回的内容启用了gzip压缩,但curl未开启自动解压,直接输出压缩内容导致乱码
- 缺少必要的请求头标识,被网站识别为非浏览器请求
以下是修复后的代码:
<?php function visit_url_with_curl($url) { $ch = curl_init($url); // Set cURL options curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, false); curl_setopt($ch, CURLOPT_SSL_VERIFYHOST, false); // 开启自动解压gzip/deflate压缩内容 curl_setopt($ch, CURLOPT_ENCODING, ''); // 添加模拟浏览器的请求头 curl_setopt($ch, CURLOPT_HTTPHEADER, [ 'User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language: en-US,en;q=0.5', 'Accept-Encoding: gzip, deflate, br' ]); // Execute cURL and store the response $response = curl_exec($ch); // Check for cURL errors if (curl_errno($ch)) { echo 'cURL error: ' . curl_error($ch); } // Close cURL session curl_close($ch); return $response; } $url = 'https://www.realestate.com.au/nsw/karuah-2324/'; $html = visit_url_with_curl($url); // Print the HTML content of the page echo $html; ?>
关键修改说明:
CURLOPT_ENCODING: '':让curl自动识别并解压服务器返回的压缩内容- 添加
User-Agent等请求头:模拟真实浏览器的请求特征,避免被网站的反爬机制拦截
内容的提问来源于stack exchange,提问作者Zoey Malkov
相关产品推荐
相关产品推荐

