curl返回部分页面DOM问题求助:爬虫脚本获取不完整HTML响应
Hey there, let's figure out why your cURL script is pulling incomplete HTML when the browser shows the full DOM (and you've confirmed no AJAX is involved). I've dealt with this exact scenario a few times, so here are actionable fixes to try:
1. Add Missing Request Headers to Mimic a Real Browser
Servers often flag requests as "non-human" if they lack common browser headers. Your current setup only includes a User-Agent and Expect: header—let's beef that up to match what a real Safari request sends:
curl_setopt($curl, CURLOPT_HTTPHEADER, array( 'Expect:', 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language: en-US,en;q=0.5', 'Connection: keep-alive', 'Upgrade-Insecure-Requests: 1' ));
2. Fix Encoding Handling
Your CURLOPT_ENCODING is set to 'identity', which tells the server you don't accept compressed content. Some servers behave unexpectedly when they can't compress responses—try setting it to accept gzip/deflate instead (cURL will automatically decompress the response for you):
curl_setopt($curl, CURLOPT_ENCODING, 'gzip, deflate');
3. Increase Timeout Limits
You only set a connection timeout (CURLOPT_CONNECTTIMEOUT), but cURL might be timing out mid-download before it gets the full HTML. Add a global timeout for the entire request:
curl_setopt($curl, CURLOPT_TIMEOUT, 30); // 30 seconds total timeout
4. Check for cURL Errors
You're not checking if cURL ran into any errors during execution—add this right after curl_exec() to catch hidden issues like truncated connections:
$response['str'] = curl_exec($curl); if (curl_errno($curl)) { echo 'cURL Error: ' . curl_error($curl); exit; }
5. Include Browser Cookies (If Needed)
Servers sometimes serve partial content if you don't have valid session cookies (which your browser automatically stores). Grab the cookie string from your browser's dev tools (under Network > Request Headers > Cookie) and add it to your cURL request:
curl_setopt($curl, CURLOPT_COOKIE, 'your_cookie_string_here');
Modified Full Script
Here's your code with all the above suggestions integrated:
$url = "https://example.com"; $curl = curl_init(); curl_setopt($curl, CURLOPT_URL, $url); curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); curl_setopt($curl, CURLOPT_CONNECTTIMEOUT, 10); curl_setopt($curl, CURLOPT_TIMEOUT, 30); // Added global timeout curl_setopt($curl, CURLOPT_MAXREDIRS, 10); curl_setopt($curl, CURLOPT_FOLLOWLOCATION, true); curl_setopt($curl,CURLOPT_USERAGENT,'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_9_3) AppleWebKit/537.75.14 (KHTML, like Gecko) Version/7.0.3 Safari/7046A194A'); curl_setopt($curl, CURLOPT_HTTPHEADER, array( 'Expect:', 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language: en-US,en;q=0.5', 'Connection: keep-alive', 'Upgrade-Insecure-Requests: 1' )); curl_setopt( $curl, CURLOPT_SSL_VERIFYPEER, false ); curl_setopt($curl, CURLOPT_ENCODING, 'gzip, deflate'); // Updated encoding $response['str'] = curl_exec($curl); // Check for cURL errors if (curl_errno($curl)) { echo 'cURL Error: ' . curl_error($curl); exit; } var_dump($response); $response['header'] = curl_getinfo($curl, CURLINFO_HTTP_CODE); curl_close($curl);
Start with adding the error checking first—it might immediately tell you why the response is truncated. If that doesn't work, try the header and encoding changes next.
内容的提问来源于stack exchange,提问作者Avi

