如何修改PHP cURL代码以获取Google搜索结果页面的源码
How to Scrape Google Search Results Without Getting Blocked/Redirected to Consent Page
Hey there, let's fix that frustrating issue where your cURL request keeps getting sent to Google's consent page instead of returning the actual search results. The core problem is that Google can tell your request is coming from a script (not a real browser) and throws the consent wall at you. Here's how to tweak your code to mimic a legitimate browser request:
Key Fixes to Implement
- Add a real User-Agent header: This makes your request look like it's coming from a standard web browser instead of a script.
- Enable automatic redirect following: Let cURL handle 302 redirects automatically, including navigating past the consent page once we mimic a browser that would accept the consent.
- Handle cookies: Google uses cookies to track consent status, so we need to store and reuse cookies between requests to maintain that state.
Modified Working Code
<!DOCTYPE html> <html> <body> <!-- This program saves source code of Google search results to an external file --> <?php ini_set('display_errors', true); error_reporting(E_ALL); // Target Google search URL $url = 'https://www.google.com/search?q=blue+car'; // Initialize cURL session $ch = curl_init($url); // Set cURL options to mimic a browser curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); // Follow redirects automatically (handles 302s) curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); // Set a valid User-Agent (matches a modern Chrome browser) curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'); // Enable cookie handling to persist consent status curl_setopt($ch, CURLOPT_COOKIEJAR, 'cookies.txt'); // Save cookies to file curl_setopt($ch, CURLOPT_COOKIEFILE, 'cookies.txt'); // Load cookies from file // Increase timeout to accommodate multiple redirects curl_setopt($ch, CURLOPT_TIMEOUT, 30); $html = curl_exec($ch); if(empty($html)) { echo "<pre>cURL request failed: " . curl_error($ch) . "</pre>"; } else { // Save the actual search results to file $myfile = fopen("google_search_results.txt", "w") or die("Unable to open file!"); fwrite($myfile, $html); fclose($myfile); echo "<pre>Successfully saved Google search results to google_search_results.txt</pre>"; } // Close cURL session curl_close($ch); ?> </body> </html>
Explanation of Changes
CURLOPT_FOLLOWLOCATION: This tells cURL to automatically follow any redirects (like the 302 you encountered) instead of returning the redirect page itself.CURLOPT_USERAGENT: Google checks this header to identify the client. Using a real browser's User-Agent string makes your request look legitimate. You can update this to match your current browser's UA if needed.- Cookie Handling: The
CURLOPT_COOKIEJARandCURLOPT_COOKIEFILEoptions let cURL save and reuse cookies. Once the consent is "accepted" via the browser-like request, the cookie will tell Google you've already given consent, so it won't redirect you again. - Timeout: Added a longer timeout to ensure the request has enough time to complete, especially with multiple redirects.
Quick Notes
- Make sure your server has permission to write to
cookies.txtand the output file (google_search_results.txt). You may need to adjust file permissions if you run into write errors. - Google's anti-scraping measures can evolve over time, so if this stops working, try updating the User-Agent or adding additional headers (like
Accept-Language) to make the request even more realistic. - Always respect Google's Terms of Service when scraping their content.
内容的提问来源于stack exchange,提问作者Jane N.
相关产品推荐
相关产品推荐

