使用PHP cURL请求外部网站时遭遇403错误的解决求助
Hey there, let's dig into that 403 error you're facing with TransUnion's site. Even after setting a custom User-Agent, their anti-scraping defenses are likely flagging your request as automated—financial sites like this have pretty strict measures to block bots. Let's walk through some adjustments to your code to get past this:
Common Issues & Fixes
1. Update Your User-Agent to a Modern Value
The IE6 User-Agent you're using is extremely outdated, and most modern sites will immediately flag it as suspicious. Swap it for a recent, widely used browser's UA string. For example:
$agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36';
2. Add Realistic Request Headers
Browsers send more than just a User-Agent—sites expect headers like Accept, Accept-Language, and Referer to validate legitimate requests. Add these to your cURL options:
curl_setopt($ch, CURLOPT_HTTPHEADER, [ 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language: en-US,en;q=0.5', 'Referer: https://www.google.com/', // Simulate coming from a search engine 'DNT: 1', // Do Not Track header (optional but adds realism) 'Connection: keep-alive', 'Upgrade-Insecure-Requests: 1' ]);
3. Enable Cookie Handling
Many sites use cookies to track session validity. Enable cURL to store and send cookies by adding these options:
// Create a temporary cookie file (make sure the directory is writable) $cookieFile = tempnam(sys_get_temp_dir(), 'curl_cookies'); curl_setopt($ch, CURLOPT_COOKIEJAR, $cookieFile); curl_setopt($ch, CURLOPT_COOKIEFILE, $cookieFile);
Don't forget to clean up the cookie file after use if you don't need it anymore.
4. Check for cURL Errors
Adding error checking will help you debug if the 403 persists—you'll get specific details about what's going wrong instead of just a blank failure.
Modified Full Code
Here's your updated function with all the above fixes:
$agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'; function file_get_contents_curl($url, $agent) { $ch = curl_init(); $cookieFile = tempnam(sys_get_temp_dir(), 'curl_cookies'); curl_setopt($ch, CURLOPT_AUTOREFERER, TRUE); curl_setopt($ch, CURLOPT_HEADER, 0); curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); curl_setopt($ch, CURLOPT_URL, $url); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, TRUE); curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, FALSE); curl_setopt($ch, CURLOPT_USERAGENT, $agent); curl_setopt($ch, CURLOPT_VERBOSE, true); // Add realistic headers curl_setopt($ch, CURLOPT_HTTPHEADER, [ 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language: en-US,en;q=0.5', 'Referer: https://www.google.com/', 'DNT: 1', 'Connection: keep-alive', 'Upgrade-Insecure-Requests: 1' ]); // Enable cookie handling curl_setopt($ch, CURLOPT_COOKIEJAR, $cookieFile); curl_setopt($ch, CURLOPT_COOKIEFILE, $cookieFile); $data = curl_exec($ch); // Check for cURL errors if(curl_errno($ch)) { echo 'cURL Error: ' . curl_error($ch); } curl_close($ch); unlink($cookieFile); // Clean up cookie file return $data; } $homepage = file_get_contents_curl("https://www.transunion.com/", $agent); echo $homepage;
Important Note
Before proceeding, make sure you're complying with TransUnion's Terms of Service. Scraping financial sites can violate their policies, and they may have legal protections against automated access. Always get explicit permission if you're planning to use this for anything beyond personal testing.
内容的提问来源于stack exchange,提问作者user18059799

