能否通过cron每日执行PHP脚本,用file_get_contents模拟用户访问?
file_get_contents() via cron simulate a real user visit? Great question! Let’s break this down clearly so you know exactly what’s happening here:
First off: No, this won’t be identical to a real user’s browser visit—but it can mimic a basic HTTP request. Here’s the full breakdown:
Key differences from a real user’s visit
- User-Agent mismatch: By default,
file_get_contents()sends a generic User-Agent string likePHP/7.4.3(depending on your PHP version). Real browsers send unique, detailed strings (e.g.,Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36...), so the target server can instantly tell this isn’t a human-driven browser. - No cookie persistence: A real user’s browser automatically stores and sends cookies to maintain sessions or personalized content. Your current script doesn’t save or reuse cookies, so any session-based pages (like logged-in dashboards) won’t load as they would for a real user.
- No JavaScript execution: Browsers parse and run JavaScript, which often loads dynamic content or modifies the page after initial load.
file_get_contents()only fetches the raw, static HTML source—any JS-rendered content won’t be included, unlike what a real user sees. - Missing standard headers: Real requests include headers like
Accept-Language,Accept-Encoding, andRefererthat your defaultfile_get_contents()call doesn’t send. This makes the request look even more "bot-like."
How to make it closer to a real user visit
If you want to mimic a browser more accurately, swap file_get_contents() for PHP’s cURL extension. It lets you add browser-like headers, handle cookies, and follow redirects. Here’s an improved example:
// Initialize cURL session $ch = curl_init('http://www.example.com/'); // Set options to mimic a modern Chrome browser curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); // Follow redirects like a browser curl_setopt($ch, CURLOPT_COOKIEJAR, 'cookies.txt'); // Save cookies between requests curl_setopt($ch, CURLOPT_COOKIEFILE, 'cookies.txt'); // Send saved cookies curl_setopt($ch, CURLOPT_HTTPHEADER, [ 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language: en-US,en;q=0.5', 'Accept-Encoding: gzip, deflate', 'Connection: keep-alive', 'Upgrade-Insecure-Requests: 1' ]); // Execute the request and close the session $page = curl_exec($ch); curl_close($ch); echo $page;
Even with this setup, it still won’t execute JavaScript. If you need to fully replicate a user’s experience (including dynamic JS content), you’ll need tools like headless browsers (e.g., Puppeteer with PHP bindings or Selenium) that control a real browser instance in the background.
Final takeaway
Your current file_get_contents() script will make a valid HTTP request and fetch raw HTML, but it’s not equivalent to a real user’s visit. For simple tasks like monitoring static pages, it might be enough—but if you need to mimic a human’s browsing experience (including sessions, dynamic content, etc.), you’ll need to enhance the request or use a headless browser tool.
内容的提问来源于stack exchange,提问作者Daniel Garcia Sanchez

