如何用simple_html_dom处理重定向并解析跳转页?RSS字段获取需求
Hey there, let's work through your two main problems: fixing the first-failure HTML parsing issue, and grabbing that missing location field even though it's not in the RSS feed.
1. Fixing the HTML Parsing Redirect/Cookie Problem
The reason your simple_html_dom only works on the second visit is almost certainly because the first request hits a redirect (like a session-based check or cookie gate) that file_get_html doesn't handle by default. To fix this, we'll use cURL to fetch the page content—cURL lets us follow redirects and persist cookies between requests, mimicking the successful "second visit" behavior.
Here's how to adapt your code to use cURL with simple_html_dom:
// Initialize cURL session $ch = curl_init('YOUR_TARGET_PAGE_URL'); // Set cURL options to handle redirects and cookies curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); // Follow 3xx redirects automatically curl_setopt($ch, CURLOPT_COOKIEJAR, 'cookie.txt'); // Save cookies to a file for persistence curl_setopt($ch, CURLOPT_COOKIEFILE, 'cookie.txt'); // Load saved cookies for subsequent requests curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); // Return content as a string instead of printing it // Execute the request and get the HTML content $htmlContent = curl_exec($ch); curl_close($ch); // Load content into simple_html_dom require_once('simple_html_dom.php'); $html = str_get_html($htmlContent); // Now you can parse the HTML normally, e.g., grab location: $location = $html->find('.your-location-selector', 0)->plaintext; // Replace with the actual CSS selector from the page
This setup will handle redirects and retain cookies between requests, so your first request will behave exactly like the successful second one you're seeing.
2. Getting the Missing location Field from RSS
Since the RSS feed doesn't include location data, you'll need to use the link from each RSS item to fetch the full job detail page, then extract the location using simple_html_dom (with the cURL setup above).
Here's how to combine RSS parsing with detail page scraping:
// First, fetch the RSS feed $xml = simplexml_load_file('https://www.globalsportsjobs.com/index.php/page/adv_rss/command/getfeed/feed/4a2a139494ad1c02b7df790877737575'); $parsed_results_array = array(); // Initialize cURL once for reusability $ch = curl_init(); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); curl_setopt($ch, CURLOPT_COOKIEJAR, 'cookie.txt'); curl_setopt($ch, CURLOPT_COOKIEFILE, 'cookie.txt'); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); foreach($xml as $entry) { foreach($entry->item as $item) { $items['title'] = (string) $item->title; $items['description'] = (string) $item->description; $items['link'] = (string) $item->link; // Fetch the job detail page to extract location curl_setopt($ch, CURLOPT_URL, $items['link']); $detailHtml = curl_exec($ch); $detailDom = str_get_html($detailHtml); // Extract location (replace the selector with the actual one from the job page) $items['location'] = $detailDom->find('.job-location-element', 0)->plaintext ?? 'Location not found'; $parsed_results_array[] = $items; // Clean up the DOM object to avoid memory leaks $detailDom->clear(); unset($detailDom); } } curl_close($ch); // Now $parsed_results_array contains all fields, including location!
Just replace .job-location-element with the actual CSS selector for the location element on the job detail page (you can find this using your browser's dev tools by inspecting the location text).
Quick Tips
- Ensure the
cookie.txtfile is writable by your server (usechmod 644 cookie.txtif needed) so cURL can save/load cookies. - Add error handling for cURL requests (check
curl_error($ch)ifcurl_execreturns false) to catch network issues. - Respect the site's
robots.txtand terms of service—add small delays between requests if you're scraping multiple pages to avoid overwhelming the server.
内容的提问来源于stack exchange,提问作者William

