Laravel中获取外部网站figure元素内图片URL的问题求助
Fixing PHP DOM/XPath to Extract Product Image URL from Matches Fashion
Let's break down why your current code isn't working and how to fix it to get the image URL you need:
Issues with Your Current Code
- Incorrect XPath Syntax: You used
"in your XPath selector, which is HTML entity encoding — this won't work in PHP's XPath evaluation. You should use regular single or double quotes properly. - Targeting the Wrong Element: The
<figure class="iiz">element itself isn't an image; it's a container for the lazy-loaded image. Matches Fashion uses the Image in Zoom (iiz) library, which stores the actual image URL in a data attribute on the figure or its child img tag.
Step-by-Step Solution
1. Handle Potential Anti-Scraping Measures
First, many retail sites block default HTTP requests from DOMDocument (since it uses a generic User-Agent). We'll use file_get_contents with a browser-like User-Agent to fetch the page content instead.
2. Correct XPath to Target the Image URL
For the target product page, the main product image URL is stored in the data-iiz-src attribute of the <figure class="iiz"> element. Alternatively, you can target the child <img> tag's data-src attribute (both work, but data-iiz-src gives the high-res version).
Fixed Code
// Set up a browser-like User-Agent to avoid being blocked $opts = [ 'http' => [ 'header' => 'User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' ] ]; $context = stream_context_create($opts); // Fetch the page content $html = file_get_contents('https://www.matchesfashion.com/products/Adidas-By-Stella-McCartney-Metallic-zebra-print-Primegreen-leggings-1424516', false, $context); $doc = new DOMDocument; $doc->preserveWhiteSpace = false; $doc->strictErrorChecking = false; $doc->recover = true; $doc->loadHTML($html); $xpath = new DOMXPath($doc); // Option 1: Get high-res image from figure's data-iiz-src attribute $imageUrl = $xpath->evaluate('string(//figure[@class="iiz"]/@data-iiz-src)'); // Option 2: Get image from child img's data-src attribute (fallback) // $imageUrl = $xpath->evaluate('string(//figure[@class="iiz"]/img/@data-src)'); if (!empty($imageUrl)) { echo "Image URL: " . $imageUrl; } else { echo "Image URL not found — check if the page structure has changed."; }
Explanation
- User-Agent Spoofing: The custom HTTP header makes the request look like it's coming from a Chrome browser, which reduces the chance of being blocked by Matches Fashion's anti-scraping systems.
- XPath Adjustments:
//figure[@class="iiz"]/@data-iiz-srcdirectly targets the data attribute on the figure that holds the high-resolution image URL.- The fallback option targets the child img tag's
data-srcif the first selector fails (page structures can change over time).
- Error Handling: The code checks if the URL is retrieved successfully, so you can debug if the page structure has been updated.
内容的提问来源于stack exchange,提问作者Devin Y
相关产品推荐
相关产品推荐

