You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Laravel中获取外部网站figure元素内图片URL的问题求助

Fixing PHP DOM/XPath to Extract Product Image URL from Matches Fashion

Let's break down why your current code isn't working and how to fix it to get the image URL you need:

Issues with Your Current Code

  • Incorrect XPath Syntax: You used " in your XPath selector, which is HTML entity encoding — this won't work in PHP's XPath evaluation. You should use regular single or double quotes properly.
  • Targeting the Wrong Element: The <figure class="iiz"> element itself isn't an image; it's a container for the lazy-loaded image. Matches Fashion uses the Image in Zoom (iiz) library, which stores the actual image URL in a data attribute on the figure or its child img tag.

Step-by-Step Solution

1. Handle Potential Anti-Scraping Measures

First, many retail sites block default HTTP requests from DOMDocument (since it uses a generic User-Agent). We'll use file_get_contents with a browser-like User-Agent to fetch the page content instead.

2. Correct XPath to Target the Image URL

For the target product page, the main product image URL is stored in the data-iiz-src attribute of the <figure class="iiz"> element. Alternatively, you can target the child <img> tag's data-src attribute (both work, but data-iiz-src gives the high-res version).

Fixed Code

// Set up a browser-like User-Agent to avoid being blocked
$opts = [
    'http' => [
        'header' => 'User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    ]
];
$context = stream_context_create($opts);

// Fetch the page content
$html = file_get_contents('https://www.matchesfashion.com/products/Adidas-By-Stella-McCartney-Metallic-zebra-print-Primegreen-leggings-1424516', false, $context);

$doc = new DOMDocument; 
$doc->preserveWhiteSpace = false; 
$doc->strictErrorChecking = false; 
$doc->recover = true; 
$doc->loadHTML($html); 

$xpath = new DOMXPath($doc); 

// Option 1: Get high-res image from figure's data-iiz-src attribute
$imageUrl = $xpath->evaluate('string(//figure[@class="iiz"]/@data-iiz-src)');

// Option 2: Get image from child img's data-src attribute (fallback)
// $imageUrl = $xpath->evaluate('string(//figure[@class="iiz"]/img/@data-src)');

if (!empty($imageUrl)) {
    echo "Image URL: " . $imageUrl;
} else {
    echo "Image URL not found — check if the page structure has changed.";
}

Explanation

  • User-Agent Spoofing: The custom HTTP header makes the request look like it's coming from a Chrome browser, which reduces the chance of being blocked by Matches Fashion's anti-scraping systems.
  • XPath Adjustments:
    • //figure[@class="iiz"]/@data-iiz-src directly targets the data attribute on the figure that holds the high-resolution image URL.
    • The fallback option targets the child img tag's data-src if the first selector fails (page structures can change over time).
  • Error Handling: The code checks if the URL is retrieved successfully, so you can debug if the page structure has been updated.

内容的提问来源于stack exchange,提问作者Devin Y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 10:54:09