如何用Python Scrapy XPath访问text/x-magento-init脚本标签并提取商品图片
Hey there! Let's figure out how to get those two product images even if you can't access the <script type="text/x-magento-init"> tag. Here are a few practical approaches you can try:
1. 解析页面图片元素并调整URL获取高清图
Most e-commerce sites (including Lidl's) serve resized images with dimension parameters in the URL. Start by locating the <img> tags for the product images using browser dev tools or a parser like BeautifulSoup.
For example, if you find an image URL like:https://sortiment.lidl.ch/media/catalog/product/cache/.../path/to/image-200x200.jpg
You can often replace the size suffix (like 200x200) with a larger dimension (e.g., 1000x1000) or remove the cache/size part entirely to get the full-resolution image. Test different variations to match Lidl's image CDN pattern.
2. 提取Open Graph元标签中的图片链接
Check the page's HTML <head> section for Open Graph (OG) tags. These are used for content previews and usually include a high-quality product image. Look for tags like:
<meta property="og:image" content="https://full-size-image-url.jpg">
You can scrape this tag directly—no need to interact with the magento-init script. Sometimes there's also a og:image:secure_url tag for HTTPS links.
3. 使用无头浏览器渲染页面获取动态内容
If the text/x-magento-init script loads dynamically via JavaScript (common in Magento 2 sites), a static HTML parser won't catch it. Use a headless browser tool like Playwright or Puppeteer to fully render the page, wait for all JS to execute, then extract the script content.
Here's a simplified Puppeteer example:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto('https://sortiment.lidl.ch/de/papier-tragetasche-fsc-0123997.html', { waitUntil: 'networkidle2' }); // Extract the magento-init script content const magentoInitData = await page.evaluate(() => { const script = document.querySelector('script[type="text/x-magento-init"]'); return script ? JSON.parse(script.textContent) : null; }); // Pull full images from gallery configuration if (magentoInitData) { const galleryConfig = magentoInitData['[data-gallery-role="gallery"]']['mage/gallery']; const fullImages = galleryConfig.data.map(item => item.img); console.log(fullImages); } await browser.close(); })();
Once you have the magentoInitData, look for the gallery configuration—it will contain an array of image objects with full-size URLs.
4. 抓取商品媒体API请求
Open your browser's DevTools (F12), go to the Network tab, reload the page, and filter for XHR or Fetch requests. Look for endpoints related to product media (e.g., something like /rest/V1/products/0123997/media—the SKU 0123997 is in your URL).
These API responses usually return a JSON array with all product images, including full URLs, alt text, and image types. You can replicate this request in your scraper or manually copy the URLs from the response.
内容的提问来源于stack exchange,提问作者akmal Khan

