如何使用Scrapy获取所有锚点href?解决Scrapy Shell中无法提取图片href的问题
Hey there! Let's fix your issue step by step.
First, your initial code using response.css("div.slick-track").getall() just returns the raw HTML of that div container—it doesn't extract the actual image links. To get the image URLs, you need to target the img tags inside that container and pull their source attributes.
Extracting Product Image URLs from the Trendyol Page
On that specific Trendyol product page, the product images are nested inside the slick-track div, usually in img tags. Many e-commerce sites use lazy loading, so the image URL might be in the data-src attribute instead of the standard src. Here's how to get them:
- In Scrapy Shell, after loading the page, run this to get all image sources:
If some images use# Get image URLs from data-src (lazy loaded) response.css("div.slick-track img::attr(data-src)").getall()srcinstead, you can combine both selectors to cover all cases:
This will give you a clean list of all product image URLs on that page.# Combine data-src and src attributes image_urls = response.css("div.slick-track img::attr(data-src)").getall() + response.css("div.slick-track img::attr(src)").getall() # Filter out empty strings image_urls = [url for url in image_urls if url]
Collecting All Anchor HREFs in General
If you want to extract every anchor tag's href from any page, use this simple selector:
# Get all anchor hrefs all_hrefs = response.css("a::attr(href)").getall()
Note that some href values might be relative paths (like /product/123). To convert them to absolute URLs, use response.urljoin() for each:
# Convert relative hrefs to absolute URLs absolute_hrefs = [response.urljoin(href) for href in all_hrefs if href]
Quick Tip for Debugging
If you're unsure about the right selector, use Scrapy Shell's view(response) command to open the page in your browser and inspect the elements directly. This helps you confirm which attributes hold the data you need.
内容的提问来源于stack exchange,提问作者Fazlul Hoque Sawrav

