如何抓取无SKU的MichaelKors页面?Scrapy新手技术求助
Hey there! Let's walk through the key things you should check to fix this scraping snag on Michael Kors' site—since you're new to Scrapy, these tips will help you troubleshoot and adapt to different page structures.
1. Verify the Real HTML Structure (Don't Trust What You See at First Glance)
First, open your browser's DevTools (F12) and head to the Elements tab. Use the selector tool (the arrow icon in the top-left) to click directly on the description text you want to scrape. This will show you the actual HTML markup for that element:
- Double-check the class name: Your XPath uses
look-description-desktop hide-on-mobile, but maybe the actual class is split differently (likelook-description desktop hide-on-mobile) or has a typo. If it's multiple classes, adjust your XPath to usecontains()for better flexibility:desc = response.xpath('//p[contains(@class, "look-description") and contains(@class, "desktop")]/text()').getall() - Also, check if the text is nested inside another tag (like a
<span>within the<p>) that you're missing in your selector.
2. Check if Content is Dynamically Loaded via JavaScript
Since you already used window.initial_state for other pages, you know Michael Kors relies heavily on JS-rendered content. For this combination page:
- Go to the Network tab in DevTools, filter for XHR/Fetch requests, and refresh the page. Look for API calls that might return the look's description (search for keywords like "look", "description", or the page's ID
L-MSTR101179). - In the browser's Console, try typing
document.querySelector('p.look-description-desktop'). If it returnsnull, that means the element doesn't exist in the static HTML—it's being added by JS after page load. You'll need to either extract the data directly from the relevant API response, or use a tool likeScrapy-SplashorPlaywrightto render the JS before scraping.
3. Try Alternative Locators (Don't Rely Solely on Class Names)
Class names can change frequently (especially on e-commerce sites for A/B testing or design updates). Switch to more stable ways to target elements:
- Use parent container classes: If the description is inside a
<div class="look-details-container">, your XPath could be:desc = response.xpath('//div[@class="look-details-container"]/p/text()').getall() - Use text context: If there's a preceding label like "Look Description:", target it with:
desc = response.xpath('//p[text()="Look Description:"]/following-sibling::p/text()').getall() - Give CSS selectors a try—they're often more readable for multi-class elements:
desc = response.css('p.look-description-desktop.hide-on-mobile::text').getall()
4. Debug with Scrapy Shell (Your New Best Friend)
As a Scrapy newbie, the shell is invaluable for testing selectors in real time. Run this command in your terminal:
scrapy shell https://www.michaelkors.com/zip-hoodie-embellished-skirt-manhattan-crossbody-goldie-moto-boot/_/L-MSTR101179
Once inside, test your XPath/CSS selectors directly. If they return empty, tweak them step by step (e.g., first get the parent div, then narrow down to the <p>) to find where the issue is.
5. Watch Out for Anti-Scraping Measures
Michael Kors might block or modify content for non-browser requests. Try these quick fixes:
- Update your
USER_AGENTinsettings.pyto match your browser's real user agent (you can find this in DevTools' Network tab, under any request's headers). Example:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' - Check if you need to send cookies with your request. Some e-commerce sites require session cookies to render full content—you can scrape cookies from the homepage first and reuse them in your request to the target page.
6. Look for Other Window Variables
Since window.initial_state worked for single-product pages, check if this combination page stores data in other window variables. In the browser's Console, type window. and scroll through the autocomplete suggestions—look for names like lookData, collectionDetails, or productBundle that might contain the description and other info you need.
内容的提问来源于stack exchange,提问作者Newbie

