Scrapy Splash:选择下拉选项并截图技术问题求助
Fixing Scrapy Splash Dropdown Selection & Screenshot for Chewy
Hey Alex, let’s get this sorted out! The problem you’re hitting is that Chewy uses a JavaScript-powered custom dropdown (not a native <select> element), so regular Scrapy requests or even SplashFormRequest won’t cut it—we need to simulate actual user clicks with a Lua script in Splash. Here’s a working solution tailored to your target page:
Working Code Example
import scrapy from scrapy_splash import SplashRequest class ChewyDogToySpider(scrapy.Spider): name = "chewy_dog_toy" target_url = "https://www.chewy.com/kong-classic-dog-toy-large/dp/38485" def start_requests(self): # Lua script to handle dropdown interaction and screenshot splash_script = """ function main(splash, args) -- Mimic a real browser user agent to avoid being blocked splash:set_user_agent("Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") -- Load the target page assert(splash:go(args.url)) -- Wait for the size dropdown trigger to be visible assert(splash:wait_for_selector('.js-select-dropdown')) -- Click the dropdown to reveal size options local dropdown = splash:select('.js-select-dropdown') dropdown:mouse_click() -- Wait for the "Small" option to appear assert(splash:wait_for_selector('.js-select-option:contains("Small")')) -- Click the Small option local small_option = splash:select('.js-select-option:contains("Small")') small_option:mouse_click() -- Give the page a second to update (e.g., refresh selected size text) splash:wait(1.5) -- Capture the full page screenshot local screenshot = splash:png() -- Return both screenshot and updated HTML for verification return { screenshot = screenshot, html = splash:html() } end """ yield SplashRequest( url=self.target_url, callback=self.parse_results, endpoint="execute", args={ "lua_source": splash_script, "wait": 2, # Extra wait time for initial page load } ) def parse_results(self, response): # Save the screenshot to a local file with open("chewy_kong_small.png", "wb") as img_file: img_file.write(response.data["screenshot"]) self.log("✅ Screenshot saved as chewy_kong_small.png") # Optional: Verify the size was actually changed selector = scrapy.Selector(text=response.data["html"]) current_size = selector.css(".js-selected-text::text").get().strip() self.log(f"🔍 Confirmed selected size: {current_size}")
Key Explanations
- Lua Script is Critical: Chewy’s dropdown isn’t a native
<select>, so we need to simulate the exact user flow (click dropdown → select option) using Splash’s browser automation capabilities. - Selector Targets: I used Chewy’s actual frontend classes (
.js-select-dropdownfor the trigger,.js-select-option:contains("Small")for the option). You can double-check these with your browser’s dev tools if they ever change. - Waiting Steps:
wait_for_selectorensures we only interact with elements once they’re fully loaded, and the finalwait(1.5)gives the page time to update the selected size before capturing the screenshot. - User Agent: Setting a real user agent helps avoid being flagged as a bot by Chewy’s anti-scraping measures.
Quick Debug Tips
- If the dropdown doesn’t click: Verify the CSS selector for the dropdown trigger (right-click the element in dev tools → Copy → Copy selector).
- If the screenshot shows the wrong size: Increase the final
waittime (try 2-3 seconds) — some pages take longer to update dynamic content. - Make sure Splash is running locally (default at
http://localhost:8050) — you’ll need to have it installed and started before running the spider.
内容的提问来源于stack exchange,提问作者Alex D
相关产品推荐
相关产品推荐

