You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy已启用Cookie仍提示“请启用Cookie”问题排查求助

Fixing the "Please Enable Cookies" Block in Your Scrapy Shoe Scraper

Hey there, let's break down why you're hitting that frustrating "Please enable cookies" wall even with COOKIES_ENABLED=True and COOKIES_DEBUG=True set up. Here are the most common culprits and actionable fixes:

1. The Site Needs JavaScript to Set/Validate Cookies

Most modern e-commerce sites use JavaScript to generate critical cookies (like session tokens or anti-bot checks) that Scrapy can't handle natively—since Scrapy doesn't execute JS. That's likely why your first 3-4 requests work (you get initial server-side cookies) but subsequent ones fail once the site expects JS-generated cookies.

Fix: Integrate a Headless Browser

Use scrapy-selenium or scrapy-playwright to simulate a real browser that runs JS and manages cookies properly. Here's a quick setup with scrapy-selenium:

  • Install the package first:
    pip install scrapy-selenium
    
  • Update your settings.py to enable the middleware:
    DOWNLOADER_MIDDLEWARES = {
        'scrapy_selenium.SeleniumMiddleware': 800
    }
    
    SELENIUM_DRIVER_NAME = 'chrome'
    SELENIUM_DRIVER_EXECUTABLE_PATH = '/path/to/your/chromedriver'
    SELENIUM_DRIVER_ARGUMENTS = ['--headless=new']  # Runs in background without opening a window
    
  • Modify your spider to use SeleniumRequest instead of regular scrapy.Request:
    from scrapy_selenium import SeleniumRequest
    
    # Inside your parse method:
    yield SeleniumRequest(url=link, callback=self.parse_page, dont_filter=True)
    

This mimics a real Chrome browser, letting the site set and validate cookies exactly like a human user would.

Some sites require you to first visit the main homepage to get foundational session cookies before accessing specific pages like the upcoming shoes list. Your current code jumps straight to the ?page=1 URL, which might miss this critical step.

Fix: Start with the Main Site

Adjust your spider to fetch the homepage first, then proceed to the shoes page once cookies are set:

def start_requests(self):
    # First get initial cookies from the main site
    yield scrapy.Request('https://www.kicksonfire.com/', callback=self.start_scraping)

def start_scraping(self, response):
    # Now request the upcoming shoes page with the cookies we just collected
    yield scrapy.Request('https://www.kicksonfire.com/app/upcoming?page=1', callback=self.parse)

Check your COOKIES_DEBUG logs after this change—you should see the site setting initial session cookies that carry over to subsequent requests.

3. Your User-Agent is Flagged as a Bot

Scrapy's default User-Agent (Scrapy/VERSION (+https://scrapy.org)) is a dead giveaway for bots. Many sites will block you or refuse to set valid cookies if they detect this.

Fix: Spoof a Real Browser User-Agent

Either set a static, common UA in settings.py:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'

Or use the scrapy-user-agents middleware to rotate random UAs:

pip install scrapy-user-agents

Update settings.py to enable it:

DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None,
    'scrapy_user_agents.middlewares.RandomUserAgentMiddleware': 400,
}

With COOKIES_DEBUG=True enabled, closely check your logs to confirm:

  • The site is setting cookies in the initial response
  • Those cookies are being sent back in subsequent requests
  • No cookies are being dropped or modified by Scrapy's cookie jar

If cookies aren't passing correctly, you can manually attach them to requests:

def parse(self, response):
    # Extract cookies from the current response
    cookies = response.headers.getlist('Set-Cookie')
    shoe_links = response.xpath('//div/a/@href').extract()[1:]
    for link in shoe_links:
        yield scrapy.Request(link, callback=self.parse_page, dont_filter=True, cookies=cookies)

Quick Test Tip

Manually visit the site in incognito mode, open DevTools > Network tab, and watch the cookies being set when you load the upcoming shoes page and individual shoe links. Compare this to your Scrapy logs—this will help you spot missing cookies or headers the site expects.

内容的提问来源于stack exchange,提问作者Ficofi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 08:17:34