You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取返回0结果求助:无法获取目标网站项目名称

解决Scrapy爬取F6s项目名称的问题

Hey there! Let's break down why your Scrapy spider isn't crawling any pages or items, and fix it step by step.

1. First: Bypass the Robots.txt Restriction

Your log clearly points to the biggest issue right now:

2018-05-15 14:20:13 [scrapy.downloadermiddlewares.robotstxt] DEBUG: Forbidden by robots.txt: <GET https://www.f6s.com/programs?type[]=accelerator&sort=open>

Scrapy automatically follows a website's robots.txt rules by default, and this site is blocking crawlers from accessing that page. To fix this, open your settings.py file and find the line for ROBOTSTXT_OBEY—change it to:

ROBOTSTXT_OBEY = False

2. Fix Your Selector & Output Format

Right now, your selector might not be targeting the right elements, and using extract() will return a list instead of the single string format you want ('Program': 'Program Name').

Step 2.1: Verify the Selector

Use your browser's DevTools (F12) to inspect the project name elements on the target page. Double-check that the selector div.title a.action.main.no.line::text actually matches the text you want. If the page structure is different, adjust the selector accordingly.

Step 2.2: Adjust the Parse Method

Modify your parse function to get a single, clean string instead of a list:

def parse(self, response):
    # Loop through each title container
    for title in response.css('div.title'):
        # Extract the first matching text (avoid lists)
        program_name = title.css('div.title a.action.main.no.line::text').extract_first()
        # Only yield if we found a name (skip empty entries)
        if program_name:
            yield {
                'Program': program_name.strip()  # Strip extra whitespace
            }

3. Make Sure You're Running the Right Spider

Your spider has the name "Company" (defined in name = "Company"), so when you run your crawler, use this command:

scrapy crawl Company

Also, double-check your spiders directory—you have byub.py and F6sSpider.py there. Ensure the code you modified is in the correct file (the one containing the CompanySpider class).

4. Extra Tips to Avoid Future Issues

  • Check if you're actually getting the page: Add a quick print statement in parse to confirm the response is valid:
    def parse(self, response):
        print("Received response! Page length:", len(response.text))
        # Rest of your code here
    
  • Set a User-Agent: Many sites block default Scrapy user agents. Add this to settings.py to mimic a real browser:
    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
    

内容的提问来源于stack exchange,提问作者paperelephant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:01:46