Scrapy爬取返回0结果求助:无法获取目标网站项目名称
Hey there! Let's break down why your Scrapy spider isn't crawling any pages or items, and fix it step by step.
1. First: Bypass the Robots.txt Restriction
Your log clearly points to the biggest issue right now:
2018-05-15 14:20:13 [scrapy.downloadermiddlewares.robotstxt] DEBUG: Forbidden by robots.txt: <GET https://www.f6s.com/programs?type[]=accelerator&sort=open>
Scrapy automatically follows a website's robots.txt rules by default, and this site is blocking crawlers from accessing that page. To fix this, open your settings.py file and find the line for ROBOTSTXT_OBEY—change it to:
ROBOTSTXT_OBEY = False
2. Fix Your Selector & Output Format
Right now, your selector might not be targeting the right elements, and using extract() will return a list instead of the single string format you want ('Program': 'Program Name').
Step 2.1: Verify the Selector
Use your browser's DevTools (F12) to inspect the project name elements on the target page. Double-check that the selector div.title a.action.main.no.line::text actually matches the text you want. If the page structure is different, adjust the selector accordingly.
Step 2.2: Adjust the Parse Method
Modify your parse function to get a single, clean string instead of a list:
def parse(self, response): # Loop through each title container for title in response.css('div.title'): # Extract the first matching text (avoid lists) program_name = title.css('div.title a.action.main.no.line::text').extract_first() # Only yield if we found a name (skip empty entries) if program_name: yield { 'Program': program_name.strip() # Strip extra whitespace }
3. Make Sure You're Running the Right Spider
Your spider has the name "Company" (defined in name = "Company"), so when you run your crawler, use this command:
scrapy crawl Company
Also, double-check your spiders directory—you have byub.py and F6sSpider.py there. Ensure the code you modified is in the correct file (the one containing the CompanySpider class).
4. Extra Tips to Avoid Future Issues
- Check if you're actually getting the page: Add a quick print statement in
parseto confirm the response is valid:def parse(self, response): print("Received response! Page length:", len(response.text)) # Rest of your code here - Set a User-Agent: Many sites block default Scrapy user agents. Add this to
settings.pyto mimic a real browser:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
内容的提问来源于stack exchange,提问作者paperelephant

