You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取初创企业信息无输出问题求助

Scrapy爬虫无输出问题排查与修复

我正在构建一个爬取初创企业信息的Scrapy爬虫,脚本计划访问目标网站并将信息存入字典,但始终无法获取输出结果。以下是我的代码:

import scrapy


class StartupsSpider(scrapy.Spider):
    name = 'startups'
    #name of the spider

    allowed_domains = ['www.bmwk.de/Navigation/DE/InvestDB/INVEST-DB_Liste/investdb.html']
    #list of allowed domains

    start_urls = ['https://bmwk.de/Navigation/DE/InvestDB/INVEST-DB_Liste/investdb.html']
    #starting url

    def parse(self, response):
        
        startups = response.xpath('//*[contains(@class,"card-link-overlay")]/@href').getall()
        #parse initial start URL for the specific startup URL

        for startup in startups:
            
            absolute_url =  response.urljoin(startup)

            yield scrapy.Request(absolute_url, callback=self.parse_startup)
            #parse the actual startup information

        next_page_url = response.xpath('//*[@class ="pagination-link"]/@href').get()
        #link to next page
        
        absolute_next_page_url = response.urljoin(next_page_url)
        #go through all pages on start URL
        yield scrapy.Request(absolute_next_page_url)
    
    def parse_startup(self, response):
    #get information regarding startup
        startup_name = response.css('h1::text').get()
        startup_hompage = response.xpath('//*[@class="document-info-item"]/a/@href').get()
        startup_description = response.css('div.document-info-item::text')[16].get()
        branche = response.css('div.document-info-item::text')[4].get()
        founded = response.xpath('//*[@class="date"]/text()')[0].getall()
        employees = response.css('div.document-info-item::text')[9].get()
        capital = response.css('div.document-info-item::text')[11].get()
        applied_for_invest = response.xpath('//*[@class="date"]/text()')[1].getall()

        contact_name = response.css('p.card-title-subtitle::text').get()
        contact_phone = response.css('p.tel > span::text').get()
        contact_mail = response.xpath('//*[@class ="person-contact"]/p/a/span/text()').get()
        contact_address_street = response.xpath('//*[@class ="adr"]/text()').get()
        contact_address_plz = response.xpath('//*[@class ="locality"]/text()').getall()
        contact_state = response.xpath('//*[@class ="country-name"]/text()').get()

        yield{'Startup':startup_name,
              'Homepage': startup_hompage,
              'Description': startup_description,
              'Branche': branche,
              'Gründungsdatum': founded,
              'Anzahl Mitarbeiter':employees,
              'Kapital Bedarf':capital,
              'Datum des Förderbescheids':applied_for_invest,
              'Contact': contact_name,
              'Telefon':contact_phone,
              'E-Mail':contact_mail,
              'Adresse': contact_address_street + contact_address_plz + contact_state}

问题排查与修复方案

1. 域名配置错误

allowed_domains需填写根域名,而非完整URL路径。原配置会导致后续请求被Scrapy的域名过滤机制拦截,修改为:

allowed_domains = ['bmwk.de']

2. 分页逻辑未做空值判断

爬取到最后一页时,next_page_url会返回None,直接调用response.urljoin会抛出异常中断爬虫。需添加判断逻辑,并显式指定回调函数:

next_page_url = response.xpath('//*[@class="pagination-link"]/@href').get()
if next_page_url:
    absolute_next_page_url = response.urljoin(next_page_url)
    yield scrapy.Request(absolute_next_page_url, callback=self.parse)

3. 数据提取依赖固定索引,稳定性差

原代码用[4]、[16]这类固定索引定位元素,页面结构稍有变化就会失效。建议通过内容关联的选择器定位,例如提取行业信息:

# 替换原branche提取逻辑
branche = response.xpath('//div[contains(text(), "Branche")]/following-sibling::div/text()').get().strip()

其他字段(如员工数、成立时间)也需调整为类似的精准选择器,避免依赖索引。

4. 字符串拼接类型错误

contact_address_plz是getall()返回的列表,直接与字符串拼接会触发TypeError,需先转为字符串:

# 替换原地址相关提取逻辑
contact_address_street = response.xpath('//*[@class="adr"]/text()').get().strip() if response.xpath('//*[@class="adr"]/text()').get() else ''
contact_address_plz = ' '.join(response.xpath('//*[@class="locality"]/text()').getall()).strip()
contact_state = response.xpath('//*[@class="country-name"]/text()').get().strip() if response.xpath('//*[@class="country-name"]/text()').get() else ''

# 地址拼接
'Adresse': f"{contact_address_street} {contact_address_plz} {contact_state}".strip()

5. User-Agent配置缺失

部分网站会拦截Scrapy默认的User-Agent,导致请求被拒绝。需在settings.py中添加:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'

内容的提问来源于stack exchange,提问作者Martin Bodzwag

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 21:10:45