Scrapy运行爬虫报ValueError: Missing scheme in request url: h如何解决
问题修复方案
核心报错原因:allowed_domains参数配置错误
Scrapy的allowed_domains字段要求传入仅包含域名的列表,不允许带http/https协议头,也不能直接传字符串。你当前代码里写的是allowed_domains = 'https://weather.com',属于字符串类型,Scrapy内部遍历该字段时会把字符串按单个字符拆分,第一个字符是h,就会触发你遇到的ValueError: Missing scheme in request url: h报错。
修正写法:
allowed_domains = ['weather.com']
其他需要修正的代码问题
- XPath语法错误:你提取城市的XPath写法有误,
.text()是错误写法,正确的XPath文本提取应该用/text()
错误写法:
修正写法:city = response.xpath('//h1[contains(@class,"location")].text()').get()city = response.xpath('//h1[contains(@class,"location")]/text()').get() - 回调方法适配:Scrapy默认的请求回调方法是
parse,你当前定义的是parse_url,直接把方法名改为parse即可适配默认逻辑,无需额外手动指定回调。
修正后的完整代码
import scrapy import re from weather_parent.weather_spider.items import WeatherItem class WeatherSpiderSpider(scrapy.Spider): name = "weather_spider2" allowed_domains = ['weather.com'] start_urls = ['https://weather.com/en-MT/weather/today/l/bf01d09009561812f3f95abece23d16e123d8c08fd0b8ec7ffc9215c0154913c'] def parse(self, response): city = response.xpath('//h1[contains(@class,"location")]/text()').get() temp = response.xpath('//span[@data-testid="TemperatureValue"]/text()').get() air_quality = response.xpath('//span[@data-testid="AirQualityCategory"]/text()').get() cond = response.xpath('//div[@data-testid="wxPhrase"]/text()').get() item = WeatherItem() item["city"] = city item["temp"] = temp item["air_quality"] = air_quality item["cond"] = cond yield item
修正后重新执行scrapy crawl weather_spider2 -o output.json即可正常运行。
内容的提问来源于stack exchange,提问作者nickcarter
相关产品推荐
相关产品推荐

