You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫仅循环一次:无法从weather.com获取全部天气数据

Scrapy爬虫循环仅执行一次的问题分析与修复

问题根源

你的代码存在三个关键问题,导致最终只输出一条数据,看起来像循环只执行了一次:

  1. Item实例化位置错误
    你在循环外创建了唯一的weather_item实例,循环中只是不断覆盖这个item的字段值,最后仅yield一次,自然只会得到最后一次循环修改后的结果。

  2. XPath使用绝对路径且固定ID
    Rain_chance的XPath用了//开头的绝对路径,还硬编码了id="detailIndex1",这会导致每次循环都抓取页面中第一个匹配该ID的元素,而不是当前循环项i下的对应数据。

  3. 目标选择器范围错误
    response.css('div.DailyForecast--DisclosureList--nosQS')匹配的是整个预报列表的容器,而非单个每日预报项,所以x本身可能只有一个元素,循环自然只执行一次。

修复后的代码

class WeatherSpider(scrapy.Spider):
    name = "weather"
    allowed_domains = ["weather.com"]
    start_urls = ["https://weather.com/weather/tenday/l/Homewood+AL?canonicalCityId=ee632098bb6c46fd10d48efa5cf1550a9e8a2d593da04926653f01e690d40ba2"]

    def parse(self, response):
        # 选择单个每日预报项,而非整个容器
        daily_forecasts = response.css('div.DailyForecast--DisclosureList--nosQS div.DailyForecast--DisclosureListItem--2nd_j')
        
        for forecast in daily_forecasts:
            # 每次循环创建新的Item实例
            weather_item = WeatherApiItem()
            weather_item['Hight_temp'] = forecast.css('span.DetailsSummary--highTempValue--3PjlX::text').get()
            weather_item['Low_temp'] = forecast.css('span.DetailsSummary--lowTempValue--2tesQ::text').get()
            # 使用相对路径的XPath,基于当前循环项forecast定位
            weather_item['Rain_chance'] = forecast.xpath('.//span[@data-testid="PercentageValue"]/text()').get()
            yield weather_item

关键修改说明

  • Item实例化移至循环内:确保每次循环生成独立的Item对象,避免字段覆盖。
  • 修正选择器范围:直接定位到单个每日预报项(DailyForecast--DisclosureListItem--2nd_j),保证循环遍历所有日期。
  • 使用相对路径提取降雨概率:XPath以.开头,限定在当前循环项的节点内查找,同时用更通用的属性选择器替代硬编码ID,提高稳定性。

内容的提问来源于stack exchange,提问作者Ethan Koch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 09:35:57