You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Scrapy爬取第三方软件版本号并构建版本仓库?(附Firefox示例)

Fixing Your Scrapy Spider to Grab Firefox's Latest Version

Hey there! Let's work through your Scrapy spider issue for fetching Firefox's current version number. First, I noticed your parse method was cut off, so I'll fill that in, fix some key issues, and share more reliable approaches tailored to this specific task.

Full Working Code Example

import scrapy
import html2text
import re
from scrapy.selector import Selector  # HtmlXPathSelector is deprecated in newer Scrapy versions, though still compatible with Python 2.7's Scrapy builds

class MozillaSpider(scrapy.Spider):
    name = 'mozilla'
    allowed_domains = ['mozilla.org']  # Heads up: your original domain was wrong (it's mozilla.org, not .com)
    start_urls = ['https://www.mozilla.org/en-US/firefox/notes/']

    def parse(self, response):
        # Method 1: CSS Selector (targets the version heading directly)
        version_css = response.css('.c-release-version h2::text').extract_first()
        if version_css:
            pure_version = version_css.strip().split(' ')[-1]
            self.logger.info("Got Firefox version via CSS: %s" % pure_version)
            yield {'version': pure_version}

        # Method 2: XPath (fallback if CSS selector breaks)
        version_xpath = response.xpath('//div[contains(@class, "c-release-version")]/h2/text()').extract_first()
        if version_xpath:
            pure_version = version_xpath.strip().split(' ')[-1]
            self.logger.info("Got Firefox version via XPath: %s" % pure_version)
            yield {'version': pure_version}

        # Method 3: Plain text parsing with html2text (good for layout changes)
        h = html2text.HTML2Text()
        h.ignore_links = True
        plain_text = h.handle(response.text)
        
        version_match = re.search(r'Firefox (\d+\.\d+\.?\d*)', plain_text)
        if version_match:
            pure_version = version_match.group(1)
            self.logger.info("Got Firefox version via text match: %s" % pure_version)
            yield {'version': pure_version}

Key Fixes & Explanations

  • Domain Correction: Your original allowed_domains used mozilla.com, but the target site is mozilla.org—this would have caused Scrapy to block your request, so that's a critical fix.
  • Selector Updates: HtmlXPathSelector is deprecated in newer Scrapy versions. While Python 2.7's compatible Scrapy (v1.x) still supports it, using response.css() or response.xpath() directly is cleaner.
  • Version Targeting: Firefox's latest version lives in a .c-release-version container's <h2> tag on the release notes page. I used browser dev tools (F12) to confirm this selector works right now.
  • Python 2.7 Compatibility: I avoided Python 3+ features like f-strings since you're on 2.7—stick with % formatting or .format() for strings.

Even More Reliable Approach (API Instead of HTML)

If you want to avoid breaking your spider every time Mozilla updates their page layout, use their official version API instead. It returns structured JSON, which is way more stable:

import scrapy
import json

class MozillaSpider(scrapy.Spider):
    name = 'mozilla'
    allowed_domains = ['product-details.mozilla.org']
    start_urls = ['https://product-details.mozilla.org/1.0/firefox_versions.json']

    def parse(self, response):
        version_data = json.loads(response.text)
        latest_version = version_data['LATEST_FIREFOX_VERSION']
        self.logger.info("Got latest Firefox version via API: %s" % latest_version)
        yield {'version': latest_version}

Quick Notes

  1. Python 2.7 EOL: Just a heads-up—Python 2.7 is no longer supported, and Scrapy dropped support for it in v2.0. If you can, upgrade to Python 3.x for better security and features.
  2. Robots.txt: Always check the site's robots.txt first—Mozilla allows crawling both the release notes page and the version API, so you're good to go.

内容的提问来源于stack exchange,提问作者afern13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:19:34