如何用Python Scrapy爬取第三方软件版本号并构建版本仓库?(附Firefox示例)
Fixing Your Scrapy Spider to Grab Firefox's Latest Version
Hey there! Let's work through your Scrapy spider issue for fetching Firefox's current version number. First, I noticed your parse method was cut off, so I'll fill that in, fix some key issues, and share more reliable approaches tailored to this specific task.
Full Working Code Example
import scrapy import html2text import re from scrapy.selector import Selector # HtmlXPathSelector is deprecated in newer Scrapy versions, though still compatible with Python 2.7's Scrapy builds class MozillaSpider(scrapy.Spider): name = 'mozilla' allowed_domains = ['mozilla.org'] # Heads up: your original domain was wrong (it's mozilla.org, not .com) start_urls = ['https://www.mozilla.org/en-US/firefox/notes/'] def parse(self, response): # Method 1: CSS Selector (targets the version heading directly) version_css = response.css('.c-release-version h2::text').extract_first() if version_css: pure_version = version_css.strip().split(' ')[-1] self.logger.info("Got Firefox version via CSS: %s" % pure_version) yield {'version': pure_version} # Method 2: XPath (fallback if CSS selector breaks) version_xpath = response.xpath('//div[contains(@class, "c-release-version")]/h2/text()').extract_first() if version_xpath: pure_version = version_xpath.strip().split(' ')[-1] self.logger.info("Got Firefox version via XPath: %s" % pure_version) yield {'version': pure_version} # Method 3: Plain text parsing with html2text (good for layout changes) h = html2text.HTML2Text() h.ignore_links = True plain_text = h.handle(response.text) version_match = re.search(r'Firefox (\d+\.\d+\.?\d*)', plain_text) if version_match: pure_version = version_match.group(1) self.logger.info("Got Firefox version via text match: %s" % pure_version) yield {'version': pure_version}
Key Fixes & Explanations
- Domain Correction: Your original
allowed_domainsusedmozilla.com, but the target site ismozilla.org—this would have caused Scrapy to block your request, so that's a critical fix. - Selector Updates:
HtmlXPathSelectoris deprecated in newer Scrapy versions. While Python 2.7's compatible Scrapy (v1.x) still supports it, usingresponse.css()orresponse.xpath()directly is cleaner. - Version Targeting: Firefox's latest version lives in a
.c-release-versioncontainer's<h2>tag on the release notes page. I used browser dev tools (F12) to confirm this selector works right now. - Python 2.7 Compatibility: I avoided Python 3+ features like f-strings since you're on 2.7—stick with
%formatting or.format()for strings.
Even More Reliable Approach (API Instead of HTML)
If you want to avoid breaking your spider every time Mozilla updates their page layout, use their official version API instead. It returns structured JSON, which is way more stable:
import scrapy import json class MozillaSpider(scrapy.Spider): name = 'mozilla' allowed_domains = ['product-details.mozilla.org'] start_urls = ['https://product-details.mozilla.org/1.0/firefox_versions.json'] def parse(self, response): version_data = json.loads(response.text) latest_version = version_data['LATEST_FIREFOX_VERSION'] self.logger.info("Got latest Firefox version via API: %s" % latest_version) yield {'version': latest_version}
Quick Notes
- Python 2.7 EOL: Just a heads-up—Python 2.7 is no longer supported, and Scrapy dropped support for it in v2.0. If you can, upgrade to Python 3.x for better security and features.
- Robots.txt: Always check the site's robots.txt first—Mozilla allows crawling both the release notes page and the version API, so you're good to go.
内容的提问来源于stack exchange,提问作者afern13
相关产品推荐
相关产品推荐

