Scrapy-Selenium爬取FT金融数据时URL、股票名称、代码字段为空问题求解
问题原因&修复方案
- 核心逻辑冲突:你同时用了两套请求逻辑,
start_requests里发起的SeleniumRequest请求的是FT首页https://markets.ft.com,最终解析用的response就是这个首页跳转后的https://markets.ft.com/data,自然没有对应基金的名称、代码字段,URL字段也是错的。而你在__init__方法里单独启动driver打开了目标基金页,但加载完直接关掉了driver,也没有把页面内容传递给解析逻辑,等于白跑了一遍。 - 无头模式参数错误:
chrome_options.add_argument('__headless')是错的,正确参数是--headless(两个短横)。 - XPath语法错误:提取基金名称、代码的XPath开头加了
.,代表相对当前节点(遍历的行节点)查找,而基金名称、代码是页面顶层的元素,不能加相对路径的点。 - 没有利用scrapy-selenium的官方能力:scrapy-selenium会自动把driver加载完的页面封装成response返回给回调函数,你不需要自己单独实例化driver。
修正后的代码
import scrapy from selenium.webdriver.chrome.options import Options from scrapy_selenium import SeleniumRequest from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from datetime import datetime class PriceSelSpider(scrapy.Spider): name = 'price_sel' def start_requests(self): # 直接请求目标基金历史页,交给scrapy-selenium加载 yield SeleniumRequest( url="https://markets.ft.com/data/funds/tearsheet/historical?s=GB00B4NXY349:GBP", wait_time=10, # 等待表格加载完成再返回 wait_until=EC.presence_of_element_located((By.XPATH, "//div[@class='mod-ui-table--freeze-pane__scroll-container']/table/tbody/tr")), callback=self.parse ) def parse(self, response): now = datetime.now() dt_string = now.strftime("%Y/%m/%d %H:%M:%S") # 先提取顶层的基金名称、代码,只需要提取一次 share_name = response.xpath("//h1[@class='mod-tearsheet-overview__header__name mod-tearsheet-overview__header__name--large']/text()").get() share_code = response.xpath("//div[@class='mod-tearsheet-overview__header__symbol']/span/text()").get() infos = response.xpath("//div[@class='mod-ui-table--freeze-pane__scroll-container']/table/tbody/tr") for info in infos: yield { 'Time': dt_string, 'URL': response.url, 'Name of the share': share_name, 'Code of the share': share_code, 'Date of share price ': info.xpath(".//td/span[1]/text()").get(), 'Opening price': info.xpath(".//td[2]/text()").get(), 'Highest price': info.xpath(".//td[3]/text()").get(), 'Lowest price': info.xpath(".//td[4]/text()").get(), 'Closing price': info.xpath(".//td[5]/text()").get(), 'Volume': info.xpath("(.//td/span)[3]/text()").get() }
额外注意事项
- 建议把无头模式、窗口大小等driver参数都放到
settings.py中配置,对应配置项为SELENIUM_DRIVER_ARGUMENTS = ['--headless', '--window-size=1920,1080'],不需要自己在爬虫里实例化driver。 - 尽量不要用
time.sleep(),用WebDriverWait的显式等待代替,避免页面没加载完就提取内容的问题。
内容的提问来源于stack exchange,提问作者Raisul Islam
相关产品推荐
相关产品推荐

