You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy-Selenium爬取FT金融数据时URL、股票名称、代码字段为空问题求解

问题原因&修复方案
  • 核心逻辑冲突:你同时用了两套请求逻辑,start_requests里发起的SeleniumRequest请求的是FT首页https://markets.ft.com,最终解析用的response就是这个首页跳转后的https://markets.ft.com/data,自然没有对应基金的名称、代码字段,URL字段也是错的。而你在__init__方法里单独启动driver打开了目标基金页,但加载完直接关掉了driver,也没有把页面内容传递给解析逻辑,等于白跑了一遍。
  • 无头模式参数错误:chrome_options.add_argument('__headless')是错的,正确参数是--headless(两个短横)。
  • XPath语法错误:提取基金名称、代码的XPath开头加了.,代表相对当前节点(遍历的行节点)查找,而基金名称、代码是页面顶层的元素,不能加相对路径的点。
  • 没有利用scrapy-selenium的官方能力:scrapy-selenium会自动把driver加载完的页面封装成response返回给回调函数,你不需要自己单独实例化driver。
修正后的代码
import scrapy
from selenium.webdriver.chrome.options import Options
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from datetime import datetime


class PriceSelSpider(scrapy.Spider):
    name = 'price_sel'

    def start_requests(self):
        # 直接请求目标基金历史页,交给scrapy-selenium加载
        yield SeleniumRequest(
            url="https://markets.ft.com/data/funds/tearsheet/historical?s=GB00B4NXY349:GBP",
            wait_time=10,
            # 等待表格加载完成再返回
            wait_until=EC.presence_of_element_located((By.XPATH, "//div[@class='mod-ui-table--freeze-pane__scroll-container']/table/tbody/tr")),
            callback=self.parse
        )

    def parse(self, response):
        now = datetime.now()
        dt_string = now.strftime("%Y/%m/%d %H:%M:%S")
        # 先提取顶层的基金名称、代码,只需要提取一次
        share_name = response.xpath("//h1[@class='mod-tearsheet-overview__header__name mod-tearsheet-overview__header__name--large']/text()").get()
        share_code = response.xpath("//div[@class='mod-tearsheet-overview__header__symbol']/span/text()").get()
        infos = response.xpath("//div[@class='mod-ui-table--freeze-pane__scroll-container']/table/tbody/tr")
        for info in infos:
            yield {
                'Time': dt_string,
                'URL': response.url,
                'Name of the share': share_name,
                'Code of the share': share_code,
                'Date of share price ': info.xpath(".//td/span[1]/text()").get(),
                'Opening price': info.xpath(".//td[2]/text()").get(),
                'Highest price': info.xpath(".//td[3]/text()").get(),
                'Lowest price': info.xpath(".//td[4]/text()").get(),
                'Closing price': info.xpath(".//td[5]/text()").get(),
                'Volume': info.xpath("(.//td/span)[3]/text()").get()
            }
额外注意事项
  • 建议把无头模式、窗口大小等driver参数都放到settings.py中配置,对应配置项为SELENIUM_DRIVER_ARGUMENTS = ['--headless', '--window-size=1920,1080'],不需要自己在爬虫里实例化driver。
  • 尽量不要用time.sleep(),用WebDriverWait的显式等待代替,避免页面没加载完就提取内容的问题。

内容的提问来源于stack exchange,提问作者Raisul Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 23:06:03