You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Splash爬取含Hashbang的赛马网站页面遇阻求助

解决Scrapy Splash爬取含Hashbang页面的URL转义问题

核心问题原因

Scrapy默认遵循Google的AJAX Crawling规范,会自动把URL中的#!/转换为?_escaped_fragment_=参数,导致请求页面与预期不符。

具体解决步骤

1. 关闭Scrapy的自动转义功能

在Scrapy项目的settings.py中添加配置,禁用Hashbang转义:

AJAXCRAWL_ENABLED = False

2. 用SplashRequest直接传递原始URL

在爬虫的请求方法里,直接传入带#!/的完整目标URL,同时配置Splash的渲染参数:

from scrapy_splash import SplashRequest

def start_requests(self):
    target_urls = [
        'https://www.britishhorseracing.com/racing/results/fixture-results/#!/2023/11284',
        # 其他需要爬取的URL
    ]
    for url in target_urls:
        yield SplashRequest(
            url=url,
            callback=self.parse_results,
            args={
                'wait': 2,  # 等待页面渲染完成,可根据实际情况调整时长
                'html': 1,
                'render_all': 1
            },
            headers={
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            }
        )

3. 验证请求与解析数据

在回调方法里先确认请求URL是否正确,再进行数据提取:

def parse_results(self, response):
    # 打印URL确认是否为原始带Hashbang的地址
    print("当前请求URL:", response.url)
    # 提取赛事结果示例
    race_items = response.xpath('//div[contains(@class, "race-result-item")]')
    for item in race_items:
        yield {
            'race_name': item.xpath('.//h4/text()').get().strip(),
            'horse_name': item.xpath('.//span[contains(@class, "horse-name")]/text()').get().strip(),
            # 其他需要提取的字段
        }

4. 批量构造目标URL

如果需要爬取多个赛事页面,可以批量生成带Hashbang的URL:

def start_requests(self):
    base_url = 'https://www.britishhorseracing.com/racing/results/fixture-results/#!/2023/{}'
    # 示例赛事ID列表,实际可从网站列表页提取
    race_ids = [11284, 11285, 11286]
    for race_id in race_ids:
        yield SplashRequest(
            url=base_url.format(race_id),
            callback=self.parse_results,
            args={'wait': 2}
        )

额外注意事项

  • 确保本地Splash服务正常运行(默认端口8050),可通过访问http://localhost:8050验证。
  • 根据页面加载速度调整wait参数,避免因渲染未完成导致数据提取失败。
  • 控制爬取频率,遵守网站robots.txt规则,防止IP被封禁。

内容的提问来源于stack exchange,提问作者Jay Jericho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 05:17:35