You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取FCA手册时丢失PRIN3.3.1表格数据求解决

解决Scrapy爬取FCA手册表格数据遗漏问题

问题排查

你的爬虫遗漏PRIN 3.3.1表格数据的核心原因有两个:

  • 表格分支未输出数据:原代码中检测到表格时仅处理了文本,但没有yield生成对应的Item,导致数据未被收集。
  • 表格元数据未提取:表格所在section的唯一ID、类型等关键信息未在表格分支中获取,无法和其他条目保持统一格式。

另外,原代码的表格选择器div.section-content-table不够精准,建议直接定位表格元素确保覆盖目标内容。

修改后的完整代码

import scrapy
from urllib.parse import urlencode


class HandBook(scrapy.Spider):
    name = "handbook_spider"

    custom_settings = {
        "LOG_FILE": "handbook_spider.log",
        "ITEM_PIPELINES": {
            "handbook_spider.pipelines.HandbookExcelPipeline": 300,
        },
    }

    headers = {
        "authority": "www.handbook.fca.org.uk",
        "accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9",
        "accept-language": "en,ru;q=0.9",
        "cache-control": "max-age=0",
        "sec-ch-ua": '"Chromium";v="106", "Yandex";v="22", "Not;A=Brand";v="99"',
        "sec-ch-ua-mobile": "?0",
        "sec-ch-ua-platform": '"Linux"',
        "sec-fetch-dest": "document",
        "sec-fetch-mode": "navigate",
        "sec-fetch-site": "cross-site",
        "sec-fetch-user": "?1",
        "upgrade-insecure-requests": "1",
        "user-agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 YaBrowser/22.11.3.832 (beta) Yowser/2.5 Safari/537.36",
    }

    params = {
        "date": "2030-12-01",
        "timeline": "True",
        "view": "chapter",
    }

    url = "https://www.handbook.fca.org.uk/handbook/PRIN/3/?"

    def start_requests(self):
        base_url = self.url + urlencode(self.params)
        yield scrapy.Request(
            url=base_url, headers=self.headers, callback=self.parse_details
        )

    def parse_details(self, response):
        for content in response.css("div.handbook-content"):
            chapter_ref = content.xpath(
                "./header/h1/span[@class='extended']/text()"
            ).get()
            chapter = "".join(content.xpath("./header/h1/text()").getall()).strip()
            topic = None
            for section in content.css("section"):
                header = section.css("header")
                # 检查当前section是否包含表格
                has_table = section.css("table").get() is not None
                if header:
                    topic = header.css("h2.crosstitle::text").get()
                
                if has_table:
                    # 提取表格所在section的元数据
                    uid = section.xpath(".//span[@class='extended']/text()").get()
                    section_type = section.css("span.section-type::text").get()
                    date_applicable = section.xpath(".//time/span/text()").get()
                    # 提取表格文本,按行整理
                    table_rows = []
                    for tr in section.css("table tr"):
                        row_text = " | ".join([td.strip() for td in tr.css("td ::text").getall()])
                        if row_text:
                            table_rows.append(row_text)
                    clause_text = "\n".join(table_rows)
                    # 生成表格对应的Item
                    yield {
                        "Unique_ids": uid,
                        "Chapter_ref": chapter_ref,
                        "Chapter": chapter,
                        "Topic": topic,
                        "Clause": uid.split(".")[-2] if uid else None,
                        "Sub_Clause": uid.split(".")[-1] if uid else None,
                        "Type": section_type,
                        "Date_applicable": date_applicable,
                        "Text": clause_text,
                    }
                else:
                    # 原有非表格内容处理逻辑
                    content = section.xpath(
                        ".//div[@class='section-content']//text()"
                    ).getall()
                    clause_text = " ".join(list(map(str.strip, content)))
                    uid = section.xpath(".//span[@class='extended']/text()").get()
                    if section.css("span.section-type").get() is not None:
                        yield {
                            "Unique_ids": uid,
                            "Chapter_ref": chapter_ref,
                            "Chapter": chapter,
                            "Topic": topic,
                            "Clause": uid.split(".")[-2],
                            "Sub_Clause": uid.split(".")[-1],
                            "Type": section.css("span.section-type::text").get(),
                            "Date_applicable": section.xpath(
                                ".//time/span/text()"
                            ).get(),
                            "Text": clause_text,
                        }

关键修改说明

  1. 新增表格检测逻辑:用section.css("table").get()直接判断section是否包含表格,比依赖容器类更可靠。
  2. 补充表格元数据提取:在表格分支中获取uid、类型、生效日期等信息,保持和非表格条目的格式一致。
  3. 优化表格文本整理:遍历表格的行和单元格,用|分隔单元格内容,换行分隔行,保留表格的结构可读性。
  4. 添加表格Item输出:在表格分支中添加yield语句,确保表格数据被收集到Pipeline。

内容的提问来源于stack exchange,提问作者X-something

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 00:20:30