Scrapy爬取FCA手册时丢失PRIN3.3.1表格数据求解决
解决Scrapy爬取FCA手册表格数据遗漏问题
问题排查
你的爬虫遗漏PRIN 3.3.1表格数据的核心原因有两个:
- 表格分支未输出数据:原代码中检测到表格时仅处理了文本,但没有
yield生成对应的Item,导致数据未被收集。 - 表格元数据未提取:表格所在section的唯一ID、类型等关键信息未在表格分支中获取,无法和其他条目保持统一格式。
另外,原代码的表格选择器div.section-content-table不够精准,建议直接定位表格元素确保覆盖目标内容。
修改后的完整代码
import scrapy from urllib.parse import urlencode class HandBook(scrapy.Spider): name = "handbook_spider" custom_settings = { "LOG_FILE": "handbook_spider.log", "ITEM_PIPELINES": { "handbook_spider.pipelines.HandbookExcelPipeline": 300, }, } headers = { "authority": "www.handbook.fca.org.uk", "accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9", "accept-language": "en,ru;q=0.9", "cache-control": "max-age=0", "sec-ch-ua": '"Chromium";v="106", "Yandex";v="22", "Not;A=Brand";v="99"', "sec-ch-ua-mobile": "?0", "sec-ch-ua-platform": '"Linux"', "sec-fetch-dest": "document", "sec-fetch-mode": "navigate", "sec-fetch-site": "cross-site", "sec-fetch-user": "?1", "upgrade-insecure-requests": "1", "user-agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 YaBrowser/22.11.3.832 (beta) Yowser/2.5 Safari/537.36", } params = { "date": "2030-12-01", "timeline": "True", "view": "chapter", } url = "https://www.handbook.fca.org.uk/handbook/PRIN/3/?" def start_requests(self): base_url = self.url + urlencode(self.params) yield scrapy.Request( url=base_url, headers=self.headers, callback=self.parse_details ) def parse_details(self, response): for content in response.css("div.handbook-content"): chapter_ref = content.xpath( "./header/h1/span[@class='extended']/text()" ).get() chapter = "".join(content.xpath("./header/h1/text()").getall()).strip() topic = None for section in content.css("section"): header = section.css("header") # 检查当前section是否包含表格 has_table = section.css("table").get() is not None if header: topic = header.css("h2.crosstitle::text").get() if has_table: # 提取表格所在section的元数据 uid = section.xpath(".//span[@class='extended']/text()").get() section_type = section.css("span.section-type::text").get() date_applicable = section.xpath(".//time/span/text()").get() # 提取表格文本,按行整理 table_rows = [] for tr in section.css("table tr"): row_text = " | ".join([td.strip() for td in tr.css("td ::text").getall()]) if row_text: table_rows.append(row_text) clause_text = "\n".join(table_rows) # 生成表格对应的Item yield { "Unique_ids": uid, "Chapter_ref": chapter_ref, "Chapter": chapter, "Topic": topic, "Clause": uid.split(".")[-2] if uid else None, "Sub_Clause": uid.split(".")[-1] if uid else None, "Type": section_type, "Date_applicable": date_applicable, "Text": clause_text, } else: # 原有非表格内容处理逻辑 content = section.xpath( ".//div[@class='section-content']//text()" ).getall() clause_text = " ".join(list(map(str.strip, content))) uid = section.xpath(".//span[@class='extended']/text()").get() if section.css("span.section-type").get() is not None: yield { "Unique_ids": uid, "Chapter_ref": chapter_ref, "Chapter": chapter, "Topic": topic, "Clause": uid.split(".")[-2], "Sub_Clause": uid.split(".")[-1], "Type": section.css("span.section-type::text").get(), "Date_applicable": section.xpath( ".//time/span/text()" ).get(), "Text": clause_text, }
关键修改说明
- 新增表格检测逻辑:用
section.css("table").get()直接判断section是否包含表格,比依赖容器类更可靠。 - 补充表格元数据提取:在表格分支中获取uid、类型、生效日期等信息,保持和非表格条目的格式一致。
- 优化表格文本整理:遍历表格的行和单元格,用
|分隔单元格内容,换行分隔行,保留表格的结构可读性。 - 添加表格Item输出:在表格分支中添加
yield语句,确保表格数据被收集到Pipeline。
内容的提问来源于stack exchange,提问作者X-something
相关产品推荐
相关产品推荐

