使用Python Scrapy爬取BBC科技新闻仅获第一页数据的问题排查
问题描述
- 需求:爬取文章内容及标题,用于扩展文本分类数据集。
- 目标:爬取
https://www.bbc.com/news/technology页面下所有分页的全部文章。 - 问题:代码仅提取了
https://www.bbc.com/news/technology?page=1的文章,尽管已设置遍历所有分页,当前items_scraped_count为24,预期应为24×29左右。
用户提供的爬虫代码如下:
class BBCSpider_2(scrapy.Spider): name = "bbc_tech" start_urls = ["https://www.bbc.com/news/technology"] def parse(self, response: Response, **kwargs: Any) -> Any: max_pages = response.xpath("//nav[@aria-label='Page']/div/div/div/div/ol/li[last()]/div/a//text()").get() max_pages = int(max_pages) for p in range(max_pages): page = f"https://www.bbc.com/news/technology?page={p+1}" yield response.follow(page, callback=self.parse_articles2)
进入对应页面提取文章链接:
def parse_articles2(self, response): container_to_scan = [4, 8] for box in container_to_scan: if box == 4: articles = response.xpath(f"//*[@id='main-content']/div[{box}]/div/div/ul/li") if box == 8: articles = response.xpath(f"//*[@id='main-content']/div[{box}]/div[2]/ol/li") for article_idx in range(len(articles)): if box == 4: relative_url = response.xpath(f"//*[@id='main-content']/div[4]/div/div/ul/li[{article_idx+1}]/div/div/div/div[1]/div[1]/a/@href").get() elif box == 8: relative_url = response.xpath(f"//*[@id='main-content']/div[8]/div[2]/ol/li[{article_idx+1}]/div/div/div[1]/div[1]/a/@href").get() else: relative_url = None if relative_url is not None: followup_url = "https://www.bbc.com" + relative_url yield response.follow(followup_url, callback=self.parse_article)
爬取单篇文章内容和标题:
def parse_article(response): article_text = response.xpath("//article/div[@data-component='text-block']") content = [] for box in article_text: text = box.css("div p::text").get() if text is not None: content.append(text) title = response.css("h1::text").get() yield { "title": title, "content": content, }
问题排查与修复
1. 分页逻辑失效
原代码中获取最大页数的XPath依赖固定页面结构,可能未正确定位到最后一页数字。可以打印max_pages的值,大概率是获取到了1,导致只遍历1页。
修复代码:
改用提取所有分页链接的方式获取最大页数,逻辑更健壮:
def parse(self, response: Response, **kwargs: Any) -> Any: # 提取所有分页链接中的page参数 page_links = response.xpath("//nav[@aria-label='Page']//a/@href").getall() max_pages = 1 for link in page_links: if 'page=' in link: page_num = int(link.split('page=')[1]) if page_num > max_pages: max_pages = page_num # 遍历所有分页(从1到max_pages) for p in range(1, max_pages + 1): page_url = f"https://www.bbc.com/news/technology?page={p}" yield scrapy.Request(page_url, callback=self.parse_articles2)
2. 硬编码容器索引导致漏爬
原parse_articles2方法用固定的div[4]、div[8]定位文章容器,BBC不同分页的页面结构可能变化,导致其他分页无法提取文章链接。
修复代码:
直接通过文章链接的特征定位,摆脱容器索引依赖:
def parse_articles2(self, response): # 提取所有BBC新闻文章的链接 article_links = response.xpath("//a[contains(@href, '/news/articles/')]/@href").getall() # 去重避免重复爬取 unique_links = list(set(article_links)) for link in unique_links: if link.startswith('/'): full_url = "https://www.bbc.com" + link yield response.follow(full_url, callback=self.parse_article)
3. 文章解析方法参数错误
原parse_article方法缺少self参数,不符合Scrapy Spider方法的定义规范,会导致调用失败。同时原代码仅提取每个文本块的第一个段落,内容不完整。
修复代码:
def parse_article(self, response): article_text = response.xpath("//article/div[@data-component='text-block']") content = [] for box in article_text: # 获取当前文本块下所有段落文本 texts = box.css("div p::text").getall() content.extend(texts) title = response.css("h1::text").get() yield { "title": title, "content": content, }
4. 反爬与日志配置
为避免被BBC限制,建议在settings.py中添加:
# 增加请求延迟 DOWNLOAD_DELAY = 2 # 设置模拟浏览器的User-Agent USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' # 启用DEBUG日志查看请求详情 LOG_LEVEL = 'DEBUG'
内容的提问来源于stack exchange,提问作者Seb
相关产品推荐
相关产品推荐

