You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Scrapy爬取BBC科技新闻仅获第一页数据的问题排查

问题描述
  • 需求:爬取文章内容及标题,用于扩展文本分类数据集。
  • 目标:爬取https://www.bbc.com/news/technology页面下所有分页的全部文章。
  • 问题:代码仅提取了https://www.bbc.com/news/technology?page=1的文章,尽管已设置遍历所有分页,当前items_scraped_count为24,预期应为24×29左右。

用户提供的爬虫代码如下:

class BBCSpider_2(scrapy.Spider):

    name = "bbc_tech"
    start_urls = ["https://www.bbc.com/news/technology"]


    def parse(self, response: Response, **kwargs: Any) -> Any:
        max_pages = response.xpath("//nav[@aria-label='Page']/div/div/div/div/ol/li[last()]/div/a//text()").get()
        max_pages = int(max_pages)
        for p in range(max_pages):
            page = f"https://www.bbc.com/news/technology?page={p+1}"
            yield response.follow(page, callback=self.parse_articles2)

进入对应页面提取文章链接:

def parse_articles2(self, response):
        container_to_scan = [4, 8]
        for box in container_to_scan:
            if box == 4:
                articles = response.xpath(f"//*[@id='main-content']/div[{box}]/div/div/ul/li")
            if box == 8:
                articles = response.xpath(f"//*[@id='main-content']/div[{box}]/div[2]/ol/li")
            for article_idx in range(len(articles)):
                if box == 4:
                    relative_url = response.xpath(f"//*[@id='main-content']/div[4]/div/div/ul/li[{article_idx+1}]/div/div/div/div[1]/div[1]/a/@href").get()
                elif box == 8:
                    relative_url = response.xpath(f"//*[@id='main-content']/div[8]/div[2]/ol/li[{article_idx+1}]/div/div/div[1]/div[1]/a/@href").get()
                else:
                    relative_url = None

                if relative_url is not None:
                    followup_url = "https://www.bbc.com" + relative_url
                    yield response.follow(followup_url, callback=self.parse_article)

爬取单篇文章内容和标题:

def parse_article(response):
        article_text = response.xpath("//article/div[@data-component='text-block']")
        content = []
        for box in article_text:
            text = box.css("div p::text").get()
            if text is not None:
                content.append(text)

        title = response.css("h1::text").get()

        yield {
            "title": title,
            "content": content,
        }
问题排查与修复

1. 分页逻辑失效

原代码中获取最大页数的XPath依赖固定页面结构,可能未正确定位到最后一页数字。可以打印max_pages的值,大概率是获取到了1,导致只遍历1页。

修复代码:
改用提取所有分页链接的方式获取最大页数,逻辑更健壮:

def parse(self, response: Response, **kwargs: Any) -> Any:
    # 提取所有分页链接中的page参数
    page_links = response.xpath("//nav[@aria-label='Page']//a/@href").getall()
    max_pages = 1
    for link in page_links:
        if 'page=' in link:
            page_num = int(link.split('page=')[1])
            if page_num > max_pages:
                max_pages = page_num
    # 遍历所有分页(从1到max_pages)
    for p in range(1, max_pages + 1):
        page_url = f"https://www.bbc.com/news/technology?page={p}"
        yield scrapy.Request(page_url, callback=self.parse_articles2)

2. 硬编码容器索引导致漏爬

原parse_articles2方法用固定的div[4]、div[8]定位文章容器,BBC不同分页的页面结构可能变化,导致其他分页无法提取文章链接。

修复代码:
直接通过文章链接的特征定位,摆脱容器索引依赖:

def parse_articles2(self, response):
    # 提取所有BBC新闻文章的链接
    article_links = response.xpath("//a[contains(@href, '/news/articles/')]/@href").getall()
    # 去重避免重复爬取
    unique_links = list(set(article_links))
    for link in unique_links:
        if link.startswith('/'):
            full_url = "https://www.bbc.com" + link
            yield response.follow(full_url, callback=self.parse_article)

3. 文章解析方法参数错误

原parse_article方法缺少self参数,不符合Scrapy Spider方法的定义规范,会导致调用失败。同时原代码仅提取每个文本块的第一个段落,内容不完整。

修复代码:

def parse_article(self, response):
    article_text = response.xpath("//article/div[@data-component='text-block']")
    content = []
    for box in article_text:
        # 获取当前文本块下所有段落文本
        texts = box.css("div p::text").getall()
        content.extend(texts)
    title = response.css("h1::text").get()

    yield {
        "title": title,
        "content": content,
    }

4. 反爬与日志配置

为避免被BBC限制,建议在settings.py中添加:

# 增加请求延迟
DOWNLOAD_DELAY = 2
# 设置模拟浏览器的User-Agent
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
# 启用DEBUG日志查看请求详情
LOG_LEVEL = 'DEBUG'

内容的提问来源于stack exchange,提问作者Seb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 14:15:08