You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Scrapy与BeautifulSoup爬取网页遇阻:无法获取目标页面所有文章链接

解决Scrapy和BeautifulSoup爬取CNN印度商业页面无输出的问题

嘿,我来帮你排查下问题!你遇到的没输出数据的情况,主要是几个小细节没处理好,咱们分别针对Scrapy和BeautifulSoup来修复:

一、Scrapy版本的问题与修复

你的Scrapy代码有三个关键问题:

  • 属性拼写错误:Scrapy Spider类需要的是start_urls(复数形式),而你写的是start_url,这会导致Scrapy根本不会发起请求。
  • 反爬拦截:CNN会检测请求的User-Agent,默认的Scrapy User-Agent容易被拦截,返回的页面没有目标内容。
  • 提取逻辑不对:你当前只是打印span元素,而文章链接其实在包含该span的<a>标签里。

修正后的代码:

import scrapy

class CnnSpider(scrapy.Spider):
    name = "cnn"
    # 修正为复数形式的start_urls
    start_urls = ['https://edition.cnn.com/business/india/index.html']
    
    # 设置自定义请求头,避免被反爬拦截
    custom_settings = {
        'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }

    def parse(self, response):
        # 找到包含标题span的a标签,提取链接和标题文本
        for article in response.css('a.cd__headline'):
            article_link = article.css('::attr(href)').get()
            # 拼接完整URL,因为CNN的链接是相对路径
            full_link = response.urljoin(article_link)
            article_title = article.css('span.cd__headline-text::text').get()
            # 输出结果
            print(f"标题: {article_title}, 链接: {full_link}")
            # 如果要保存到Item,可以返回字典或Item对象
            # yield {'title': article_title, 'link': full_link}

二、BeautifulSoup版本的问题与修复

你的BeautifulSoup代码主要问题是:

  • 反爬拦截:requests.get没有设置User-Agent,CNN会返回无内容的页面,导致找不到目标元素。
  • 提取逻辑偏差:链接在包含span.cd__headline-text的<a>标签中,直接找span无法获取链接。

修正后的代码:

import requests
from bs4 import BeautifulSoup

def get_cnn_links():
    url = "https://edition.cnn.com/business/india/index.html"
    # 设置请求头,模拟浏览器请求
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    # 带headers发起请求
    page = requests.get(url, headers=headers)
    # 检查请求是否成功
    if page.status_code != 200:
        print(f"请求失败,状态码: {page.status_code}")
        return
    soup = BeautifulSoup(page.content, "html.parser")
    # 找到所有包含标题span的a标签
    articles = soup.find_all('a', class_="cd__headline")
    for article in articles:
        article_link = article.get('href')
        # 拼接完整URL
        full_link = f"https://edition.cnn.com{article_link}" if article_link.startswith('/') else article_link
        article_title = article.find('span', class_="cd__headline-text").get_text(strip=True)
        print(f"标题: {article_title}, 链接: {full_link}")

# 调用函数
get_cnn_links()

额外说明

如果还是没有数据,可能是CNN启用了更严格的反爬(比如需要Cookie或JS渲染),这时候可以考虑:

  • Scrapy中添加DOWNLOAD_DELAY设置,降低请求频率
  • 使用selenium或playwright处理动态渲染的页面

内容的提问来源于stack exchange,提问作者Marco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 23:12:31