You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python Requests与Beautiful Soup的CoinDesk多文章爬取求助

如何爬取CoinDesk比特币市场新闻的所有文章

你现在的代码只提取了列表页的第一篇文章预览信息,这是因为你直接在列表页的DOM里取了第一个h3和p.desc。要爬取所有文章,你需要先从列表页提取每篇文章的详情页链接,再逐个访问这些链接去爬取完整内容。下面是具体的解决步骤和修改后的代码:

问题分析

当前代码的核心问题:

  • 你访问的是文章列表页,而非单篇文章的详情页
  • 列表页的h3和p.desc只是文章的预览内容,且你只取了第一个匹配项,所以只能拿到第一篇

解决方案步骤

  1. 从列表页提取所有文章的详情页URL:遍历列表页的所有文章卡片,提取每个卡片对应的详情页链接
  2. 遍历每个详情页URL,爬取完整内容:访问每个详情页,提取标题和正文内容
  3. 添加基础异常处理:避免单个请求失败导致整个程序中断

修改后的完整代码

import requests
from bs4 import BeautifulSoup

class Content:
    def __init__(self, url, title, body):
        self.url = url
        self.title = title
        self.body = body

def getPage(url):
    try:
        req = requests.get(url, headers={"User-Agent": "Mozilla/5.0"})
        req.raise_for_status()  # 抛出HTTP请求错误
        return BeautifulSoup(req.text, 'html.parser')
    except requests.exceptions.RequestException as e:
        print(f"请求页面失败: {e}")
        return None

# 爬取单篇CoinDesk文章的完整内容
def scrapeSingleCoindeskArticle(url):
    bs = getPage(url)
    if not bs:
        return None
    
    # 提取详情页标题
    title_tag = bs.find("h1", class_="typography__StyledTypography-owin6q-0")
    title = title_tag.text.strip() if title_tag else "无标题"
    
    # 提取详情页正文(合并所有正文段落)
    body_paragraphs = bs.find_all("p", class_="typography__StyledTypography-owin6q-0")
    body = "\n\n".join([p.text.strip() for p in body_paragraphs if p.text.strip()])
    
    return Content(url, title, body)

# 从CoinDesk列表页获取所有文章并爬取
def scrapeCoindeskArticles(url):
    bs = getPage(url)
    if not bs:
        return []
    
    # 提取列表页所有文章的详情页链接
    article_links = []
    # 匹配文章卡片中的链接(根据页面DOM结构调整选择器)
    article_cards = bs.find_all("a", class_="linkstyles__StyledLink-sc-14289xe-0")
    for card in article_cards:
        link = card.get("href")
        # 处理相对链接,拼接完整URL
        if link and link.startswith("/"):
            full_link = f"https://www.coindesk.com{link}"
            article_links.append(full_link)
    
    # 去重(避免重复链接)
    article_links = list(set(article_links))
    
    # 遍历所有链接,爬取每篇文章
    articles = []
    for link in article_links:
        print(f"正在爬取: {link}")
        article = scrapeSingleCoindeskArticle(link)
        if article:
            articles.append(article)
    
    return articles

# 主程序执行
if __name__ == "__main__":
    url = 'https://www.coindesk.com/category/markets-news/markets-markets-news/markets-bitcoin/'
    articles = scrapeCoindeskArticles(url)
    
    # 输出爬取结果
    for idx, article in enumerate(articles, 1):
        print(f"\n=== 第{idx}篇文章 ===")
        print(f"标题: {article.title}")
        print(f"URL: {article.url}")
        print(f"正文预览: {article.body[:300]}...")  # 只输出前300字符预览

关键说明

  • User-Agent设置:添加了浏览器UA头,避免被网站反爬机制拦截
  • 异常处理:在getPage函数中捕获请求异常,防止程序崩溃
  • 选择器适配:根据CoinDesk当前的DOM结构调整了选择器(如果后续网站结构变化,需要对应修改选择器)
  • 链接去重:使用set去除可能的重复链接
  • 正文提取:合并详情页所有正文段落,获取完整内容

如果需要爬取多页列表(比如下一页),还可以进一步添加翻页逻辑:找到列表页的"下一页"按钮链接,循环访问每一页的列表页即可。

内容的提问来源于stack exchange,提问作者Martin598

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:41:25