基于Python Requests与Beautiful Soup的CoinDesk多文章爬取求助
如何爬取CoinDesk比特币市场新闻的所有文章
你现在的代码只提取了列表页的第一篇文章预览信息,这是因为你直接在列表页的DOM里取了第一个h3和p.desc。要爬取所有文章,你需要先从列表页提取每篇文章的详情页链接,再逐个访问这些链接去爬取完整内容。下面是具体的解决步骤和修改后的代码:
问题分析
当前代码的核心问题:
- 你访问的是文章列表页,而非单篇文章的详情页
- 列表页的
h3和p.desc只是文章的预览内容,且你只取了第一个匹配项,所以只能拿到第一篇
解决方案步骤
- 从列表页提取所有文章的详情页URL:遍历列表页的所有文章卡片,提取每个卡片对应的详情页链接
- 遍历每个详情页URL,爬取完整内容:访问每个详情页,提取标题和正文内容
- 添加基础异常处理:避免单个请求失败导致整个程序中断
修改后的完整代码
import requests from bs4 import BeautifulSoup class Content: def __init__(self, url, title, body): self.url = url self.title = title self.body = body def getPage(url): try: req = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}) req.raise_for_status() # 抛出HTTP请求错误 return BeautifulSoup(req.text, 'html.parser') except requests.exceptions.RequestException as e: print(f"请求页面失败: {e}") return None # 爬取单篇CoinDesk文章的完整内容 def scrapeSingleCoindeskArticle(url): bs = getPage(url) if not bs: return None # 提取详情页标题 title_tag = bs.find("h1", class_="typography__StyledTypography-owin6q-0") title = title_tag.text.strip() if title_tag else "无标题" # 提取详情页正文(合并所有正文段落) body_paragraphs = bs.find_all("p", class_="typography__StyledTypography-owin6q-0") body = "\n\n".join([p.text.strip() for p in body_paragraphs if p.text.strip()]) return Content(url, title, body) # 从CoinDesk列表页获取所有文章并爬取 def scrapeCoindeskArticles(url): bs = getPage(url) if not bs: return [] # 提取列表页所有文章的详情页链接 article_links = [] # 匹配文章卡片中的链接(根据页面DOM结构调整选择器) article_cards = bs.find_all("a", class_="linkstyles__StyledLink-sc-14289xe-0") for card in article_cards: link = card.get("href") # 处理相对链接,拼接完整URL if link and link.startswith("/"): full_link = f"https://www.coindesk.com{link}" article_links.append(full_link) # 去重(避免重复链接) article_links = list(set(article_links)) # 遍历所有链接,爬取每篇文章 articles = [] for link in article_links: print(f"正在爬取: {link}") article = scrapeSingleCoindeskArticle(link) if article: articles.append(article) return articles # 主程序执行 if __name__ == "__main__": url = 'https://www.coindesk.com/category/markets-news/markets-markets-news/markets-bitcoin/' articles = scrapeCoindeskArticles(url) # 输出爬取结果 for idx, article in enumerate(articles, 1): print(f"\n=== 第{idx}篇文章 ===") print(f"标题: {article.title}") print(f"URL: {article.url}") print(f"正文预览: {article.body[:300]}...") # 只输出前300字符预览
关键说明
- User-Agent设置:添加了浏览器UA头,避免被网站反爬机制拦截
- 异常处理:在
getPage函数中捕获请求异常,防止程序崩溃 - 选择器适配:根据CoinDesk当前的DOM结构调整了选择器(如果后续网站结构变化,需要对应修改选择器)
- 链接去重:使用
set去除可能的重复链接 - 正文提取:合并详情页所有正文段落,获取完整内容
如果需要爬取多页列表(比如下一页),还可以进一步添加翻页逻辑:找到列表页的"下一页"按钮链接,循环访问每一页的列表页即可。
内容的提问来源于stack exchange,提问作者Martin598
相关产品推荐
相关产品推荐

