You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+BeautifulSoup爬取WSJ主页所有类型文章?

WSJ主页全量文章爬取优化方案

你的问题核心是现有代码分散处理不同类别的文章,不仅重复代码冗余,还因为选择器定位错误(比如评论类用了署名元素的class,而非文章容器)导致爬取不全。下面是整合后的优化方案:

关键改进点

  • 把所有可能的文章容器CSS类合并到一个选择器中,避免多次重复循环
  • 用通用标题标签选择器(h2, h3, h4)适配不同层级的文章标题
  • 自动补全相对链接为完整URL,避免无效链接
  • 增加去重逻辑,防止同一文章被多次抓取

优化后代码

from bs4 import BeautifulSoup
import requests

url = "https://www.wsj.com"
header = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36"
}

response = requests.get(url, headers=header)
soup = BeautifulSoup(response.content, "lxml")

# 整合所有文章容器的CSS类,注意替换评论文章的实际容器class(需自行通过浏览器开发者工具定位)
article_containers = soup.select(
    ".WSJTheme--headline--7VCzo7Ay, .WSJTheme--bullet-item--5c1Mqfdr, .style--opinion-container--xxxxxx"
)

# 用集合存储已抓取链接,实现去重
seen_links = set()

for container in article_containers:
    # 提取标题:适配h2/h3/h4不同层级标签
    title_elem = container.select_one("h2, h3, h4")
    if not title_elem:
        continue
    headline = title_elem.get_text(strip=True)
    
    # 提取链接并补全为完整URL
    link_elem = container.find("a")
    if not link_elem or "href" not in link_elem.attrs:
        continue
    link = link_elem["href"]
    # 处理相对路径链接
    if not link.startswith("http"):
        link = f"{url}{link}"
    
    # 去重后输出结果
    if link not in seen_links:
        seen_links.add(link)
        print(f"{headline} - {link}")

注意事项

  1. 评论类文章选择器修正:你原代码中用的.style--byline--1k-cTV4i是文章的署名元素,不是评论文章的容器。打开WSJ主页,按F12调出浏览器开发者工具,定位评论区文章的外层容器,替换代码中.style--opinion-container--xxxxxx为实际的class名称。
  2. 动态内容处理:WSJ主页部分内容是JavaScript动态加载的,requests只能抓取静态渲染的内容。如果需要爬取动态加载的文章,建议使用selenium或playwright模拟浏览器加载页面。
  3. 反爬规避:频繁请求可能触发网站反爬机制导致封禁,建议在请求间增加间隔(比如time.sleep(1)),或使用代理IP分散请求来源。

内容的提问来源于stack exchange,提问作者KS_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 23:54:27