You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Medium:如何获取全部文章标题并过滤无关内容

解决Medium文章标题爬取的两个问题:过滤无关内容+获取全部标题

问题分析

你的代码通过find_all('h2')获取所有h2标签,但Medium页面中侧边栏的“编辑推荐”“订阅提示”等无关内容也用了h2标签,同时首页的文章是滚动动态加载的,仅用requests请求一次只能拿到初始加载的部分文章。

修改后的代码

方案1:仅用requests+BeautifulSoup(获取初始页面的有效文章标题)

import requests
from bs4 import BeautifulSoup as bs

class Publication:
    def __init__(self, publication):
        self.publication = publication
        self.headers = {'User-Agent':'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36'}

    def get_articles(self):
        url = f"https://{self.publication}.com/"
        r = requests.get(url, headers=self.headers)
        soup = bs(r.text, 'lxml')
        
        # 仅筛选文章容器内的h2标题(过滤侧边栏无关内容)
        article_containers = soup.find_all('div', class_='postArticle-content')
        for container in article_containers:
            title_tag = container.find('h2')
            if title_tag:
                print(title_tag.text.strip())

publication = Publication('towardsdatascience')
publication.get_articles()

方案2:使用Selenium处理动态加载(获取所有滚动加载的文章标题)

如果需要获取首页所有文章标题,就得处理滚动加载,用Selenium模拟浏览器滚动:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time

class Publication:
    def __init__(self, publication):
        self.publication = publication
        # 配置无头浏览器(可选,不想弹出浏览器窗口就启用)
        self.chrome_options = Options()
        self.chrome_options.add_argument('--headless=new')
        self.chrome_options.add_argument('user-agent=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36')

    def get_all_articles(self):
        url = f"https://{self.publication}.com/"
        driver = webdriver.Chrome(options=self.chrome_options)
        driver.get(url)
        time.sleep(2)  # 等待初始加载

        # 模拟滚动到底部,加载所有内容
        last_height = driver.execute_script("return document.body.scrollHeight")
        while True:
            # 滚动到底部
            driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            time.sleep(3)  # 等待加载
            new_height = driver.execute_script("return document.body.scrollHeight")
            if new_height == last_height:
                break
            last_height = new_height

        # 提取所有有效文章标题
        article_titles = driver.find_elements(By.CSS_SELECTOR, 'div.postArticle-content h2')
        for title in article_titles:
            print(title.text.strip())

        driver.quit()

publication = Publication('towardsdatascience')
publication.get_all_articles()

修改说明

  1. 过滤无关内容:
    • 不再直接抓取所有h2,而是先定位文章的父容器div.postArticle-content,再从容器内找h2标题,这样能自动排除侧边栏、页脚等区域的无关h2。
  2. 获取全部文章:
    • Medium首页是滚动加载,仅用requests无法获取后续加载的内容,方案2用Selenium模拟浏览器滚动,直到页面不再加载新内容,再提取所有标题。
    • 注意:使用Selenium需要提前安装ChromeDriver并配置环境,或者用webdriver-manager自动管理驱动。

内容的提问来源于stack exchange,提问作者Karthik Bhandary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 03:53:23