使用BeautifulSoup爬取Medium:如何获取全部文章标题并过滤无关内容
解决Medium文章标题爬取的两个问题:过滤无关内容+获取全部标题
问题分析
你的代码通过find_all('h2')获取所有h2标签,但Medium页面中侧边栏的“编辑推荐”“订阅提示”等无关内容也用了h2标签,同时首页的文章是滚动动态加载的,仅用requests请求一次只能拿到初始加载的部分文章。
修改后的代码
方案1:仅用requests+BeautifulSoup(获取初始页面的有效文章标题)
import requests from bs4 import BeautifulSoup as bs class Publication: def __init__(self, publication): self.publication = publication self.headers = {'User-Agent':'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36'} def get_articles(self): url = f"https://{self.publication}.com/" r = requests.get(url, headers=self.headers) soup = bs(r.text, 'lxml') # 仅筛选文章容器内的h2标题(过滤侧边栏无关内容) article_containers = soup.find_all('div', class_='postArticle-content') for container in article_containers: title_tag = container.find('h2') if title_tag: print(title_tag.text.strip()) publication = Publication('towardsdatascience') publication.get_articles()
方案2:使用Selenium处理动态加载(获取所有滚动加载的文章标题)
如果需要获取首页所有文章标题,就得处理滚动加载,用Selenium模拟浏览器滚动:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time class Publication: def __init__(self, publication): self.publication = publication # 配置无头浏览器(可选,不想弹出浏览器窗口就启用) self.chrome_options = Options() self.chrome_options.add_argument('--headless=new') self.chrome_options.add_argument('user-agent=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36') def get_all_articles(self): url = f"https://{self.publication}.com/" driver = webdriver.Chrome(options=self.chrome_options) driver.get(url) time.sleep(2) # 等待初始加载 # 模拟滚动到底部,加载所有内容 last_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) # 等待加载 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 提取所有有效文章标题 article_titles = driver.find_elements(By.CSS_SELECTOR, 'div.postArticle-content h2') for title in article_titles: print(title.text.strip()) driver.quit() publication = Publication('towardsdatascience') publication.get_all_articles()
修改说明
- 过滤无关内容:
- 不再直接抓取所有h2,而是先定位文章的父容器
div.postArticle-content,再从容器内找h2标题,这样能自动排除侧边栏、页脚等区域的无关h2。
- 不再直接抓取所有h2,而是先定位文章的父容器
- 获取全部文章:
- Medium首页是滚动加载,仅用requests无法获取后续加载的内容,方案2用Selenium模拟浏览器滚动,直到页面不再加载新内容,再提取所有标题。
- 注意:使用Selenium需要提前安装ChromeDriver并配置环境,或者用
webdriver-manager自动管理驱动。
内容的提问来源于stack exchange,提问作者Karthik Bhandary
相关产品推荐
相关产品推荐

