You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需Selenium爬取Medium文章点赞数的方法求助

无需Selenium爬取Medium文章点赞数的方法

Medium的点赞数是通过前端JavaScript动态渲染的,直接用requests获取静态HTML只能得到占位符“--”,但页面中会内嵌包含完整文章数据的JSON脚本,我们可以直接提取这些数据来获取点赞数。

具体实现步骤

  • 发送请求时添加浏览器UA头,避免被Medium的反爬机制拦截
  • 提取页面内嵌的JSON数据,从中解析点赞数

代码示例

import requests
import json
from bs4 import BeautifulSoup

url = "https://medium.com/@xdxxxx4713/basic-settings-of-nginx-aeace532534f"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 获取页面内容
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 方法1:从schema.org元数据中提取
script_tag = soup.find("script", type="application/ld+json")
if script_tag:
    metadata = json.loads(script_tag.string)
    for stat in metadata.get("interactionStatistic", []):
        if stat.get("interactionType") == "http://schema.org/LikeAction":
            print(f"点赞数:{stat['userInteractionCount']}")
            exit()

# 方法2:从Apollo状态数据中提取(方法1失效时用)
script_tag = soup.find("script", string=lambda t: t and "__APOLLO_STATE__" in t)
if script_tag:
    # 剥离JS包裹,提取纯JSON
    raw_json = script_tag.string.split("window.__APOLLO_STATE__ = ")[1].rstrip(";")
    apollo_data = json.loads(raw_json)
    # 遍历Post节点找点赞数
    for key in apollo_data:
        if key.startswith("Post:"):
            like_count = apollo_data[key]["virtuals"]["reactions"]["count"]
            print(f"点赞数:{like_count}")
            break

注意事项

  • 必须添加User-Agent头,否则Medium会返回不包含完整数据的页面
  • 不同Medium文章的内嵌数据结构可能略有差异,若其中一种方法失效,可尝试另一种

内容的提问来源于stack exchange,提问作者Sefa Kalkan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 17:33:26