You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用requests或Playwright抓取Coinbase博客指定区块的最新资讯

Coinbase博客内容抓取解决方案

问题根因

你遇到的获取不到目标区块内容的问题,本质是Medium平台的懒加载机制导致:首页初始服务端渲染仅返回前9篇文章内容,你要的目标区块属于用户滚动到页面对应位置后,才通过异步接口请求加载的动态内容,直接请求初始页面URL无法拿到这部分数据。

优先方案:requests实现

Medium对外提供了公开的用户内容查询接口,不需要模拟浏览器即可直接获取全量文章数据,实现逻辑如下:

  1. 接口地址为 https://medium.com/@coinbase/latest,可通过limit参数自定义返回文章数量
  2. 请求头需指定 Accept: application/json,否则会返回HTML页面
  3. 接口返回的内容前带有防爬虫的干扰字符串 ])}while(1);</x>,需先删除该字符串再解析JSON

示例代码:

import requests
import json

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "application/json"
}
# limit参数可根据需要调整,设置20即可覆盖前9篇之后的目标区块内容
params = {"limit": 20}
response = requests.get("https://medium.com/@coinbase/latest", headers=headers, params=params)

# 移除干扰字符串,解析JSON
raw_content = response.text
clean_content = raw_content.replace("])}while(1);</x>", "")
data = json.loads(clean_content)

# 提取文章列表,即可获取目标区块的所有资讯
articles = data["payload"]["references"]["Post"].values()
for article in articles:
    print(f"标题:{article['title']},发布时间:{article['latestPublishedAt']},链接:https://blog.coinbase.com/{article['uniqueSlug']}")

备选方案:Playwright实现

你原来的Playwright代码没有触发页面滚动,也没有等待目标元素渲染,修改后代码如下:

import asyncio
from playwright.async_api import async_playwright

async def parser():        
    page_path = "https://blog.coinbase.com/"        
    async with async_playwright() as p:          
        browser = await p.chromium.launch(headless=True)           
        page = await browser.new_page(user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')         
        await page.goto(page_path)
        # 滚动页面触发懒加载
        await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
        # 等待目标区块元素加载完成,超时时间设为10秒
        target_locator = page.locator('div.streamItem.streamItem--section.js-streamItem[data-action-scope="_actionscope_6"]')
        await target_locator.wait_for(timeout=10000)
        # 获取完整页面内容
        page_content = await page.content()            
        await browser.close()        
        print(page_content)    
        
asyncio.get_event_loop().run_until_complete(parser())

内容的提问来源于stack exchange,提问作者Workingsolutions

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 03:27:07