You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用python-requests和BeautifulSoup抓取JS渲染页面的title内容

解决JS渲染页面title抓取问题的方案

首先明确前提:纯requests + BeautifulSoup 无法直接完成JS渲染页面的内容抓取。requests只能获取服务器返回的原始静态HTML,不会执行页面内嵌的JavaScript代码,而JS渲染页面的title是前端脚本运行后动态生成的,静态HTML中不存在最终的展示内容,必须搭配支持JS运行的无头浏览器工具实现抓取。

方案1:修正你当前使用的requests-html写法

你用的HTMLSession属于requests-html库,本身已经自带JS渲染能力,只需要调整取值逻辑即可拿到正确的title:

from requests_html import HTMLSession

session = HTMLSession()
resp = session.get("https://twitter.com/aProfile/")
# sleep参数用于等待页面JS执行完成,复杂页面可以适当调大数值
resp.html.render(sleep=2)
# 直接读取渲染后的页面title属性即可
title = resp.html.xpath('//title/text()')[0]
print(title)

方案2:结合BeautifulSoup解析渲染后的内容

如果你习惯用BeautifulSoup做解析,可以把渲染完成的完整HTML传给BeautifulSoup处理:

from requests_html import HTMLSession
from bs4 import BeautifulSoup

session = HTMLSession()
resp = session.get("https://twitter.com/aProfile/")
resp.html.render(sleep=2)
# 获取JS渲染后的完整HTML文本
rendered_html = resp.html.html
soup = BeautifulSoup(rendered_html, "html.parser")
title = soup.title.string
print(title)

方案3:替换为更稳定的无头浏览器工具

如果遇到反爬严格、渲染逻辑复杂的页面,可以使用playwright这类更成熟的无头浏览器工具,抓取成功率更高:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    # 启动无头Chromium浏览器
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    # 等待页面网络空闲后再读取内容
    page.goto("https://twitter.com/aProfile/", wait_until="networkidle")
    title = page.title()
    print(title)
    browser.close()

内容的提问来源于stack exchange,提问作者qwdoicjq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 16:09:00