You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取Chrome浏览器中已打开的页面(无需requests)

可以读取Chrome已打开页面的Python方案

方法1:Selenium连接已启动的Chrome实例

这是最常用的方案,核心是通过Chrome的远程调试端口连接已打开的浏览器:

  1. 确保Chrome开启远程调试:
    如果还没打开Chrome,用命令行启动并指定调试端口(若已打开但未开端口,需关闭所有Chrome进程后重启):

    • Windows:chrome.exe --remote-debugging-port=9222 --user-data-dir="C:\ChromeProfile"
    • Linux/macOS:google-chrome --remote-debugging-port=9222 --user-data-dir="/tmp/ChromeProfile"
  2. 用Selenium连接并获取页面内容:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

# 配置连接参数
options = Options()
options.add_experimental_option("debuggerAddress", "127.0.0.1:9222")
# 连接已打开的Chrome
driver = webdriver.Chrome(options=options)

# 获取当前页面的源码
page_html = driver.page_source

# 结合BeautifulSoup解析(如需)
from bs4 import BeautifulSoup
soup = BeautifulSoup(page_html, "html.parser")
# 后续可使用bs4方法提取目标数据

方法2:通过Chrome DevTools Protocol直接连接

可以用pychrome库直接调用Chrome的调试协议,无需依赖Selenium:

  1. 安装依赖:
pip install pychrome
  1. 连接并获取页面源码:
from pychrome import Browser

# 连接到开启调试的Chrome
browser = Browser(url="http://127.0.0.1:9222")
# 获取第一个标签页(可根据需求选择其他标签)
tab = browser.list_tab()[0]
tab.start()

# 执行JS获取页面完整HTML
result = tab.Runtime.evaluate(expression="document.documentElement.outerHTML")
page_html = result["result"]["value"]

tab.stop()

注意事项

  • 必须保证Chrome启动时添加了--remote-debugging-port参数,否则无法建立连接。
  • 指定--user-data-dir是为了避免与默认Chrome配置冲突,确保调试实例独立运行。
  • 连接成功后不仅能获取源码,还可直接操作已打开的页面(如切换标签、执行自定义JS等)。

内容的提问来源于stack exchange,提问作者FGalanB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:58:02