You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:无法用Selenium时,如何通过webbrowser关联其他库获取网页内容?

问题描述

我用以下Python代码通过webbrowser模块打开网页:

import webbrowser

url = 'http://docs.python.org/'
# MacOS
#chrome_path = 'open -a /Applications/Google\ Chrome.app %s'

# Windows
chrome_path = 'C:/Program Files (x86)/Google/Chrome/Application/chrome.exe %s'

# Linux
# chrome_path = '/usr/bin/google-chrome %s'

webbrowser.get(chrome_path).open(url)

发现webbrowser模块无法像Selenium那样获取网页内容,且某些场景下无法使用Selenium。想问是否可以获取sessionID或windowID,结合其他库来获取网页内容,甚至实现点击行展开数据的操作?

可行方案

1. 利用Chrome DevTools Protocol(CDP)连接已打开的浏览器

如果已经用webbrowser打开Chrome,只需启动Chrome时开启远程调试端口,就能通过CDP工具连接并控制浏览器:

  • 修改Chrome启动命令,添加--remote-debugging-port=9222参数,Windows下示例:
    chrome_path = 'C:/Program Files (x86)/Google/Chrome/Application/chrome.exe --remote-debugging-port=9222 %s'
    
  • 使用chrome-remote-interface库连接调试端口,获取窗口session并操作页面:
    import chrome_remote_interface as cri
    import asyncio
    
    async def operate_page():
        # 连接到远程调试端口
        async with cri.connect(host='localhost', port=9222) as conn:
            # 获取第一个页面目标
            targets = await conn.target.getTargets()
            page_target = next(t for t in targets if t['type'] == 'page')
            # 附加到目标页面
            async with cri.Page(conn, target_id=page_target['targetId']):
                # 获取页面DOM内容
                dom_content = await cri.Page.getDocument()
                print(dom_content)
                # 模拟点击操作(替换成目标元素选择器)
                await cri.Runtime.evaluate(expression="document.querySelector('.expand-row').click()")
    
    asyncio.run(operate_page())
    

2. 获取session相关数据

通过CDP的Network域可捕获浏览器请求中的Cookie、sessionID等信息:

import chrome_remote_interface as cri
import asyncio

async def get_session_info():
    async with cri.connect(host='localhost', port=9222) as conn:
        targets = await conn.target.getTargets()
        page_target = next(t for t in targets if t['type'] == 'page')
        async with cri.Page(conn, target_id=page_target['targetId']), cri.Network(conn):
            await cri.Network.enable()
            # 监听请求完成事件,提取Cookie
            async def on_request_done(event):
                cookies = await cri.Network.getCookies(urls=[event['request']['url']])
                print("Session关联Cookie:", cookies)
            conn.Network.requestFinished += on_request_done
            # 等待页面加载完成
            await cri.Page.loadEventFired()

asyncio.run(get_session_info())

3. 替代方案:用Pyppeteer直接控制浏览器

若可提前控制浏览器启动,Pyppeteer(支持有头/无头模式)比webbrowser更灵活,能直接获取内容和模拟交互:

import asyncio
from pyppeteer import launch

async def main():
    browser = await launch(executablePath='C:/Program Files (x86)/Google/Chrome/Application/chrome.exe', headless=False)
    page = await browser.newPage()
    await page.goto('http://docs.python.org/')
    # 获取页面完整HTML内容
    page_content = await page.content()
    print(page_content)
    # 模拟点击展开行(替换成实际元素选择器)
    await page.click('.expand-row')
    await browser.close()

asyncio.run(main())

注意事项

  • 确保Chrome版本与CDP类库版本兼容,避免接口不匹配问题
  • 启动Chrome时必须开启远程调试端口,否则无法连接已打开的窗口
  • Linux或MacOS用户需对应调整Chrome路径和启动参数

内容的提问来源于stack exchange,提问作者user982455

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 21:22:56