You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mac下Python脚本实现在已打开Chrome中打开URL的代码修改求助

解决Mac上用Python脚本连接已打开Chrome并打开指定URL的问题

前置操作(必须先做)

关闭所有Chrome窗口,打开终端执行以下命令启动带调试端口的Chrome:

/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222

这一步是让Chrome监听9222端口,允许脚本连接到已打开的实例。

原代码的核心问题及修改点

  • 未连接到已打开的Chrome:原代码直接启动了新的Chrome实例,需要添加debugger_address参数指定连接到本地9222端口的调试实例。
  • 缺失必要导入:原代码没写依赖的导入语句,需要补充。
  • 循环内新标签页逻辑冗余:最后一个URL处理完不需要再开新标签,避免生成多余空标签。
  • 未实现依赖函数:read_config和extract_urls需要自己实现,否则会报错。

修改后的完整代码

from selenium.webdriver.chrome.options import Options
from selenium import webdriver
import time
import json

# 实现read_config函数,读取配置文件
def read_config(config_path):
    try:
        with open(config_path, 'r') as f:
            return json.load(f)
    except FileNotFoundError:
        print(f"配置文件{config_path}未找到")
        return None

# 实现extract_urls函数(示例逻辑,可按需修改)
def extract_urls(html_content):
    from bs4 import BeautifulSoup  # 需要先安装:pip install beautifulsoup4
    soup = BeautifulSoup(html_content, 'html.parser')
    urls = []
    for a_tag in soup.find_all('a', href=True):
        url = a_tag['href']
        if url not in urls:
            urls.append(url)
    return urls

def main():
    # 接收用户输入的URL,支持空格分隔多个URL
    urls_input = input("Enter the URL to scrape: ")
    urls_list = urls_input.split()
    if not urls_list:
        print("请输入至少一个URL")
        return

    # 配置Selenium连接已打开的Chrome
    options = Options()
    # 指定连接到已启动的调试Chrome实例
    options.add_experimental_option("debuggerAddress", "127.0.0.1:9222")
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option("detach", True)

    try:
        driver = webdriver.Chrome(options=options)
    except Exception as e:
        print(f"连接Chrome失败:{e}")
        print("请确认已按照前置操作启动带调试端口的Chrome")
        return

    first_iter = True
    for idx, url in enumerate(urls_list):
        # 在当前标签打开目标URL
        driver.get(url)
        
        # 第一次运行时等待认证完成
        if first_iter:
            time_to_sleep = 80
            config_data = read_config('config.json')
            if config_data and 'time' in config_data:
                time_to_sleep = config_data['time']
            print(f"等待{time_to_sleep}秒以完成认证")
            time.sleep(time_to_sleep)
            first_iter = False
        
        # 提取页面HTML内容和页面内的URL
        html_content = driver.page_source
        extracted_urls = extract_urls(html_content)
        print(f"从{url}提取到的URL:{extracted_urls}")
        
        # 仅当不是最后一个URL时,打开新标签页
        if idx != len(urls_list) - 1:
            driver.execute_script("window.open('');")
            driver.switch_to.window(driver.window_handles[-1])

    # 可选:关闭所有多余标签页,保留第一个标签
    for handle in driver.window_handles[1:]:
        driver.switch_to.window(handle)
        driver.close()
    driver.switch_to.window(driver.window_handles[0])

if __name__ == "__main__":
    main()

额外说明

  • 未安装依赖库的话,执行:pip install selenium beautifulsoup4
  • extract_urls是示例实现,可根据自身需求调整URL提取逻辑
  • 如果连接失败,检查Chrome是否确实在9222端口运行,是否关闭所有Chrome窗口后再启动调试模式

内容的提问来源于stack exchange,提问作者Sandy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 03:20:57