You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Chrome扩展中BackgroundPage如何打开URL并实现站点内容爬取

Chrome扩展通过后台页面爬取指定站点内容实现方案

前置注意

你直接在BackgroundPage上下文调用document.querySelectorAll获取的是Background页面自身的DOM元素,并不是你要访问的目标站点的DOM,这是最容易踩的实现误区。

另外Chrome扩展Manifest V3版本已正式弃用chrome.extension.getBackgroundPage()方法,以下方案按Manifest版本区分:

Manifest V2版本实现

你可以选择两种常用实现路径:

路径1:隐藏iframe加载(无需打开可见标签)

首先在Background脚本中提前定义公用爬取方法:

// background.js 中写入
async function crawlTargetUrl(targetUrl, domSelector) {
  return new Promise((resolve, reject) => {
    const crawlIframe = document.createElement('iframe')
    crawlIframe.src = targetUrl
    crawlIframe.style.display = 'none'
    crawlIframe.onload = () => {
      const targetDoc = crawlIframe.contentDocument || crawlIframe.contentWindow.document
      const result = []
      targetDoc.querySelectorAll(domSelector).forEach(el => {
        // 此处自定义你要提取的元素内容,比如文本、指定属性值等
        result.push(el.textContent.trim())
      })
      // 用完移除iframe释放内存
      document.body.removeChild(crawlIframe)
      resolve(result)
    }
    crawlIframe.onerror = reject
    document.body.appendChild(crawlIframe)
  })
}

之后在你获取Background实例的页面直接调用即可:

var BG = chrome.extension.getBackgroundPage();
// 调用示例
async function getProductList() {
  try {
    const productList = await BG.crawlTargetUrl('替换为你要爬取的目标URL', '[role="tgk"]')
    // 这里写你拿到爬取结果后的业务逻辑
    console.log(productList)
  } catch (err) {
    console.error('爬取失败:', err)
  }
}
getProductList()

路径2:后台标签加载(适合有反爬、需要执行页面JS的场景)

同样先在Background脚本中定义方法:

// background.js 中写入
async function crawlByBackgroundTab(targetUrl, domSelector) {
  return new Promise((resolve, reject) => {
    // 创建active为false的标签,用户侧不可见
    chrome.tabs.create({url: targetUrl, active: false}, (tab) => {
      chrome.tabs.onUpdated.addListener(function updateListener(tabId, info) {
        if (tabId === tab.id && info.status === 'complete') {
          chrome.tabs.onUpdated.removeListener(updateListener)
          // 注入脚本提取内容
          chrome.tabs.executeScript(tab.id, {
            code: `Array.from(document.querySelectorAll('${domSelector}')).map(el => el.textContent.trim())`
          }, (res) => {
            // 用完关闭后台标签
            chrome.tabs.remove(tab.id)
            if (chrome.runtime.lastError) return reject(chrome.runtime.lastError)
            resolve(res[0])
          })
        }
      })
    })
  })
}

调用方式和路径1一致,直接调用BG.crawlByBackgroundTab方法传参即可。

Manifest V3版本实现

V3版本后台替换为无DOM环境的Service Worker,已不支持chrome.extension.getBackgroundPage(),直接使用后台标签+脚本注入的方案,通过chrome.runtime.sendMessage给后台发消息触发爬取逻辑即可。

必要配置

  • V2版本需要在manifest.json的permissions数组中添加你要爬取的目标站点域名权限
  • V3版本需要在manifest.json的host_permissions数组中添加目标站点域名权限
  • 爬取行为需要符合目标站点的robots协议及相关法律法规要求

内容的提问来源于stack exchange,提问作者Dawg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 21:06:10