You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取script标签内内容?有无document.getElementById替代方案

提取网页Script标签内特定字段的实现方案

能不能编程实现?

完全可以,而且有多种成熟的自动化方案,无需手动分析源码。

除正则外的技术方案

1. 浏览器自动化工具(基于JS但更可靠)

像Puppeteer、Playwright这类工具能模拟真实浏览器加载页面,直接访问页面渲染后的JS上下文,比正则解析更稳定:

  • 示例(Puppeteer):
const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch();
  const page = await browser.newPage();
  await page.goto('https://www.amazon.co.uk/dp/B0CRMCFCWB');
  
  const landingImageUrl = await page.evaluate(() => {
    const scripts = document.querySelectorAll('script');
    for (const script of scripts) {
      if (script.textContent.includes('landingImageUrl')) {
        try {
          // 匹配亚马逊script里的JSON结构
          const jsonStr = script.textContent.match(/window\['.*?'\s*=\s*({.*?});/s)[1];
          const data = JSON.parse(jsonStr);
          return data.landingImageUrl || data.image?.landingImageUrl;
        } catch (e) {
          continue;
        }
      }
    }
    return null;
  });

  console.log(landingImageUrl);
  await browser.close();
})();

这种方法能处理动态渲染内容,且直接在浏览器环境解析JS对象,出错概率远低于正则。

2. AppleScript方案(非JS工具)

Mac环境下可通过AppleScript调用Safari/Chrome执行提取逻辑:

  • 示例(Safari):
tell application "Safari"
  activate
  open location "https://www.amazon.co.uk/dp/B0CRMCFCWB"
  delay 5 -- 等待页面加载完成
  tell document 1
    set landingImageUrl to do JavaScript "
      const scripts = document.querySelectorAll('script');
      for (const script of scripts) {
        if (script.textContent.includes('landingImageUrl')) {
          try {
            const jsonStr = script.textContent.match(/window\\['.*?'\\s*=\\s*({.*?});/s)[1];
            const data = JSON.parse(jsonStr);
            return data.landingImageUrl || data.image?.landingImageUrl;
          } catch (e) {
            continue;
          }
        }
      }
      return null;
    "
    log landingImageUrl
  end tell
end tell

无需额外安装工具,直接利用Mac原生能力完成提取。

3. HTML解析库+JSON解析(比正则稳健)

用BeautifulSoup(Python)、Cheerio(Node.js)等库先提取script标签内容,再用JSON工具解析:

  • 示例(Python + BeautifulSoup):
import requests
from bs4 import BeautifulSoup
import re
import json

url = 'https://www.amazon.co.uk/dp/B0CRMCFCWB'
headers = {
  'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

landing_image_url = None
for script in soup.find_all('script'):
  if script.string and 'landingImageUrl' in script.string:
    match = re.search(r'window\[\'.*?\'\s*=\s*({.*?});', script.string, re.DOTALL)
    if match:
      try:
        data = json.loads(match.group(1))
        landing_image_url = data.get('landingImageUrl') or data.get('image', {}).get('landingImageUrl')
        break
      except json.JSONDecodeError:
        continue

print(landing_image_url)

HTML解析库能正确处理标签格式,再配合JSON解析避免正则匹配JSON的各种漏洞。

注意事项

  • 亚马逊页面结构可能变动,需定期验证提取逻辑
  • 请求时需设置合理的User-Agent,避免被反爬拦截
  • 部分script内容为压缩格式,上述方案可自动适配多数场景

内容的提问来源于stack exchange,提问作者Sockie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 20:43:34