You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用纯JavaScript跨域爬取网页遇Network Error,如何解决?

跨域爬取网络错误的纯JS解决方案疑问

问题描述

尝试爬取简单网站时,同域环境下代码可正常运行,但跨域爬取时出现网络错误:

TypeError: NetworkError when attempting to fetch resource.

使用的代码如下:

async function scrapeWebPage(url) {
  try {
    const response = await fetch(url);
    const html = await response.text();
    const parser = new DOMParser();
    const doc = parser.parseFromString(html, 'text/html');

    // 提取页面所有链接的title属性
    const titles = [];
    const links = doc.querySelectorAll('a');
    links.forEach(link => {
      titles.push(link.getAttribute('title'));
    });

    console.log(titles);
    alert(titles);
    
  } catch (error) {
    console.error('Error scraping webpage:', error);
    alert(error);
  }
}

scrapeWebPage('https://www.w3schools.com/');

解答

浏览器环境下的纯前端JavaScript无法直接解决这个跨域错误。

原因是浏览器的同源策略限制:浏览器会阻止前端脚本从一个域名请求另一个域名的资源,除非目标服务器配置了CORS(跨域资源共享)响应头,明确允许你的域名发起跨域请求。如果目标网站(比如示例中的w3schools)没有配置允许你域名的CORS规则,前端纯JS根本绕不开这个限制。

如果要实现跨域爬取,只能借助后端代理:

  1. 用JavaScript写一个后端服务(比如基于Node.js的Express框架),让后端去请求目标网站的资源;
  2. 前端代码改为请求自己的后端接口,由后端返回爬取到的数据。

举个简单的Node.js代理示例:

// 后端代码(Node.js + Express)
const express = require('express');
const fetch = require('node-fetch');
const app = express();
const port = 3000;

app.get('/scrape', async (req, res) => {
  try {
    const url = req.query.url;
    const response = await fetch(url);
    const html = await response.text();
    res.send(html);
  } catch (error) {
    res.status(500).send(error.message);
  }
});

app.listen(port, () => {
  console.log(`Proxy server running on port ${port}`);
});

然后前端代码修改为请求这个代理接口:

async function scrapeWebPage(url) {
  try {
    // 请求自己的后端代理接口
    const response = await fetch(`http://localhost:3000/scrape?url=${encodeURIComponent(url)}`);
    const html = await response.text();
    const parser = new DOMParser();
    const doc = parser.parseFromString(html, 'text/html');

    const titles = [];
    const links = doc.querySelectorAll('a');
    links.forEach(link => {
      titles.push(link.getAttribute('title'));
    });

    console.log(titles);
    alert(titles);
    
  } catch (error) {
    console.error('Error scraping webpage:', error);
    alert(error);
  }
}

scrapeWebPage('https://www.w3schools.com/');

这样就能绕过浏览器的同源限制,实现跨域爬取。


内容的提问来源于stack exchange,提问作者Learning

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:52:13