JavaScript爬虫与HTML+CSS结合及用户输入传递问题求助
问题解决指南
1. 先搞定axios未初始化的错误
这个报错大概率是axios加载/声明的顺序不对:
- 如果是在浏览器里直接写代码:得先加载axios脚本,再调用你的爬虫函数。比如先加这句:
把它放在你的自定义脚本前面,确保调用<script src="https://cdn.jsdelivr.net/npm/axios/dist/axios.min.js"></script>scrapeSite()的时候axios已经加载好了。 - 如果是Node环境:检查你是不是用
const axios = require('axios')或者import axios from 'axios',而且这条语句必须写在scrapeSite()函数定义之前,不能在函数里才声明axios。
2. 必须用Node.js服务器连接前后端
别纠结了,肯定要做后端!浏览器有跨域限制,直接在前端用axios请求Google的接口会被拦下来,根本拿不到数据。正确的流程是前端把关键词发给Node后端,后端用axios去爬,再把结果返回给前端。
具体步骤:
第一步:搭Node后端
- 先初始化项目:
npm init -y npm install axios express cors - 创建
server.js文件,写后端逻辑:const express = require('express'); const axios = require('axios'); const cors = require('cors'); const app = express(); // 解决跨域问题 app.use(cors()); // 解析前端传的参数 app.use(express.json()); // 接收前端的搜索请求 app.post('/search-images', async (req, res) => { const keyword = req.body.keyword; try { // 调用你的爬虫函数 const imageData = await scrapeSite(keyword); res.send(imageData); } catch (err) { res.status(500).send('爬取出错:' + err.message); } }); // 你的爬虫函数,放在后端执行 async function scrapeSite(keyword) { // 这里写原来的爬取逻辑,比如构造Google图片的搜索URL const searchUrl = `https://www.google.com/search?q=${encodeURIComponent(keyword)}&tbm=isch`; // 加个浏览器请求头,避免被Google反爬 const response = await axios.get(searchUrl, { headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/114.0.0.0 Safari/537.36' } }); // 这里解析返回的HTML,提取图片链接等数据 // 比如用cheerio库解析,需要先npm install cheerio // const cheerio = require('cheerio'); // const $ = cheerio.load(response.data); // const images = []; // $('img').each((i, el) => images.push($(el).attr('src'))); // return { images }; return { message: '爬取成功,这里放解析后的图片数据' }; } // 启动服务器 app.listen(3000, () => { console.log('服务器跑在http://localhost:3000'); });
第二步:改前端代码
前端不要直接调用scrapeSite(),改成发请求给后端:
<form id="searchForm"> <input type="text" id="keyword" placeholder="输入要搜的图片关键词"> <button type="submit">搜索</button> </form> <script> document.getElementById('searchForm').addEventListener('submit', async (e) => { e.preventDefault(); const keyword = document.getElementById('keyword').value; try { // 发请求给后端接口 const res = await fetch('http://localhost:3000/search-images', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ keyword }) }); const data = await res.json(); // 拿到数据后展示,比如打印到控制台或者渲染到页面 console.log(data); } catch (err) { console.error('搜索失败:', err); } }); </script>
3. 额外提醒
Google有反爬机制,直接用axios请求容易被封IP,记得加User-Agent请求头模拟浏览器;另外,解析HTML可以用cheerio库,需要单独安装。
内容的提问来源于stack exchange,提问作者Jarod
相关产品推荐
相关产品推荐

