如何将表单数据传入JS程序并实现指定URL网页数据抓取?
Hey there! Based on your existing code snippet, it looks like you want to build a tool where users input a URL via an HTML input field, then your Node.js script fetches and scrapes that target page. Let's break this down into frontend and backend parts to make it work end-to-end:
1. Frontend: HTML Page for User Input
First, we need a simple HTML interface where users can enter the URL and trigger the scrape. Since browsers block cross-origin requests directly, we'll send the URL to a Node.js backend which handles the actual scraping (avoids CORS issues).
Save this as public/index.html (we'll set up the backend to serve this file later):
<!DOCTYPE html> <html lang="en"> <head> <meta charset="UTF-8"> <title>Web Scraper Tool</title> </head> <body> <div style="margin: 20px;"> <label for="urlInput">Enter Target URL:</label> <input type="text" id="urlInput" style="width: 400px; padding: 8px;" placeholder="https://example.com"> <button onclick="startScraping()" style="padding: 8px 16px; margin-left: 10px;">Scrape Data</button> </div> <div id="scrapeResult" style="margin: 20px; padding: 15px; border: 1px solid #ddd;"></div> <script> async function startScraping() { const url = document.getElementById('urlInput').value.trim(); // Basic URL validation if (!url || !url.startsWith('http')) { alert('Please enter a valid URL starting with http/https'); return; } try { // Send URL to backend scrape endpoint const response = await fetch('/scrape', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ url }) }); const result = await response.json(); const resultDiv = document.getElementById('scrapeResult'); if (result.success) { resultDiv.innerHTML = ` <h3>Scraped Data:</h3> <pre style="white-space: pre-wrap;">${JSON.stringify(result.data, null, 2)}</pre> `; } else { resultDiv.innerHTML = `<p style="color: red;">Error: ${result.error}</p>`; } } catch (err) { document.getElementById('scrapeResult').innerHTML = `<p style="color: red;">Failed to connect to server. Please check if the backend is running.</p>`; } } </script> </body> </html>
2. Backend: Node.js Server with Scraping Logic
Your original code uses request, but that library is deprecated. I'll use axios (a modern alternative) and add Express to create a server that accepts the URL from the frontend, runs the scrape, and sends back the data.
First, install the required dependencies:
npm install express cheerio axios cors
Then create your main server file (e.g., server.js):
const express = require('express'); const cheerio = require('cheerio'); const axios = require('axios'); const cors = require('cors'); const app = express(); const PORT = 3000; // Middleware to handle cross-origin requests and JSON data app.use(cors()); app.use(express.json()); // Serve the frontend HTML file app.use(express.static('public')); // Scrape endpoint app.post('/scrape', async (req, res) => { const { url } = req.body; if (!url) { return res.status(400).json({ success: false, error: 'No URL provided' }); } try { // Fetch the target page (add a user-agent to avoid being blocked) const pageResponse = await axios.get(url, { headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } }); const $ = cheerio.load(pageResponse.data); const items = []; // Continue your scraping logic here (based on your original code) const allItems = $('.clearfix').parent().children(); allItems.each((index, element) => { // Customize this part based on what data you want to extract const itemText = $(element).text().trim(); const itemLink = $(element).find('a').attr('href'); if (itemText) { items.push({ position: index + 1, content: itemText, link: itemLink ? new URL(itemLink, url).href : null // Resolve relative links }); } }); res.json({ success: true, data: items }); } catch (err) { console.error('Scraping error:', err.message); res.status(500).json({ success: false, error: `Failed to scrape URL: ${err.message}` }); } }); // Start the server app.listen(PORT, () => { console.log(`Server running at http://localhost:${PORT}`); });
3. How to Run
- Create a
publicfolder in your project directory and place theindex.htmlfile inside it. - Save the server code as
server.js. - Run the server with
node server.js. - Open your browser and go to
http://localhost:3000– you can now enter a URL and scrape its data!
Key Notes
- Avoiding Blocking: Many sites block scrapers, so adding a valid
User-Agentheader helps mimic a real browser. - URL Validation: We added basic checks, but you can enhance this with a regex for stricter validation.
- Cheerio Logic: Adjust the
$('.clearfix').parent().children()part to match the actual elements you want to scrape from your target sites. - Error Handling: The code includes basic error handling for invalid URLs, server issues, and scraping failures.
内容的提问来源于stack exchange,提问作者Sunil kumar

