You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将表单数据传入JS程序并实现指定URL网页数据抓取?

实现用户输入URL后抓取目标网页数据的完整方案

Hey there! Based on your existing code snippet, it looks like you want to build a tool where users input a URL via an HTML input field, then your Node.js script fetches and scrapes that target page. Let's break this down into frontend and backend parts to make it work end-to-end:

1. Frontend: HTML Page for User Input

First, we need a simple HTML interface where users can enter the URL and trigger the scrape. Since browsers block cross-origin requests directly, we'll send the URL to a Node.js backend which handles the actual scraping (avoids CORS issues).

Save this as public/index.html (we'll set up the backend to serve this file later):

<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <title>Web Scraper Tool</title>
</head>
<body>
    <div style="margin: 20px;">
        <label for="urlInput">Enter Target URL:</label>
        <input type="text" id="urlInput" style="width: 400px; padding: 8px;" placeholder="https://example.com">
        <button onclick="startScraping()" style="padding: 8px 16px; margin-left: 10px;">Scrape Data</button>
    </div>
    <div id="scrapeResult" style="margin: 20px; padding: 15px; border: 1px solid #ddd;"></div>

    <script>
        async function startScraping() {
            const url = document.getElementById('urlInput').value.trim();
            
            // Basic URL validation
            if (!url || !url.startsWith('http')) {
                alert('Please enter a valid URL starting with http/https');
                return;
            }

            try {
                // Send URL to backend scrape endpoint
                const response = await fetch('/scrape', {
                    method: 'POST',
                    headers: { 'Content-Type': 'application/json' },
                    body: JSON.stringify({ url })
                });

                const result = await response.json();
                const resultDiv = document.getElementById('scrapeResult');
                
                if (result.success) {
                    resultDiv.innerHTML = `
                        <h3>Scraped Data:</h3>
                        <pre style="white-space: pre-wrap;">${JSON.stringify(result.data, null, 2)}</pre>
                    `;
                } else {
                    resultDiv.innerHTML = `<p style="color: red;">Error: ${result.error}</p>`;
                }
            } catch (err) {
                document.getElementById('scrapeResult').innerHTML = `<p style="color: red;">Failed to connect to server. Please check if the backend is running.</p>`;
            }
        }
    </script>
</body>
</html>

2. Backend: Node.js Server with Scraping Logic

Your original code uses request, but that library is deprecated. I'll use axios (a modern alternative) and add Express to create a server that accepts the URL from the frontend, runs the scrape, and sends back the data.

First, install the required dependencies:

npm install express cheerio axios cors

Then create your main server file (e.g., server.js):

const express = require('express');
const cheerio = require('cheerio');
const axios = require('axios');
const cors = require('cors');
const app = express();
const PORT = 3000;

// Middleware to handle cross-origin requests and JSON data
app.use(cors());
app.use(express.json());
// Serve the frontend HTML file
app.use(express.static('public'));

// Scrape endpoint
app.post('/scrape', async (req, res) => {
    const { url } = req.body;

    if (!url) {
        return res.status(400).json({ success: false, error: 'No URL provided' });
    }

    try {
        // Fetch the target page (add a user-agent to avoid being blocked)
        const pageResponse = await axios.get(url, {
            headers: {
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            }
        });

        const $ = cheerio.load(pageResponse.data);
        const items = [];

        // Continue your scraping logic here (based on your original code)
        const allItems = $('.clearfix').parent().children();
        allItems.each((index, element) => {
            // Customize this part based on what data you want to extract
            const itemText = $(element).text().trim();
            const itemLink = $(element).find('a').attr('href');
            
            if (itemText) {
                items.push({
                    position: index + 1,
                    content: itemText,
                    link: itemLink ? new URL(itemLink, url).href : null // Resolve relative links
                });
            }
        });

        res.json({ success: true, data: items });
    } catch (err) {
        console.error('Scraping error:', err.message);
        res.status(500).json({ success: false, error: `Failed to scrape URL: ${err.message}` });
    }
});

// Start the server
app.listen(PORT, () => {
    console.log(`Server running at http://localhost:${PORT}`);
});

3. How to Run

  1. Create a public folder in your project directory and place the index.html file inside it.
  2. Save the server code as server.js.
  3. Run the server with node server.js.
  4. Open your browser and go to http://localhost:3000 – you can now enter a URL and scrape its data!

Key Notes

  • Avoiding Blocking: Many sites block scrapers, so adding a valid User-Agent header helps mimic a real browser.
  • URL Validation: We added basic checks, but you can enhance this with a regex for stricter validation.
  • Cheerio Logic: Adjust the $('.clearfix').parent().children() part to match the actual elements you want to scrape from your target sites.
  • Error Handling: The code includes basic error handling for invalid URLs, server issues, and scraping failures.

内容的提问来源于stack exchange,提问作者Sunil kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:02:38