You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

创建HTML页面提取指定类元素内容数组 加速免费数据库录入

Simple HTML Tool to Extract Specific Elements from a Target Page

Got it, let's build this tool together. Since all your target pages share the same structure, we can create a self-contained HTML page that handles URL input, fetches the page, and pulls out the content you need into an array.

How it works

We'll use vanilla JavaScript (no frameworks needed) to:

  1. Accept a URL input from you
  2. Fetch the target page's HTML
  3. Parse the HTML and select all elements with your target class (like .FamilyName)
  4. Extract their text content into a clean array
  5. Display the results clearly

Full Code

Here's the complete HTML file you can save and run directly in your browser:

<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width, initial-scale=1.0">
    <title>Element Extractor Tool</title>
    <style>
        body {
            max-width: 800px;
            margin: 2rem auto;
            padding: 0 1rem;
            font-family: Arial, sans-serif;
        }
        .input-section {
            margin-bottom: 1.5rem;
        }
        input[type="url"] {
            width: 70%;
            padding: 0.8rem;
            font-size: 1rem;
        }
        button {
            padding: 0.8rem 1.5rem;
            font-size: 1rem;
            background: #007bff;
            color: white;
            border: none;
            border-radius: 4px;
            cursor: pointer;
        }
        button:hover {
            background: #0056b3;
        }
        .results-section {
            margin-top: 2rem;
            padding: 1rem;
            border: 1px solid #eee;
            border-radius: 4px;
        }
        .error {
            color: #dc3545;
        }
        .success {
            color: #28a745;
        }
        pre {
            background: #f8f9fa;
            padding: 1rem;
            border-radius: 4px;
            overflow-x: auto;
        }
    </style>
</head>
<body>
    <h1>Target Page Element Extractor</h1>
    
    <div class="input-section">
        <label for="targetUrl">Enter Target Page URL:</label><br>
        <input type="url" id="targetUrl" placeholder="https://example.com/target-page" required>
        <button id="extractBtn">Extract Elements</button>
    </div>

    <div class="results-section">
        <h2>Extracted Content Array</h2>
        <div id="output"></div>
    </div>

    <script>
        // Grab DOM elements
        const urlInput = document.getElementById('targetUrl');
        const extractBtn = document.getElementById('extractBtn');
        const outputDiv = document.getElementById('output');

        extractBtn.addEventListener('click', async () => {
            const url = urlInput.value.trim();
            if (!url) {
                outputDiv.innerHTML = '<p class="error">Please enter a valid URL.</p>';
                return;
            }

            try {
                outputDiv.innerHTML = '<p>Fetching and parsing page...</p>';
                
                // Fetch the target page
                const response = await fetch(url);
                if (!response.ok) {
                    throw new Error(`HTTP error! Status: ${response.status}`);
                }
                const htmlText = await response.text();

                // Parse HTML into a manipulable DOM document
                const parser = new DOMParser();
                const doc = parser.parseFromString(htmlText, 'text/html');

                // Target the specific elements using your class structure
                // Adjust this selector to match your exact needs (e.g., .section .article .author .FamilyName)
                const targetElements = doc.querySelectorAll('.FamilyName');
                
                if (targetElements.length === 0) {
                    outputDiv.innerHTML = '<p class="error">No elements found with the specified class.</p>';
                    return;
                }

                // Convert elements to a clean text array
                const contentArray = Array.from(targetElements).map(el => el.textContent.trim());

                // Display the formatted result
                outputDiv.innerHTML = `
                    <p class="success">Successfully extracted ${contentArray.length} elements:</p>
                    <pre>${JSON.stringify(contentArray, null, 2)}</pre>
                `;

            } catch (error) {
                outputDiv.innerHTML = `<p class="error">Error: ${error.message}</p>`;
                console.error('Extraction failed:', error);
            }
        });
    </script>
</body>
</html>

Key Tips

  • Tweak the selector: For your example structure, replace .FamilyName with .section .article .author .FamilyName in the querySelectorAll call to ensure you only pick up elements from the exact hierarchy you want.
  • CORS heads-up: Most public websites block cross-origin requests for security. If you hit a CORS error, try:
    • Running this tool as a browser extension (extensions have relaxed CORS rules)
    • Using a simple local proxy server to fetch the page for you
    • Adding CORS headers to the target pages if you control them
  • Clean up text: The code uses trim() to strip extra whitespace—you can remove this or add more formatting logic if you need to preserve line breaks or other spacing.

Save this code as element-extractor.html, open it in your browser, and you'll be able to skip those tedious copy-paste tasks in no time!

内容的提问来源于stack exchange,提问作者Ádám Nagy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:03:23