创建HTML页面提取指定类元素内容数组 加速免费数据库录入
Simple HTML Tool to Extract Specific Elements from a Target Page
Got it, let's build this tool together. Since all your target pages share the same structure, we can create a self-contained HTML page that handles URL input, fetches the page, and pulls out the content you need into an array.
How it works
We'll use vanilla JavaScript (no frameworks needed) to:
- Accept a URL input from you
- Fetch the target page's HTML
- Parse the HTML and select all elements with your target class (like
.FamilyName) - Extract their text content into a clean array
- Display the results clearly
Full Code
Here's the complete HTML file you can save and run directly in your browser:
<!DOCTYPE html> <html lang="en"> <head> <meta charset="UTF-8"> <meta name="viewport" content="width=device-width, initial-scale=1.0"> <title>Element Extractor Tool</title> <style> body { max-width: 800px; margin: 2rem auto; padding: 0 1rem; font-family: Arial, sans-serif; } .input-section { margin-bottom: 1.5rem; } input[type="url"] { width: 70%; padding: 0.8rem; font-size: 1rem; } button { padding: 0.8rem 1.5rem; font-size: 1rem; background: #007bff; color: white; border: none; border-radius: 4px; cursor: pointer; } button:hover { background: #0056b3; } .results-section { margin-top: 2rem; padding: 1rem; border: 1px solid #eee; border-radius: 4px; } .error { color: #dc3545; } .success { color: #28a745; } pre { background: #f8f9fa; padding: 1rem; border-radius: 4px; overflow-x: auto; } </style> </head> <body> <h1>Target Page Element Extractor</h1> <div class="input-section"> <label for="targetUrl">Enter Target Page URL:</label><br> <input type="url" id="targetUrl" placeholder="https://example.com/target-page" required> <button id="extractBtn">Extract Elements</button> </div> <div class="results-section"> <h2>Extracted Content Array</h2> <div id="output"></div> </div> <script> // Grab DOM elements const urlInput = document.getElementById('targetUrl'); const extractBtn = document.getElementById('extractBtn'); const outputDiv = document.getElementById('output'); extractBtn.addEventListener('click', async () => { const url = urlInput.value.trim(); if (!url) { outputDiv.innerHTML = '<p class="error">Please enter a valid URL.</p>'; return; } try { outputDiv.innerHTML = '<p>Fetching and parsing page...</p>'; // Fetch the target page const response = await fetch(url); if (!response.ok) { throw new Error(`HTTP error! Status: ${response.status}`); } const htmlText = await response.text(); // Parse HTML into a manipulable DOM document const parser = new DOMParser(); const doc = parser.parseFromString(htmlText, 'text/html'); // Target the specific elements using your class structure // Adjust this selector to match your exact needs (e.g., .section .article .author .FamilyName) const targetElements = doc.querySelectorAll('.FamilyName'); if (targetElements.length === 0) { outputDiv.innerHTML = '<p class="error">No elements found with the specified class.</p>'; return; } // Convert elements to a clean text array const contentArray = Array.from(targetElements).map(el => el.textContent.trim()); // Display the formatted result outputDiv.innerHTML = ` <p class="success">Successfully extracted ${contentArray.length} elements:</p> <pre>${JSON.stringify(contentArray, null, 2)}</pre> `; } catch (error) { outputDiv.innerHTML = `<p class="error">Error: ${error.message}</p>`; console.error('Extraction failed:', error); } }); </script> </body> </html>
Key Tips
- Tweak the selector: For your example structure, replace
.FamilyNamewith.section .article .author .FamilyNamein thequerySelectorAllcall to ensure you only pick up elements from the exact hierarchy you want. - CORS heads-up: Most public websites block cross-origin requests for security. If you hit a CORS error, try:
- Running this tool as a browser extension (extensions have relaxed CORS rules)
- Using a simple local proxy server to fetch the page for you
- Adding CORS headers to the target pages if you control them
- Clean up text: The code uses
trim()to strip extra whitespace—you can remove this or add more formatting logic if you need to preserve line breaks or other spacing.
Save this code as element-extractor.html, open it in your browser, and you'll be able to skip those tedious copy-paste tasks in no time!
内容的提问来源于stack exchange,提问作者Ádám Nagy
相关产品推荐
相关产品推荐

