如何用Shell/Python批量下载需登录的论坛帖子及静态资源
Hey there, let's work through your Futura Sciences scraping challenge—you've got two key roadblocks: getting past that login with the hidden password field, and properly archiving all those posts (with CSS included) locally. Here's a practical, step-by-step solution:
display:none Password Field) That hidden password field is a classic anti-scraping trick, but Selenium can work around it with a bit of JavaScript magic. The problem is Selenium won't interact with elements that aren't visible by default, so we need to modify the element's styling first, or set its value directly via JS.
Here's a code snippet that should work:
from selenium import webdriver from selenium.webdriver.common.by import By import time driver = webdriver.Chrome() driver.get("https://futura-sciences.com/login-url") # Replace with the actual login page URL # Fill in the visible username field username_field = driver.find_element(By.ID, "username") # Update selector to match the site's HTML username_field.send_keys("your-account-username") # Handle the hidden password field password_field = driver.find_element(By.ID, "password") # Update selector to match the site's HTML # Option 1: Make the field visible first so Selenium can interact with it driver.execute_script("arguments[0].style.display = 'block';", password_field) password_field.send_keys("your-account-password") # Option 2: Skip visibility checks and set the password directly via JS (sometimes more reliable) # driver.execute_script("arguments[0].value = 'your-account-password';", password_field) # Submit the login form login_button = driver.find_element(By.XPATH, "//button[@type='submit']") # Adjust selector as needed login_button.click() # Wait for post-login page to load driver.implicitly_wait(10)
Pro tip: If there's a CAPTCHA or 2-step verification, add a manual pause (time.sleep(30)) to complete that step manually before the script proceeds.
Once logged in, you need to systematically crawl post lists, handle multi-page posts, and save everything with CSS intact.
2.1 Crawl Post List Pages
First, iterate through the 6-7 list pages to collect all post URLs:
base_list_url = "https://futura-sciences.com/forum/category/page-" post_urls = [] for page_num in range(1, 8): # Covers pages 1 to 7 driver.get(f"{base_list_url}{page_num}") time.sleep(2) # Add delay to avoid triggering anti-scraping measures # Grab all post links on the page (update selector to match the site's HTML) post_links = driver.find_elements(By.CSS_SELECTOR, "a.post-title-link") for link in post_links: post_urls.append(link.get_attribute("href")) # Remove duplicate URLs just in case post_urls = list(set(post_urls))
2.2 Handle Multi-Page Posts
For each post, check for pagination and scrape all pages:
def scrape_full_post(driver, post_url): driver.get(post_url) time.sleep(2) all_post_content = [] while True: # Grab the current page's post content (update selector to match the site's HTML) content_block = driver.find_element(By.CSS_SELECTOR, "div.post-thread-container").get_attribute("outerHTML") all_post_content.append(content_block) # Check for a "Next Page" button try: next_button = driver.find_element(By.CSS_SELECTOR, "a.pagination-next") next_button.click() time.sleep(2) except: break # No more pages to scrape return "\n".join(all_post_content)
2.3 Save Pages with CSS & Assets
Wget's recursive download can get messy (grabbing unrelated site assets), so it's better to manually save HTML and fetch only necessary CSS files:
- Export Login Cookies (Optional, for faster non-Selenium requests)
After logging in with Selenium, export cookies to use withrequests(faster than Selenium for bulk scraping):import requests cookies = driver.get_cookies() session = requests.Session() for cookie in cookies: session.cookies.set(cookie['name'], cookie['value']) - Save HTML & Download CSS
For each post, save the HTML, download linked CSS files to a local folder, and update HTML links to point to local paths:import os from bs4 import BeautifulSoup # Create storage folders os.makedirs("futura_sciences_posts", exist_ok=True) os.makedirs("futura_sciences_posts/css", exist_ok=True) for idx, post_url in enumerate(post_urls): # Use requests for faster access (or stick with Selenium if needed) response = session.get(post_url) soup = BeautifulSoup(response.text, 'html.parser') # Download and localize CSS files for css_link in soup.find_all("link", rel="stylesheet"): css_url = css_link['href'] # Fix relative URLs if not css_url.startswith("http"): css_url = f"https://futura-sciences.com{css_url}" # Download the CSS file css_response = session.get(css_url) css_filename = os.path.basename(css_url) with open(f"futura_sciences_posts/css/{css_filename}", "w", encoding="utf-8") as f: f.write(css_response.text) # Update the HTML link to point to local CSS css_link['href'] = f"css/{css_filename}" # Save the modified HTML file with open(f"futura_sciences_posts/post_{idx+1}.html", "w", encoding="utf-8") as f: f.write(str(soup))
# Important Notes - **Rate Limiting**: Always add delays (`time.sleep(2-3)`) between requests to avoid getting IP-banned. Consider adding random delays for more natural behavior. - **Terms of Service**: Double-check the site's robots.txt and terms of service to ensure scraping is allowed. - **Selector Adjustments**: All CSS/XPath selectors in the code are examples—inspect the Futura Sciences site's HTML to find the actual selectors for fields, links, and content containers. 内容的提问来源于stack exchange,提问作者user1773603

