You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Shell/Python批量下载需登录的论坛帖子及静态资源

Hey there, let's work through your Futura Sciences scraping challenge—you've got two key roadblocks: getting past that login with the hidden password field, and properly archiving all those posts (with CSS included) locally. Here's a practical, step-by-step solution:

1. Fixing the Login Issue (Selenium + display:none Password Field)

That hidden password field is a classic anti-scraping trick, but Selenium can work around it with a bit of JavaScript magic. The problem is Selenium won't interact with elements that aren't visible by default, so we need to modify the element's styling first, or set its value directly via JS.

Here's a code snippet that should work:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time

driver = webdriver.Chrome()
driver.get("https://futura-sciences.com/login-url") # Replace with the actual login page URL

# Fill in the visible username field
username_field = driver.find_element(By.ID, "username") # Update selector to match the site's HTML
username_field.send_keys("your-account-username")

# Handle the hidden password field
password_field = driver.find_element(By.ID, "password") # Update selector to match the site's HTML
# Option 1: Make the field visible first so Selenium can interact with it
driver.execute_script("arguments[0].style.display = 'block';", password_field)
password_field.send_keys("your-account-password")

# Option 2: Skip visibility checks and set the password directly via JS (sometimes more reliable)
# driver.execute_script("arguments[0].value = 'your-account-password';", password_field)

# Submit the login form
login_button = driver.find_element(By.XPATH, "//button[@type='submit']") # Adjust selector as needed
login_button.click()

# Wait for post-login page to load
driver.implicitly_wait(10)

Pro tip: If there's a CAPTCHA or 2-step verification, add a manual pause (time.sleep(30)) to complete that step manually before the script proceeds.

2. Batch Scraping & Static Local Storage

Once logged in, you need to systematically crawl post lists, handle multi-page posts, and save everything with CSS intact.

2.1 Crawl Post List Pages

First, iterate through the 6-7 list pages to collect all post URLs:

base_list_url = "https://futura-sciences.com/forum/category/page-"
post_urls = []

for page_num in range(1, 8): # Covers pages 1 to 7
    driver.get(f"{base_list_url}{page_num}")
    time.sleep(2) # Add delay to avoid triggering anti-scraping measures
    
    # Grab all post links on the page (update selector to match the site's HTML)
    post_links = driver.find_elements(By.CSS_SELECTOR, "a.post-title-link")
    for link in post_links:
        post_urls.append(link.get_attribute("href"))

# Remove duplicate URLs just in case
post_urls = list(set(post_urls))

2.2 Handle Multi-Page Posts

For each post, check for pagination and scrape all pages:

def scrape_full_post(driver, post_url):
    driver.get(post_url)
    time.sleep(2)
    
    all_post_content = []
    while True:
        # Grab the current page's post content (update selector to match the site's HTML)
        content_block = driver.find_element(By.CSS_SELECTOR, "div.post-thread-container").get_attribute("outerHTML")
        all_post_content.append(content_block)
        
        # Check for a "Next Page" button
        try:
            next_button = driver.find_element(By.CSS_SELECTOR, "a.pagination-next")
            next_button.click()
            time.sleep(2)
        except:
            break # No more pages to scrape
    
    return "\n".join(all_post_content)

2.3 Save Pages with CSS & Assets

Wget's recursive download can get messy (grabbing unrelated site assets), so it's better to manually save HTML and fetch only necessary CSS files:

  1. Export Login Cookies (Optional, for faster non-Selenium requests)
    After logging in with Selenium, export cookies to use with requests (faster than Selenium for bulk scraping):
    import requests
    
    cookies = driver.get_cookies()
    session = requests.Session()
    for cookie in cookies:
        session.cookies.set(cookie['name'], cookie['value'])
    
  2. Save HTML & Download CSS
    For each post, save the HTML, download linked CSS files to a local folder, and update HTML links to point to local paths:
    import os
    from bs4 import BeautifulSoup
    
    # Create storage folders
    os.makedirs("futura_sciences_posts", exist_ok=True)
    os.makedirs("futura_sciences_posts/css", exist_ok=True)
    
    for idx, post_url in enumerate(post_urls):
        # Use requests for faster access (or stick with Selenium if needed)
        response = session.get(post_url)
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # Download and localize CSS files
        for css_link in soup.find_all("link", rel="stylesheet"):
            css_url = css_link['href']
            # Fix relative URLs
            if not css_url.startswith("http"):
                css_url = f"https://futura-sciences.com{css_url}"
            
            # Download the CSS file
            css_response = session.get(css_url)
            css_filename = os.path.basename(css_url)
            with open(f"futura_sciences_posts/css/{css_filename}", "w", encoding="utf-8") as f:
                f.write(css_response.text)
            
            # Update the HTML link to point to local CSS
            css_link['href'] = f"css/{css_filename}"
        
        # Save the modified HTML file
        with open(f"futura_sciences_posts/post_{idx+1}.html", "w", encoding="utf-8") as f:
            f.write(str(soup))
    
# Important Notes
- **Rate Limiting**: Always add delays (`time.sleep(2-3)`) between requests to avoid getting IP-banned. Consider adding random delays for more natural behavior.
- **Terms of Service**: Double-check the site's robots.txt and terms of service to ensure scraping is allowed.
- **Selector Adjustments**: All CSS/XPath selectors in the code are examples—inspect the Futura Sciences site's HTML to find the actual selectors for fields, links, and content containers.

内容的提问来源于stack exchange,提问作者user1773603
相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:12:15