You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R抓取Java网站弹窗数据:ImmGen细胞类型信息提取需求

解决方案:抓取immgen网站细胞类型弹窗信息

Got it, let's walk through how to scrape those Short Name, Long Name, and Description fields from the popups on the Population Comparison page. Since the content loads dynamically via JavaScript, we'll use Selenium to simulate a real browser interaction—this is way more reliable than static scraping tools here.

步骤1:安装依赖

First off, make sure you have the required packages installed. Open your terminal and run:

pip install selenium webdriver-manager
  • selenium: Handles browser automation to interact with dynamic content
  • webdriver-manager: Automatically fetches and manages the correct browser driver for your system, so you don't have to manually download it

步骤2:编写抓取脚本

Here's a complete Python script that does the job, with comments to break down each part:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
import csv

# Initialize Chrome browser (works with Firefox too—swap to GeckoDriverManager if needed)
driver = webdriver.Chrome(ChromeDriverManager().install())
driver.get("http://rstats.immgen.org/PopulationComparison/")

# Wait for the main dropdown to load (adjust timeout if your internet is slow)
wait = WebDriverWait(driver, 10)
dropdown_trigger = wait.until(EC.element_to_be_clickable((By.ID, "populationSelect")))
dropdown_trigger.click()

# Grab all population code options from the dropdown
population_options = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#populationSelect option")))
population_codes = [opt.get_attribute("value") for opt in population_options if opt.get_attribute("value")]

# Prepare to store our scraped data
results = []

for code in population_codes:
    try:
        # Re-open the dropdown (it closes after each selection)
        dropdown_trigger = wait.until(EC.element_to_be_clickable((By.ID, "populationSelect")))
        dropdown_trigger.click()
        
        # Find and click the specific population code
        target_option = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, f"#populationSelect option[value='{code}']")))
        target_option.click()
        
        # Wait for the popup to load, then extract the three fields
        popup = wait.until(EC.presence_of_element_located((By.ID, "populationInfoModal")))
        
        short_name = popup.find_element(By.CSS_SELECTOR, ".modal-body p:nth-child(1)").text.replace("Short Name: ", "")
        long_name = popup.find_element(By.CSS_SELECTOR, ".modal-body p:nth-child(2)").text.replace("Long Name: ", "")
        description = popup.find_element(By.CSS_SELECTOR, ".modal-body p:nth-child(3)").text.replace("Description: ", "")
        
        # Add the data to our results list
        results.append({
            "Code": code,
            "Short Name": short_name,
            "Long Name": long_name,
            "Description": description
        })
        
        # Close the popup to move to the next code
        close_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".modal-footer .btn")))
        close_btn.click()
        
        print(f"Successfully scraped data for {code}")
        
    except Exception as e:
        print(f"Failed to scrape {code}: {str(e)}")
        continue

# Save all results to a CSV file for easy viewing/analysis
with open("immgen_population_data.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["Code", "Short Name", "Long Name", "Description"])
    writer.writeheader()
    writer.writerows(results)

# Clean up the browser session
driver.quit()
print(f"Scraping complete! Data saved to immgen_population_data.csv")

关键细节提醒

  • Dynamic Loading Handling: We use WebDriverWait instead of hardcoded time.sleep()—this ensures the script waits only as long as needed for elements to load, making it faster and more stable.
  • Dropdown Re-opening: The dropdown closes after each selection, so we have to re-trigger it for every code to avoid errors.
  • Error Resilience: The try/except block lets the script keep running even if one population code fails to load (e.g., due to a temporary glitch).
  • Browser Flexibility: If you prefer Firefox, just replace ChromeDriverManager with GeckoDriverManager and webdriver.Chrome with webdriver.Firefox.

无浏览器替代方案

If you don't want to use a browser automation tool, you can inspect the site's network traffic (via your browser's DevTools > Network tab) to find the AJAX endpoint that returns the popup data. Once you have that endpoint, you can use the requests library to call it directly with each population code—this is faster but requires a bit of digging into the site's backend logic.

内容的提问来源于stack exchange,提问作者Atakan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:45:58