You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy与Python抓取JS生成的NCBI页面FTP列链接?

Hey, I've dealt with NCBI's dynamic pages like this before—those FTP links are definitely loaded client-side via JavaScript, so plain Scrapy won't pick them up out of the box. Here are three reliable approaches to get those links, ordered from easiest to most efficient:

Method 1: Use Scrapy with Playwright (Simplest Setup)

Playwright is a modern browser automation tool that integrates smoothly with Scrapy, and it's my go-to for dynamic content these days.

  1. Install dependencies
    First, grab the required packages:

    pip install scrapy-playwright playwright
    playwright install chromium
    
  2. Configure Scrapy to use Playwright
    Add these settings to your settings.py:

    DOWNLOAD_HANDLERS = {
        "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    }
    
    DOWNLOADER_MIDDLEWARES = {
        "scrapy.downloadermiddlewares.useragent.UserAgentMiddleware": None,
        "scrapy_playwright.middleware.PlaywrightMiddleware": 543,
    }
    
    PLAYWRIGHT_LAUNCH_OPTIONS = {
        "headless": True,  # Set to False if you want to see the browser in action
        "timeout": 30000,
    }
    
  3. Write your spider
    In your spider, tell Scrapy to use Playwright for the target URL by adding meta={"playwright": True} to your Request. Then you can parse the fully rendered page:

    import scrapy
    
    class NCBIGenomeSpider(scrapy.Spider):
        name = "ncbi_genome"
        start_urls = ["https://www.ncbi.nlm.nih.gov/genome/genomes/971"]
    
        def start_requests(self):
            for url in self.start_urls:
                yield scrapy.Request(url, meta={"playwright": True})
    
        def parse(self, response):
            # Adjust the selector to match the actual FTP link elements on the page
            for ftp_link in response.css("td a[href^='ftp://']::attr(href)"):
                yield {
                    "ftp_url": ftp_link.get()
                }
    

Method 2: Use Selenium with Scrapy

If you're more familiar with Selenium, you can set up a custom downloader middleware to handle JS rendering.

  1. Install dependencies

    pip install scrapy selenium
    # Don't forget to download ChromeDriver matching your Chrome browser version
    
  2. Create a Selenium middleware
    Add a new file middlewares.py in your project with this code:

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    from scrapy.http import HtmlResponse
    
    class SeleniumMiddleware:
        def __init__(self):
            chrome_options = Options()
            chrome_options.add_argument("--headless=new")
            self.driver = webdriver.Chrome(options=chrome_options)
    
        def process_request(self, request, spider):
            self.driver.get(request.url)
            body = self.driver.page_source
            return HtmlResponse(
                self.driver.current_url,
                body=body,
                encoding='utf-8',
                request=request
            )
    
        def __del__(self):
            self.driver.quit()
    
  3. Configure settings.py
    Add the middleware to your project settings:

    DOWNLOADER_MIDDLEWARES = {
        "your_project_name.middlewares.SeleniumMiddleware": 800,
    }
    
  4. Write the spider
    Your spider can now be a standard Scrapy spider—no extra meta needed, since the middleware handles rendering automatically:

    import scrapy
    
    class NCBIGenomeSpider(scrapy.Spider):
        name = "ncbi_genome"
        start_urls = ["https://www.ncbi.nlm.nih.gov/genome/genomes/971"]
    
        def parse(self, response):
            for ftp_link in response.css("td a[href^='ftp://']::attr(href)"):
                yield {"ftp_url": ftp_link.get()}
    

Method 3: Bypass JS by Calling the Backend API (Most Efficient)

This is the fastest method because you skip full browser rendering entirely—you just fetch raw data directly from NCBI's backend API.

  1. Find the API endpoint
    Open your browser's DevTools (F12), go to the Network tab, reload the page, and look for XHR/fetch requests returning JSON data. For this NCBI page, you'll spot an API endpoint that serves the genome data, including FTP links.

  2. Directly scrape the API
    Once you have the API URL, request it directly in your spider and parse the JSON response:

    import scrapy
    import json
    
    class NCBIGenomeAPISpider(scrapy.Spider):
        name = "ncbi_genome_api"
        # Replace with the actual API endpoint you found in DevTools
        start_urls = ["https://api.ncbi.nlm.nih.gov/genome/..."]
    
        def parse(self, response):
            data = json.loads(response.body)
            # Adjust this path to match the structure of the API's JSON response
            for genome in data["genomes"]:
                if "ftp_url" in genome:
                    yield {"ftp_url": genome["ftp_url"]}
    

    Pro tip: NCBI's APIs follow predictable patterns, so this method will be way faster than browser rendering—especially if you're scraping large datasets.


内容的提问来源于stack exchange,提问作者Sergii

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:21:37