如何用Scrapy与Python抓取JS生成的NCBI页面FTP列链接?
Hey, I've dealt with NCBI's dynamic pages like this before—those FTP links are definitely loaded client-side via JavaScript, so plain Scrapy won't pick them up out of the box. Here are three reliable approaches to get those links, ordered from easiest to most efficient:
Method 1: Use Scrapy with Playwright (Simplest Setup)
Playwright is a modern browser automation tool that integrates smoothly with Scrapy, and it's my go-to for dynamic content these days.
Install dependencies
First, grab the required packages:pip install scrapy-playwright playwright playwright install chromiumConfigure Scrapy to use Playwright
Add these settings to yoursettings.py:DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } DOWNLOADER_MIDDLEWARES = { "scrapy.downloadermiddlewares.useragent.UserAgentMiddleware": None, "scrapy_playwright.middleware.PlaywrightMiddleware": 543, } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, # Set to False if you want to see the browser in action "timeout": 30000, }Write your spider
In your spider, tell Scrapy to use Playwright for the target URL by addingmeta={"playwright": True}to your Request. Then you can parse the fully rendered page:import scrapy class NCBIGenomeSpider(scrapy.Spider): name = "ncbi_genome" start_urls = ["https://www.ncbi.nlm.nih.gov/genome/genomes/971"] def start_requests(self): for url in self.start_urls: yield scrapy.Request(url, meta={"playwright": True}) def parse(self, response): # Adjust the selector to match the actual FTP link elements on the page for ftp_link in response.css("td a[href^='ftp://']::attr(href)"): yield { "ftp_url": ftp_link.get() }
Method 2: Use Selenium with Scrapy
If you're more familiar with Selenium, you can set up a custom downloader middleware to handle JS rendering.
Install dependencies
pip install scrapy selenium # Don't forget to download ChromeDriver matching your Chrome browser versionCreate a Selenium middleware
Add a new filemiddlewares.pyin your project with this code:from selenium import webdriver from selenium.webdriver.chrome.options import Options from scrapy.http import HtmlResponse class SeleniumMiddleware: def __init__(self): chrome_options = Options() chrome_options.add_argument("--headless=new") self.driver = webdriver.Chrome(options=chrome_options) def process_request(self, request, spider): self.driver.get(request.url) body = self.driver.page_source return HtmlResponse( self.driver.current_url, body=body, encoding='utf-8', request=request ) def __del__(self): self.driver.quit()Configure settings.py
Add the middleware to your project settings:DOWNLOADER_MIDDLEWARES = { "your_project_name.middlewares.SeleniumMiddleware": 800, }Write the spider
Your spider can now be a standard Scrapy spider—no extra meta needed, since the middleware handles rendering automatically:import scrapy class NCBIGenomeSpider(scrapy.Spider): name = "ncbi_genome" start_urls = ["https://www.ncbi.nlm.nih.gov/genome/genomes/971"] def parse(self, response): for ftp_link in response.css("td a[href^='ftp://']::attr(href)"): yield {"ftp_url": ftp_link.get()}
Method 3: Bypass JS by Calling the Backend API (Most Efficient)
This is the fastest method because you skip full browser rendering entirely—you just fetch raw data directly from NCBI's backend API.
Find the API endpoint
Open your browser's DevTools (F12), go to the Network tab, reload the page, and look for XHR/fetch requests returning JSON data. For this NCBI page, you'll spot an API endpoint that serves the genome data, including FTP links.Directly scrape the API
Once you have the API URL, request it directly in your spider and parse the JSON response:import scrapy import json class NCBIGenomeAPISpider(scrapy.Spider): name = "ncbi_genome_api" # Replace with the actual API endpoint you found in DevTools start_urls = ["https://api.ncbi.nlm.nih.gov/genome/..."] def parse(self, response): data = json.loads(response.body) # Adjust this path to match the structure of the API's JSON response for genome in data["genomes"]: if "ftp_url" in genome: yield {"ftp_url": genome["ftp_url"]}Pro tip: NCBI's APIs follow predictable patterns, so this method will be way faster than browser rendering—especially if you're scraping large datasets.
内容的提问来源于stack exchange,提问作者Sergii

