You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用URL类处理NCBI基因网站分页并解析页面时遇到技术问题

Troubleshooting Your NCBI Gene Scraping & Pagination Code

Hey there! I’ve tackled similar NCBI web scraping tasks before, so let’s work through the issues you’re likely hitting with your current code and get you back on track.

Common Pitfalls in Your Current Approach

Your code snippet has a few gaps that are probably causing problems:

  • NCBI’s anti-scraping measures: Plain URLConnection sends minimal request headers, which NCBI’s servers often flag as non-browser traffic and block.
  • No pagination handling: You’re only fetching a single page, but you mentioned needing to handle page turns—your code doesn’t account for NCBI’s pagination parameters.
  • Brittle HTML parsing: Reading raw HTML line-by-line is error-prone; small changes to NCBI’s page structure will break your data extraction.
  • Incomplete resource management: You don’t have proper catch/finally blocks to close connections or readers, leading to resource leaks.

Fixes & Improved Implementation

Let’s rewrite and enhance your code to address these issues:

1. Bypass Anti-Scraping with Proper Request Headers

First, update your connection to mimic a real browser—this will avoid getting blocked by NCBI:

import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.net.HttpURLConnection;
import java.net.URL;

public class NCBIScraper {
    public static void main(String[] args) {
        HttpURLConnection conn = null;
        BufferedReader br = null;
        
        try {
            String fullUrl = Vars.searchLink + Vars.searchWord;
            URL url = new URL(fullUrl);
            conn = (HttpURLConnection) url.openConnection();
            
            // Mimic a Chrome browser request
            conn.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36");
            conn.setRequestProperty("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8");
            
            br = new BufferedReader(new InputStreamReader(conn.getInputStream()));
            StringBuilder pageContent = new StringBuilder();
            String inputLine;
            
            while ((inputLine = br.readLine()) != null) {
                pageContent.append(inputLine);
            }
            
            // Pass content to your parsing logic here
            parseNCBITable(pageContent.toString());
            
        } catch (IOException e) {
            // Handle specific errors (e.g., connection refused, 403 Forbidden)
            System.err.println("Error fetching page: " + e.getMessage());
            e.printStackTrace();
        } finally {
            // Clean up resources properly
            if (br != null) {
                try {
                    br.close();
                } catch (IOException e) {
                    e.printStackTrace();
                }
            }
            if (conn != null) {
                conn.disconnect();
            }
        }
    }
    
    private static void parseNCBITable(String htmlContent) {
        // Replace with your table parsing logic (or use Jsoup as shown below)
    }
}

2. Simplify HTML Parsing with Jsoup

Manual line-by-line parsing is unreliable—use the Jsoup library to easily extract table data. First, add the dependency (Maven example):

<dependency>
    <groupId>org.jsoup</groupId>
    <artifactId>jsoup</artifactId>
    <version>1.17.2</version>
</dependency>

Then, update your code to handle pagination and table extraction:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

import java.io.IOException;

public class NCBIPaginationScraper {
    private static final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36";

    public static void main(String[] args) throws IOException {
        String baseSearchUrl = Vars.searchLink + Vars.searchWord;
        int totalPages = getTotalPages(baseSearchUrl);
        
        // Iterate through all pages
        for (int page = 1; page <= totalPages; page++) {
            String pageUrl = baseSearchUrl + "&page=" + page;
            Document doc = Jsoup.connect(pageUrl)
                                .userAgent(USER_AGENT)
                                .timeout(10000) // Add timeout to avoid hanging
                                .get();
            
            // Extract table rows (update the selector to match NCBI's actual table class)
            Elements tableRows = doc.select("table.usa-table tr");
            
            for (Element row : tableRows) {
                Elements columns = row.select("td");
                if (columns.size() >= 2) { // Ensure we have enough columns
                    String geneId = columns.get(0).text().trim();
                    String geneName = columns.get(1).text().trim();
                    
                    // Process your extracted data here
                    System.out.println("Gene ID: " + geneId + ", Name: " + geneName);
                }
            }
            
            // Add a delay to avoid overwhelming NCBI's servers (critical to avoid bans)
            try {
                Thread.sleep(1000); // 1 second delay between requests
            } catch (InterruptedException e) {
                Thread.currentThread().interrupt();
            }
        }
    }

    // Helper method to get total number of pages from NCBI's pagination bar
    private static int getTotalPages(String baseUrl) throws IOException {
        Document doc = Jsoup.connect(baseUrl)
                            .userAgent(USER_AGENT)
                            .get();
        
        // Update this selector to match NCBI's pagination text (e.g., "Page 1 of 10")
        String paginationText = doc.select("div.pagination span").text();
        
        if (paginationText.contains("of")) {
            return Integer.parseInt(paginationText.split("of")[1].trim());
        }
        return 1; // Default to 1 page if pagination isn't found
    }
}

Key Notes to Avoid Issues

  • Follow NCBI’s Terms of Service: Don’t send too many requests too quickly—add delays as shown. Consider using NCBI’s official Entrez API instead of scraping; it’s more stable and compliant.
  • Update Selectors: NCBI may change its page structure over time, so periodically verify that your CSS selectors (like table.usa-table tr) still target the correct elements.
  • Handle Edge Cases: Add checks for empty pages, 403/404 errors, and unexpected HTML structures to make your code robust.

内容的提问来源于stack exchange,提问作者Orel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:53:58