You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Java/Python爬取指定链接时页面存在却返回404状态码求助

Hey there! Let's tackle this 404 issue you're facing even though the page clearly exists—super frustrating, right? I've run into this plenty of times when scraping, so here are the most common fixes and things to check, split out for both Java and Python:

Common Reasons & Fixes for False 404s When Scraping

1. Missing or Incorrect Request Headers

Most modern websites block scrapers by checking for browser-like request headers. If your code sends a bare-bones request without a proper User-Agent, the server might reject it with a 404 (even if the page exists).

Java Fix (Using Apache HttpClient)

import org.apache.http.client.methods.HttpGet;
import org.apache.http.impl.client.CloseableHttpClient;
import org.apache.http.impl.client.HttpClients;

public class Scraper {
    public static void main(String[] args) throws Exception {
        CloseableHttpClient client = HttpClients.createDefault();
        HttpGet request = new HttpGet("YOUR_TARGET_URL");
        
        // Mimic a real Chrome browser's headers
        request.addHeader("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");
        request.addHeader("Accept-Language", "en-US,en;q=0.9");
        request.addHeader("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8");
        
        var response = client.execute(request);
        System.out.println("Status Code: " + response.getStatusLine().getStatusCode());
        client.close();
    }
}

Python Fix (Using Requests Library)

import requests

url = "YOUR_TARGET_URL"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"
}

response = requests.get(url, headers=headers)
print(f"Status Code: {response.status_code}")

2. Redirects Are Not Being Followed

Some pages use redirects (e.g., from http:// to https://, or a temporary redirect to the actual page). If your scraper isn't configured to follow these, you might get a 404 instead of the redirected content.

Java Note

Apache HttpClient follows redirects by default, but if you've customized your client, double-check this setting:

CloseableHttpClient client = HttpClients.custom()
    .setRedirectStrategy(new DefaultRedirectStrategy())
    .build();

Python Note

Requests also follows redirects by default, but you can explicitly enable it to be sure:

response = requests.get(url, headers=headers, allow_redirects=True)

3. IP Blocking or Rate Limiting

If you've made too many requests too quickly, the site might block your IP with a 404 (instead of a more obvious 429). Try adding delays between requests or using a proxy.

Java Proxy Example

import org.apache.http.HttpHost;

HttpHost proxy = new HttpHost("YOUR_PROXY_IP", YOUR_PROXY_PORT);
CloseableHttpClient client = HttpClients.custom()
    .setProxy(proxy)
    .build();

Python Delay Example

import time

# Add a 2-second delay before each request
time.sleep(2)
response = requests.get(url, headers=headers)

4. URL Encoding Issues

If your target URL has special characters (spaces, &, =, or non-ASCII characters), failing to properly encode it can lead to a 404.

Java Encoding Fix

import java.net.URLEncoder;
import java.nio.charset.StandardCharsets;

String rawUrl = "YOUR_RAW_URL_WITH_SPECIAL_CHARS";
String encodedUrl = URLEncoder.encode(rawUrl, StandardCharsets.UTF_8);
HttpGet request = new HttpGet(encodedUrl);

Python Encoding Fix

from urllib.parse import quote

raw_url = "YOUR_RAW_URL_WITH_SPECIAL_CHARS"
# Keep common URL characters like : / ? & unencoded
encoded_url = quote(raw_url, safe=':/?&=')
response = requests.get(encoded_url, headers=headers)

Some sites require you to have a valid session cookie (e.g., after visiting the homepage) before accessing specific pages. Without these cookies, the server might return a 404.

Java Session Example

CloseableHttpClient client = HttpClients.createDefault();

// First visit the homepage to get session cookies
HttpGet homeRequest = new HttpGet("YOUR_TARGET_SITE_HOME_URL");
client.execute(homeRequest);

// Now request the target page - cookies are automatically stored
HttpGet targetRequest = new HttpGet("YOUR_TARGET_URL");
targetRequest.addHeader("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");
var response = client.execute(targetRequest);

Python Session Example

session = requests.Session()

# First visit the homepage to set session cookies
session.get("YOUR_TARGET_SITE_HOME_URL", headers=headers)

# Now request the target page using the same session
response = session.get("YOUR_TARGET_URL", headers=headers)

Quick Pro Tip

Use your browser's dev tools (press F12, go to the Network tab) to capture a real request to the target page. Copy all the headers from that request into your scraper code—this is the fastest way to match what a browser sends, and it usually fixes the 404 issue immediately.

内容的提问来源于stack exchange,提问作者Partiban

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:12:34