使用Java/Python爬取指定链接时页面存在却返回404状态码求助
Hey there! Let's tackle this 404 issue you're facing even though the page clearly exists—super frustrating, right? I've run into this plenty of times when scraping, so here are the most common fixes and things to check, split out for both Java and Python:
1. Missing or Incorrect Request Headers
Most modern websites block scrapers by checking for browser-like request headers. If your code sends a bare-bones request without a proper User-Agent, the server might reject it with a 404 (even if the page exists).
Java Fix (Using Apache HttpClient)
import org.apache.http.client.methods.HttpGet; import org.apache.http.impl.client.CloseableHttpClient; import org.apache.http.impl.client.HttpClients; public class Scraper { public static void main(String[] args) throws Exception { CloseableHttpClient client = HttpClients.createDefault(); HttpGet request = new HttpGet("YOUR_TARGET_URL"); // Mimic a real Chrome browser's headers request.addHeader("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"); request.addHeader("Accept-Language", "en-US,en;q=0.9"); request.addHeader("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"); var response = client.execute(request); System.out.println("Status Code: " + response.getStatusLine().getStatusCode()); client.close(); } }
Python Fix (Using Requests Library)
import requests url = "YOUR_TARGET_URL" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8" } response = requests.get(url, headers=headers) print(f"Status Code: {response.status_code}")
2. Redirects Are Not Being Followed
Some pages use redirects (e.g., from http:// to https://, or a temporary redirect to the actual page). If your scraper isn't configured to follow these, you might get a 404 instead of the redirected content.
Java Note
Apache HttpClient follows redirects by default, but if you've customized your client, double-check this setting:
CloseableHttpClient client = HttpClients.custom() .setRedirectStrategy(new DefaultRedirectStrategy()) .build();
Python Note
Requests also follows redirects by default, but you can explicitly enable it to be sure:
response = requests.get(url, headers=headers, allow_redirects=True)
3. IP Blocking or Rate Limiting
If you've made too many requests too quickly, the site might block your IP with a 404 (instead of a more obvious 429). Try adding delays between requests or using a proxy.
Java Proxy Example
import org.apache.http.HttpHost; HttpHost proxy = new HttpHost("YOUR_PROXY_IP", YOUR_PROXY_PORT); CloseableHttpClient client = HttpClients.custom() .setProxy(proxy) .build();
Python Delay Example
import time # Add a 2-second delay before each request time.sleep(2) response = requests.get(url, headers=headers)
4. URL Encoding Issues
If your target URL has special characters (spaces, &, =, or non-ASCII characters), failing to properly encode it can lead to a 404.
Java Encoding Fix
import java.net.URLEncoder; import java.nio.charset.StandardCharsets; String rawUrl = "YOUR_RAW_URL_WITH_SPECIAL_CHARS"; String encodedUrl = URLEncoder.encode(rawUrl, StandardCharsets.UTF_8); HttpGet request = new HttpGet(encodedUrl);
Python Encoding Fix
from urllib.parse import quote raw_url = "YOUR_RAW_URL_WITH_SPECIAL_CHARS" # Keep common URL characters like : / ? & unencoded encoded_url = quote(raw_url, safe=':/?&=') response = requests.get(encoded_url, headers=headers)
5. Session/Cookie Requirements
Some sites require you to have a valid session cookie (e.g., after visiting the homepage) before accessing specific pages. Without these cookies, the server might return a 404.
Java Session Example
CloseableHttpClient client = HttpClients.createDefault(); // First visit the homepage to get session cookies HttpGet homeRequest = new HttpGet("YOUR_TARGET_SITE_HOME_URL"); client.execute(homeRequest); // Now request the target page - cookies are automatically stored HttpGet targetRequest = new HttpGet("YOUR_TARGET_URL"); targetRequest.addHeader("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"); var response = client.execute(targetRequest);
Python Session Example
session = requests.Session() # First visit the homepage to set session cookies session.get("YOUR_TARGET_SITE_HOME_URL", headers=headers) # Now request the target page using the same session response = session.get("YOUR_TARGET_URL", headers=headers)
Quick Pro Tip
Use your browser's dev tools (press F12, go to the Network tab) to capture a real request to the target page. Copy all the headers from that request into your scraper code—this is the fastest way to match what a browser sends, and it usually fixes the 404 issue immediately.
内容的提问来源于stack exchange,提问作者Partiban

