使用Java HttpURLConnection编写爬虫时遇406错误求助
Fixing the 406 Error with Java HttpURLConnection Web Crawler
Hey there, let’s break down that 406 error you’re hitting. A 406 Not Acceptable response means the server can’t send back content in a format that matches what your request told it you’d accept. Most likely, your HttpURLConnection request isn’t sending the right Accept header—so the server has no clue what kind of content to return to you.
Here’s how to fix it step by step:
1. Add Key Request Headers
Web servers often validate requests using two critical headers:
Accept: Tells the server what content types your client can handle (like HTML, which is standard for web pages)User-Agent: Makes your request look like it’s coming from a real browser (many sites block requests without this to prevent bots)
2. Updated Code Snippet
Modify your sendGet method to include these headers. Here’s a working example:
public static String sendGet(String url) throws IOException { URL obj = new URL(url); HttpURLConnection connection = (HttpURLConnection) obj.openConnection(); // Set HTTP GET method connection.setRequestMethod("GET"); // Add headers to resolve 406 and mimic a browser connection.setRequestProperty("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8"); connection.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"); int responseCode = connection.getResponseCode(); System.out.println("Response Code: " + responseCode); if (responseCode == HttpURLConnection.HTTP_OK) { BufferedReader in = new BufferedReader(new InputStreamReader(connection.getInputStream())); String inputLine; StringBuilder response = new StringBuilder(); while ((inputLine = in.readLine()) != null) { response.append(inputLine); } in.close(); return response.toString(); } else { throw new IOException("Server returned HTTP response code: " + responseCode + " for URL: " + url); } }
Why This Works
- The
Acceptheader explicitly tells the server you accept HTML and other common web formats. Theq=0.9andq=0.8values are quality scores, letting the server know your preference if multiple formats are available. - The
User-Agentheader mimics a Chrome browser, which avoids triggering many sites’ anti-bot filters that block generic crawler requests.
Extra Tips
- If you still run into issues, use your browser’s DevTools (Network tab) to check the headers sent during a normal page load, then copy those exact headers into your code.
- Some sites may require additional headers like
RefererorCookie, but fixingAcceptandUser-Agentshould resolve the 406 error first.
内容的提问来源于stack exchange,提问作者Anna Noukou
相关产品推荐
相关产品推荐

