使用HttpClient 4.5.8获取验证码图片返回202,浏览器却可正常获取
验证码图片下载问题:HttpClient返回202但浏览器正常获取
问题描述
我尝试用HttpClient 4.5.8下载验证码图片,接口地址是https://wbca.cde.org.cn/wbca/jcaptcha,但HttpClient返回状态码202,无法正常保存图片;而用浏览器访问却能正常显示。相关代码如下:
String url = "https://wbca.cde.org.cn/wbca/jcaptcha"; HttpClient httpClient = new DefaultHttpClient(); HttpGet getMethod = new HttpGet(url); try { HttpResponse response; response = httpClient.execute(getMethod, new BasicHttpContext()); System.out.println(response.getStatusLine()); HttpEntity entity = response.getEntity(); InputStream instream = entity.getContent(); OutputStream outstream = new FileOutputStream(new File("D:\\123.jpg")); int l = -1; byte[] tmp = new byte[2048]; while ((l = instream.read(tmp)) != -1) { outstream.write(tmp); } outstream.close(); } catch (ClientProtocolException e) { e.printStackTrace(); } catch (IOException e) { e.printStackTrace(); }finally { getMethod.releaseConnection(); }
请问为什么会出现这种差异,该怎么解决?
原因分析
出现202(Accepted)状态码且浏览器能正常访问的核心原因,是验证码接口依赖会话上下文和浏览器特征请求头:
- 会话Cookie缺失:验证码接口通常会和用户会话绑定,浏览器访问时会自动携带之前建立会话的Cookie(比如JSESSIONID),而你当前的HttpClient代码没有启用Cookie管理,每次请求都是全新的无状态请求,服务器无法识别会话,返回202表示请求已接收但暂不处理。
- 请求头特征不符:服务器可能会校验请求的
User-Agent、Referer等头部,判断请求是否来自合法浏览器。你的HttpClient请求没有设置这些头部,被服务器判定为非浏览器请求,返回202而非正常的图片响应。
解决方法
针对上述问题,你需要修改代码,添加Cookie管理并模拟浏览器请求头:
1. 启用Cookie存储管理
使用BasicCookieStore和HttpClientBuilder来维护会话Cookie,确保请求携带服务器分配的会话标识。
2. 添加浏览器特征请求头
设置User-Agent、Referer等头部,让请求更接近浏览器的行为。
修正后的代码如下:
import org.apache.http.client.CookieStore; import org.apache.http.client.config.CookieSpecs; import org.apache.http.client.config.RequestConfig; import org.apache.http.client.methods.CloseableHttpResponse; import org.apache.http.client.methods.HttpGet; import org.apache.http.client.protocol.HttpClientContext; import org.apache.http.impl.client.BasicCookieStore; import org.apache.http.impl.client.CloseableHttpClient; import org.apache.http.impl.client.HttpClients; import org.apache.http.util.EntityUtils; import java.io.FileOutputStream; import java.io.IOException; public class CaptchaDownloader { public static void main(String[] args) { String url = "https://wbca.cde.org.cn/wbca/jcaptcha"; // 初始化CookieStore,用于保存会话Cookie CookieStore cookieStore = new BasicCookieStore(); // 构建HttpClient,启用Cookie管理 CloseableHttpClient httpClient = HttpClients.custom() .setDefaultCookieStore(cookieStore) .build(); // 设置请求配置,兼容常见Cookie规范 RequestConfig requestConfig = RequestConfig.custom() .setCookieSpec(CookieSpecs.DEFAULT) .build(); HttpGet getMethod = new HttpGet(url); // 添加浏览器请求头 getMethod.setHeader("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"); getMethod.setHeader("Referer", "https://wbca.cde.org.cn/"); // 填写验证码所在页面的Referer getMethod.setConfig(requestConfig); HttpClientContext context = HttpClientContext.create(); context.setCookieStore(cookieStore); try (CloseableHttpResponse response = httpClient.execute(getMethod, context)) { System.out.println(response.getStatusLine()); if (response.getStatusLine().getStatusCode() == 200) { byte[] imageBytes = EntityUtils.toByteArray(response.getEntity()); try (FileOutputStream outstream = new FileOutputStream("D:\\123.jpg")) { outstream.write(imageBytes); System.out.println("验证码图片保存成功"); } } else { System.out.println("请求失败,状态码:" + response.getStatusLine().getStatusCode()); } } catch (IOException e) { e.printStackTrace(); } finally { try { httpClient.close(); } catch (IOException e) { e.printStackTrace(); } } } }
额外说明
- 注意
Referer头的值要填写验证码图片所在页面的URL,确保和浏览器访问时的Referer一致; - 如果还是返回202,可以尝试先请求验证码所在的页面,获取会话Cookie后再请求验证码接口,确保会话上下文完全匹配浏览器的访问流程;
DefaultHttpClient在HttpClient 4.3+已经被标记为过时,推荐使用HttpClients.custom()构建CloseableHttpClient,更符合最新的API规范。
内容的提问来源于stack exchange,提问作者Greatvia
相关产品推荐
相关产品推荐

