libcurl调用Pixabay接口后发起后续图片下载请求的实现方法
Pixabay图片下载问题排查与解决
实现目标
- 在Pixabay平台搜索图片(本次以搜索黄色花朵为例)
- 提交查询请求后,获取包含图片详情信息的JSON数组
- 解析返回数据并存储数组内容
- 基于数组数据中提供的URL,发起后续curl请求获取/下载对应图片
当前开发进度
已完成前3步开发,第3步通过将接口返回数据存储到本地文件后再解析实现,目前卡在第4步图片下载环节。
接口返回单条数据示例
"hits":[ { "id":2295434, "pageURL":"https://pixabay.com/photos/spring-bird-bird-tit-spring-blue-2295434/", "type":"photo", "tags":"spring bird, bird, tit", "previewURL":"https://cdn.pixabay.com/photo/2017/05/08/13/15/spring-bird-2295434_150.jpg", "previewWidth":150, "previewHeight":99, "webformatURL":"https://pixabay.com/get/gc323739b5570ab1afac9cff34f0ed431beffcf004e0660fcaff96a7b1780933e31c982b26cdb3234c4a757a9e7e8b824bda6059340ead3c6b755f1265f7ace52_640.jpg", "webformatWidth":640, "webformatHeight":426, "largeImageURL":"https://pixabay.com/get/g58a377c4cffe13bffd15cb7b455ecec08329352a443689296c5f15565b3a7f11bcd5ec6e4d2945bb3d32b2a63ee33d6c9a9925119944d128cab4bca663f87620_1280.jpg", "imageWidth":5363, "imageHeight":3575, "imageSize":2938651, "views":488125, "downloads":261832, "collections":1911, "likes":1846, "comments":221, "user_id":334088, "user":"JillWellington", "userImageURL":"https://cdn.pixabay.com/user/2018/06/27/01-23-02-27_250x250.jpg" } ]
现有代码
main.cc
#include "baseurihandler/base_uri_handler_pixabay.h" #include "json_parser.h" #include "download_image.h" #include <stdio.h> #include <iostream> #include <curl/curl.h> #include <iostream> #include <chrono> #include <string> #include <thread> #include <vector> void JsonParser(); int main() { CURL *curl; CURLcode res; std::vector<std::pair<std::string, std::string>> query_param; query_param.push_back(std::make_pair("q", "yellow+flowers")); query_param.push_back(std::make_pair("image_type", "photo")); std::string s; baseuri::PixabayURIHandler pixabay_uri_handler; pixabay_uri_handler.SetQueryParameters(query_param); //auto str = pixabay_uri_handler.GetURI(); auto str = std::string("https://pixabay.com/api/?key=xxxxxxx-xxxxxxxxxxx&q=yellow+flowers&image_type=photo"); std::cout << str << std::endl; curl = curl_easy_init(); if(curl) { curl_easy_setopt(curl, CURLOPT_URL, str.c_str()); /* example.com is redirected, so we tell libcurl to follow redirection */ curl_easy_setopt(curl, CURLOPT_FOLLOWLOCATION, 1L); curl_easy_setopt(curl, CURLOPT_WRITEFUNCTION, baseuri::PixabayURIHandler::CurlWrite_CallbackFunc_StdString); curl_easy_setopt(curl, CURLOPT_WRITEDATA, &pixabay_uri_handler); /* Perform the request, res will get the return code */ res = curl_easy_perform(curl); /* Check for errors */ if (res != CURLE_OK) fprintf(stderr, "curl_easy_perform() failed: %s\n", curl_easy_strerror(res)); /* always cleanup */ curl_easy_cleanup(curl); } //std::cout << pixabay_uri_handler.GetResultJsonString() << std::endl; std::cout << "calling jsonparser.. " << std::endl; JsonParser(); //This will return vec of URLs - TBD char *jpg_test = "https://pixabay.com/get/g58a377c4cffe13bffd15cb7b455ecec08329352a443689296c5f15565b3a7f11bcd5ec6e4d2945bb3d32b2a63ee33d6c9a9925119944d128cab4bca663f87620_1280.jpg"; if (!download_jpeg(jpg_test)) { printf("!! Failed to download file!\n" ); return -1; } return 0; }
download_jpeg函数实现
#include <stdio.h> #include <curl/curl.h> size_t callbackfunction(void *ptr, size_t size, size_t nmemb, void* userdata) { FILE* stream = (FILE*)userdata; if (!stream) { printf("!!! No stream\n"); return 0; } size_t written = fwrite((FILE*)ptr, size, nmemb, stream); return written; } bool download_jpeg(char* url) { FILE* fp = fopen("out.jpg", "wb"); if (!fp) { printf("!!! Failed to create file on the disk\n"); return false; } CURL* curlCtx = curl_easy_init(); curl_easy_setopt(curlCtx, CURLOPT_URL, url); curl_easy_setopt(curlCtx, CURLOPT_WRITEDATA, fp); curl_easy_setopt(curlCtx, CURLOPT_WRITEFUNCTION, callbackfunction); curl_easy_setopt(curlCtx, CURLOPT_FOLLOWLOCATION, 1); CURLcode rc = curl_easy_perform(curlCtx); if (rc) { printf("!!! Failed to download: %s\n", url); return false; } long res_code = 0; curl_easy_getinfo(curlCtx, CURLINFO_RESPONSE_CODE, &res_code); if (!((res_code == 200 || res_code == 201) && rc != CURLE_ABORTED_BY_CALLBACK)) { printf("!!! Response code: %d\n", res_code); return false; } curl_easy_cleanup(curlCtx); fclose(fp); return true; }
遇到的问题
尝试直接将largeImageURL传入download_jpeg()函数下载失败,在largeImageURL后拼接API密钥作为查询参数返回400响应码,未在官方文档找到对应解决方案。
解决方案
不要给图片直链拼接API密钥
API密钥仅在调用/api/搜索接口时需要校验,接口返回的largeImageURL、webformatURL、previewURL都是自带签名的临时有效直链,直接发起请求即可,额外拼接参数会破坏签名有效性,直接触发400错误。补全curl请求的User-Agent头
下载失败的核心原因是Pixabay的CDN会拦截默认无标识的libcurl请求。当前curl配置没有设置User-Agent,请求会被服务端识别为异常爬虫直接拦截。
在download_jpeg函数的curl配置段添加如下配置即可:
curl_easy_setopt(curlCtx, CURLOPT_USERAGENT, "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36");
- 其他注意事项
- 代码里写死的测试URL是示例中的鸟类图片链接,这类带签名的直链存在有效期,如果是很早之前拿到的URL,签名过期也会导致下载失败,测试时请使用本次搜索接口实时返回的URL。
- 批量下载时不要将所有文件都命名为
out.jpg,会出现文件覆盖问题,可以使用接口返回的图片id作为文件名,避免重名。
内容的提问来源于stack exchange,提问作者Vineet Dwivedi
相关产品推荐
相关产品推荐

