使用与浏览器相同请求头时,R的httr包POST请求超时问题
问题:使用httr请求Franklin Templeton ETF持仓数据超时的解决办法
背景
我需要用R的httr包获取Franklin Templeton网站上ClearBridge All-Cap Growth ESG ETF(代码CACG)的持仓数据,页面的“Portfolio Holdings”板块提供“Download Holdings (XLS)”下载按钮,浏览器可正常触发下载。
通过浏览器开发者工具抓包,确认数据由POST请求返回,我完全复制了浏览器的请求头和请求负载,编写了如下httr代码,但运行时始终出现超时错误。同一机器、同一IP下浏览器访问该网站及触发下载均无异常。
dta <- POST("https://www.franklintempleton.com/api/pds/price-and-performance?apikey=4ef35821-5244-41bc-a699-0192d002c3d1p&op=Holdings&id=14", body='{"operationName":"Holdings","variables":{"countrycode":"US","languagecode":"en_US","fundid":"91616"},"query":"query Holdings($fundid: String!, $countrycode: String!, $languagecode: String!) { Portfolio( fundid: $fundid countrycode: $countrycode languagecode: $languagecode ) { fundname producttype assetclass portfolio { topholdings { asofdate hldngname geocode sectorname brkdwnpct frequency calcbasislocal calctypelocal allocflag hasderivatives } dailyholdings { asofdate frequency secticker isinsecnbr cusipnbr secname quantityshrpar origcouponrate sctrname pctofnetassets mktvalue notionalmktvalue secexpdate assetclasscatg mktcurr contracts } fullholdings { asofdate asofdatestd frequency secticker isinsecnbr cusipnbr secname quantityshrpar origcouponrate sctrname pctofnetassets mktvalue notionalmktvalue secexpdate assetclasscatg mktcurr contracts finalmaturitydate investmentcategory } } } } "}', add_headers(.headers=c( 'Host'='www.franklintempleton.com', 'User-Agent'='Mozilla/5.0 (Windows NT 10.0; rv:121.0) Gecko/20100101 Firefox/121.0', 'Accept'='application/json', 'Accept-Language'='en-US,en;q=0.5', 'Accept-Encoding'='gzip, deflate, br', 'Content-Type'='application/json', 'Content-Length'=1401, 'Origin'='https://www.franklintempleton.com', 'DNT'=1, 'Connection'='keep-alive', 'Referer'='https://www.franklintempleton.com/investments/options/exchange-traded-funds/products/91616/SINGLCLASS/clearbridge-all-cap-growth-esg-etf/CACG', 'Sec-Fetch-Dest'='empty', 'Sec-Fetch-Mode'='cors', 'Sec-Fetch-Site'='same-origin', 'Sec-GPC'=1, 'Pragma'='no-cache', 'Cache-Control'='no-cache', 'TE'='trailers')))
疑问与需求
- 为何完全复制浏览器的请求头和负载后,httr请求仍会超时?服务器是否通过其他机制识别出非浏览器请求并拦截?
- 有没有办法通过httr或其他类似的R包解决该问题,避免使用RSelenium这类模拟浏览器的工具?
内容的提问来源于stack exchange,提问作者Kevin
相关产品推荐
相关产品推荐

