遍历列表处理302重定向时Rescue块仅首次执行问题
问题与解决方案:遍历ID时Rescue块仅首次触发及S3授权错误处理
问题现象
- 遍历ID列表调用
download_doc方法时,Rescue块仅在第一次迭代时触发,后续迭代无响应 - 原GET请求重定向到S3时触发错误:
Only one auth mechanism allowed; only the X-Amz-Algorithm query parameter, Signature query string parameter or the Authorization header should be specified,因此禁用了自动重定向,尝试用wget获取返回的Location,但问题未解决
原代码
def download_doc(gem_id) url = @base_url + ATTACHMENT_ENDPOINT + gem_id + "\/attachment" headers = { 'Authorization' => @token } attempts = 0 begin response = RestClient::Request.execute( method: :get, url: url, headers: { Authorization: "Bearer #{@token}" }, max_redirects: 0 ) rescue RestClient::Found => found puts "FOUND #{gem_id}" exec "wget \\\"#{found.response.headers[:location]}\\\" -O \\\"#{gem_id}.pdf\\\" -q --show-progress" rescue => ex puts ex.response end end ids = File.read(@input_file).split(",") arra.each do |id| download_doc(id) end
问题分析
- Rescue块仅首次触发的核心原因:代码中使用了
exec命令,该命令会直接替换当前Ruby进程,执行完wget后进程立即退出,后续迭代根本没有机会执行。 - S3授权错误原因:RestClient默认会在重定向时携带原请求的Authorization头,但S3的预签名URL已经包含了签名参数,同时携带Authorization头会触发冲突,导致报错。
解决方案
1. 修复Rescue块仅触发一次的问题
将exec替换为system或反引号(`),这两个方法会在子进程中执行命令,不会终止当前Ruby进程,确保后续迭代能正常运行:
# 替换exec为system system "wget \"#{found.response.headers[:location]}\" -O \"#{gem_id}.pdf\" -q --show-progress" # 或者使用反引号 `wget \"#{found.response.headers[:location]}\" -O \"#{gem_id}.pdf\" -q --show-progress`
2. 绕过S3授权错误的优化方案(无需依赖wget)
直接用Ruby处理重定向,请求S3地址时移除Authorization头,避免冲突:
def download_doc(gem_id) url = @base_url + ATTACHMENT_ENDPOINT + gem_id + "/attachment" begin # 第一步:发起请求获取重定向地址,禁用自动重定向 response = RestClient::Request.execute( method: :get, url: url, headers: { Authorization: "Bearer #{@token}" }, max_redirects: 0 ) rescue RestClient::Found => found puts "FOUND #{gem_id}" s3_url = found.response.headers[:location] # 第二步:请求S3地址,不携带Authorization头 s3_response = RestClient.get(s3_url) # 保存文件到本地 File.write("#{gem_id}.pdf", s3_response.body) rescue => ex puts ex.response || ex.message end end # 修正遍历变量名错误,同时清理ID中的空白字符 ids = File.read(@input_file).split(",").map(&:strip) ids.each do |id| download_doc(id) end
额外修正点
- 原代码遍历的变量名写错:
arra.each应为ids.each - 读取的ID可能包含空格,添加
strip方法清理空白字符,避免无效请求
内容的提问来源于stack exchange,提问作者boberczus
相关产品推荐
相关产品推荐

