Python Flask REST API传递网址作为参数含/字符报错怎么解决
问题根源
- Flask默认的字符串路由转换器不会匹配路径分隔符
/,传入带/的完整URL时,路由无法命中,自然返回无效响应 - 爬虫函数硬编码拼接
https://前缀,若传入的参数本身带http:///https://,会生成无效的请求地址 - 如果你后续把服务部署到Apache上,才需要调整
AllowEncodedSlashes On配置允许编码后的斜杠传递,本地用Flask内置调试服务器的话,修改路由转换器即可解决问题
修复方案
1. 修改Flask路由配置
把路由参数改用path转换器,允许参数包含/字符:
import json from flask import Flask, render_template from imagescraper import image_scraper app = Flask(__name__) @app.route("/", methods = ['GET']) def home(): return render_template('index.html') # 仅修改此行路由规则,使用path转换器 @app.route("/<path:site>", methods = ['GET']) def get_image(site): return json.dumps(image_scraper(site)) if __name__ == '__main__': app.run(host='0.0.0.0', port=2434, debug=True)
2. 修复爬虫函数的前缀拼接逻辑
判断传入的站点地址是否已经带HTTP前缀,避免重复拼接:
import requests from bs4 import BeautifulSoup def image_scraper(site): """scrapes user inputed url for all images on a website and :param http url ex. https://www.cookinglight.com :return dictionary key:alt text; value: source link""" search = site.strip() search = search.replace(' ', '+') # 新增前缀判断逻辑,避免重复拼接HTTP头 if not search.startswith(('http://', 'https://')): website = 'https://' + search else: website = search # 可按需添加超时参数,避免接口卡住 response = requests.get(website, timeout=10) soup = BeautifulSoup(response.text, 'html.parser') img_tags = soup.find_all('img') # create dictionary to add image alt tag and source link images = {} for img in img_tags: try: name = img['alt'] link = img['src'] images[name] = link except: pass return images
额外优化建议
- 可以新增异常捕获逻辑,处理目标站点无法访问、返回错误状态码的情况,避免接口直接抛出500错误
- 相对路径的图片链接可以补充域名拼接,保证返回的图片地址可以直接访问
内容的提问来源于stack exchange,提问作者jaclynpgh
相关产品推荐
相关产品推荐

