You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Flask REST API传递网址作为参数含/字符报错怎么解决

问题根源
  • Flask默认的字符串路由转换器不会匹配路径分隔符/,传入带/的完整URL时,路由无法命中,自然返回无效响应
  • 爬虫函数硬编码拼接https://前缀,若传入的参数本身带http:///https://,会生成无效的请求地址
  • 如果你后续把服务部署到Apache上,才需要调整AllowEncodedSlashes On配置允许编码后的斜杠传递,本地用Flask内置调试服务器的话,修改路由转换器即可解决问题
修复方案

1. 修改Flask路由配置

把路由参数改用path转换器,允许参数包含/字符:

import json
from flask import Flask, render_template

from imagescraper import image_scraper

app = Flask(__name__)

@app.route("/", methods = ['GET'])
def home():
    return render_template('index.html')

# 仅修改此行路由规则,使用path转换器
@app.route("/<path:site>", methods = ['GET'])
def get_image(site):
    return json.dumps(image_scraper(site))


if __name__ == '__main__':
    app.run(host='0.0.0.0', port=2434, debug=True)

2. 修复爬虫函数的前缀拼接逻辑

判断传入的站点地址是否已经带HTTP前缀,避免重复拼接:

import requests
from bs4 import BeautifulSoup


def image_scraper(site):
    """scrapes user inputed url for all images on a website and
    :param http url ex. https://www.cookinglight.com
    :return dictionary key:alt text; value: source link"""
    search = site.strip()
    search = search.replace(' ', '+')

    # 新增前缀判断逻辑,避免重复拼接HTTP头
    if not search.startswith(('http://', 'https://')):
        website = 'https://' + search
    else:
        website = search
    # 可按需添加超时参数,避免接口卡住
    response = requests.get(website, timeout=10)

    soup = BeautifulSoup(response.text, 'html.parser')
    img_tags = soup.find_all('img')
    # create dictionary to add image alt tag and source link
    images = {}
    for img in img_tags:
        try:
            name = img['alt']
            link = img['src']
            images[name] = link
        except:
            pass
    return images
额外优化建议
  • 可以新增异常捕获逻辑,处理目标站点无法访问、返回错误状态码的情况,避免接口直接抛出500错误
  • 相对路径的图片链接可以补充域名拼接,保证返回的图片地址可以直接访问

内容的提问来源于stack exchange,提问作者jaclynpgh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 23:27:03