You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

跟随Python网页爬取教程时遇MissingSchema错误求助

问题原因与解决方法

错误原因

你抓取到的图片src属性是相对路径(例如media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg),而requests.get()要求传入的必须是包含http:///https://协议的完整URL,因此触发MissingSchema错误。

解决步骤

  1. 拼接完整URL:用Python标准库urllib.parse.urljoin,将网站基础URL与图片相对路径拼接成合法的完整请求地址。
  2. 确保目录存在:提前创建images文件夹,避免写入文件时出现“找不到目录”的额外错误。

修改后的完整代码

import requests
import os
from bs4 import BeautifulSoup
from urllib.parse import urljoin

URL = "http://books.toscrape.com/"
# 统一复用UA头,降低被网站拦截的概率
headers = {"User-Agent":"Mozilla/5.0"}
getURL = requests.get(URL, headers=headers)
print(getURL.status_code)

soup = BeautifulSoup(getURL.text, 'html.parser')
images = soup.find_all('img')

imageSources = [image.get("src") for image in images]
print(imageSources)

# 自动创建images目录(不存在时)
if not os.path.exists('images'):
    os.makedirs('images')

# 遍历下载图片
for image_src in imageSources:
    full_image_url = urljoin(URL, image_src)
    webs = requests.get(full_image_url, headers=headers)
    # 提取纯文件名,避免路径嵌套导致的命名问题
    filename = os.path.basename(image_src)
    with open(f"images/{filename}", "wb") as f:
        f.write(webs.content)

额外提示

  • 统一复用headers参数,避免被目标网站识别为爬虫拦截。
  • 使用with open语法操作文件,能自动处理文件关闭,避免资源泄漏。

内容的提问来源于stack exchange,提问作者slow_learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 02:30:46