You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取URL保存至HerHoops目录报错:FileNotFoundError问题排查

问题排查:保存爬取内容时的FileNotFoundError错误

我已成功从herhoopstats.com爬取链接,想要将其保存至已创建的本地文件夹“HerHoops”中,以供后续解析。此前我曾成功完成此类操作,但本次目标网站的链接需要进行额外清理。

需求说明:保留链接中“box_score”之后的所有内容,使保存的文件名包含日期和对阵球队信息,且以“w+”模式写入。

原代码

import requests
from bs4 import BeautifulSoup
import time

url = f"https://herhoopstats.com/stats/wnba/schedule_date/2004/6/1/"
data = requests.get(url)
soup = BeautifulSoup(data.text)
matchup_table = soup.find_all("div", {"class": "schedule"})[0]

links = matchup_table.find_all('a')
links = [l.get("href") for l in links]
links = [l for l in links if '/box_score/' in l]

box_scores_urls = [f"https://herhoopstats.com{l}" for l in links]

for box_scores_url in box_scores_urls:
      data = requests.get(box_scores_url)
      # within loop opening up page and saving to folder in write mode
      with open("HerHoops/{}".format(box_scores_url[46:]), "w+") as f:
         # write to the files
         f.write(data.text) 
      time.sleep(3)

报错信息

FileNotFoundError: [Errno 2] No such file or directory: 'HerHoops/2004/06/01/new-york-liberty-vs-charlotte-sting/'

问题原因

  1. 嵌套文件夹未创建:报错路径中的2004/06/01是嵌套子文件夹,open()函数无法自动创建这些不存在的层级目录。
  2. 路径末尾斜杠导致识别错误:链接末尾的斜杠会被当作文件夹路径,而非文件名,系统找不到对应的文件。
  3. 固定索引提取路径风险:用box_scores_url[46:]提取路径依赖固定位置,一旦链接结构变化就会出错。

解决方案

修改代码,添加文件夹创建逻辑,修正路径处理方式:

import requests
from bs4 import BeautifulSoup
import time
import os

url = f"https://herhoopstats.com/stats/wnba/schedule_date/2004/6/1/"
data = requests.get(url)
soup = BeautifulSoup(data.text)
matchup_table = soup.find_all("div", {"class": "schedule"})[0]

links = matchup_table.find_all('a')
links = [l.get("href") for l in links]
links = [l for l in links if '/box_score/' in l]

box_scores_urls = [f"https://herhoopstats.com{l}" for l in links]

for box_scores_url in box_scores_urls:
    data = requests.get(box_scores_url)
    # 提取box_score之后的路径部分,去掉末尾斜杠
    path_part = box_scores_url.split('/box_score/')[-1].rstrip('/')
    # 拼接完整保存路径,添加.html后缀明确是网页文件
    save_path = os.path.join("HerHoops", path_part) + ".html"
    # 创建所有必要的父文件夹
    os.makedirs(os.path.dirname(save_path), exist_ok=True)
    # 写入文件
    with open(save_path, "w+", encoding="utf-8") as f:
        f.write(data.text)
    time.sleep(3)

关键修改点

  • 用split('/box_score/')提取路径:避免固定索引的脆弱性,准确获取目标部分。
  • rstrip('/')去掉末尾斜杠:确保路径指向文件而非文件夹。
  • os.makedirs(..., exist_ok=True):自动创建所有缺失的父文件夹,exist_ok=True避免文件夹已存在时报错。
  • 添加.html后缀:明确保存的是网页文件,方便后续解析。
  • 指定encoding="utf-8":避免写入时出现乱码问题。

内容的提问来源于stack exchange,提问作者kc_balr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 18:15:42