爬取URL保存至HerHoops目录报错:FileNotFoundError问题排查
问题排查:保存爬取内容时的FileNotFoundError错误
我已成功从herhoopstats.com爬取链接,想要将其保存至已创建的本地文件夹“HerHoops”中,以供后续解析。此前我曾成功完成此类操作,但本次目标网站的链接需要进行额外清理。
需求说明:保留链接中“box_score”之后的所有内容,使保存的文件名包含日期和对阵球队信息,且以“w+”模式写入。
原代码
import requests from bs4 import BeautifulSoup import time url = f"https://herhoopstats.com/stats/wnba/schedule_date/2004/6/1/" data = requests.get(url) soup = BeautifulSoup(data.text) matchup_table = soup.find_all("div", {"class": "schedule"})[0] links = matchup_table.find_all('a') links = [l.get("href") for l in links] links = [l for l in links if '/box_score/' in l] box_scores_urls = [f"https://herhoopstats.com{l}" for l in links] for box_scores_url in box_scores_urls: data = requests.get(box_scores_url) # within loop opening up page and saving to folder in write mode with open("HerHoops/{}".format(box_scores_url[46:]), "w+") as f: # write to the files f.write(data.text) time.sleep(3)
报错信息
FileNotFoundError: [Errno 2] No such file or directory: 'HerHoops/2004/06/01/new-york-liberty-vs-charlotte-sting/'
问题原因
- 嵌套文件夹未创建:报错路径中的
2004/06/01是嵌套子文件夹,open()函数无法自动创建这些不存在的层级目录。 - 路径末尾斜杠导致识别错误:链接末尾的斜杠会被当作文件夹路径,而非文件名,系统找不到对应的文件。
- 固定索引提取路径风险:用
box_scores_url[46:]提取路径依赖固定位置,一旦链接结构变化就会出错。
解决方案
修改代码,添加文件夹创建逻辑,修正路径处理方式:
import requests from bs4 import BeautifulSoup import time import os url = f"https://herhoopstats.com/stats/wnba/schedule_date/2004/6/1/" data = requests.get(url) soup = BeautifulSoup(data.text) matchup_table = soup.find_all("div", {"class": "schedule"})[0] links = matchup_table.find_all('a') links = [l.get("href") for l in links] links = [l for l in links if '/box_score/' in l] box_scores_urls = [f"https://herhoopstats.com{l}" for l in links] for box_scores_url in box_scores_urls: data = requests.get(box_scores_url) # 提取box_score之后的路径部分,去掉末尾斜杠 path_part = box_scores_url.split('/box_score/')[-1].rstrip('/') # 拼接完整保存路径,添加.html后缀明确是网页文件 save_path = os.path.join("HerHoops", path_part) + ".html" # 创建所有必要的父文件夹 os.makedirs(os.path.dirname(save_path), exist_ok=True) # 写入文件 with open(save_path, "w+", encoding="utf-8") as f: f.write(data.text) time.sleep(3)
关键修改点
- 用
split('/box_score/')提取路径:避免固定索引的脆弱性,准确获取目标部分。 rstrip('/')去掉末尾斜杠:确保路径指向文件而非文件夹。os.makedirs(..., exist_ok=True):自动创建所有缺失的父文件夹,exist_ok=True避免文件夹已存在时报错。- 添加
.html后缀:明确保存的是网页文件,方便后续解析。 - 指定
encoding="utf-8":避免写入时出现乱码问题。
内容的提问来源于stack exchange,提问作者kc_balr
相关产品推荐
相关产品推荐

