使用Python从网站下载指定PDF文件的问题排查
问题与解决方案
我正在攻读数据科学硕士学位,目前在完成一项文本挖掘作业。本项目中我需要从网站下载若干PDF文件,具体目标是爬取并保存名为"Prospectus(招股说明书)"的文档。以下是我使用的Python代码,我想要下载的招股说明书已在截图中展示,但当前脚本下载的是网页上的其他文档,请问我需要对脚本做哪些修改?
原代码
import os import requests from urllib.parse import urljoin from bs4 import BeautifulSoup url = "https://www.ishares.com/us/products/239726/ishares-core-sp-500-etf" # If there is no such folder, the script will create one automatically folder_location = r'.\Output' if not os.path.exists(folder_location): os.mkdir(folder_location) response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") for link in soup.select("a[href$='.pdf']"): # Name the pdf files using the last portion of each link which are unique in this case filename = os.path.join(folder_location, link['href'].split('/')[-1]) with open(filename, 'wb') as f: f.write(requests.get(urljoin(url, link['href'])).content)
目标文档截图

修改方案
当前脚本会下载页面上所有PDF链接,要只获取目标招股说明书,需通过链接的文本内容过滤。从截图可知,目标链接文本包含"Prospectus",据此调整代码即可:
修改后的代码
import os import requests from urllib.parse import urljoin from bs4 import BeautifulSoup url = "https://www.ishares.com/us/products/239726/ishares-core-sp-500-etf" # 自动创建输出文件夹 folder_location = r'.\Output' if not os.path.exists(folder_location): os.mkdir(folder_location) response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # 仅处理文本含"Prospectus"的PDF链接 for link in soup.select("a[href$='.pdf']"): link_text = link.get_text(strip=True) # 匹配目标关键词,支持不区分大小写 if 'prospectus' in link_text.lower(): filename = os.path.join(folder_location, link['href'].split('/')[-1]) with open(filename, 'wb') as f: f.write(requests.get(urljoin(url, link['href'])).content) print(f"已完成下载:{filename}")
核心修改点
- 新增文本过滤逻辑:通过
link.get_text(strip=True)提取链接文本,用lower()统一转小写后匹配"prospectus",避免大小写差异导致漏抓 - 优化代码格式:补充分号、缩进,提升可读性
- 增加下载提示:打印已下载文件路径,方便确认结果
内容的提问来源于stack exchange,提问作者Data_Science_Mick
相关产品推荐
相关产品推荐

