如何用BeautifulSoup从基础URL嵌套爬取并下载指定文件?
使用BeautifulSoup定位嵌套URL并下载ABS统计文件
以下是具体实现步骤,全程用Python的requests和BeautifulSoup完成:
步骤1:从基础页面抓取最新子页面链接
ABS的基础页面会优先展示最新季度的报告入口,我们可以通过文本特征定位到目标子页面链接:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin base_url = "https://www.abs.gov.au/statistics/economy/national-accounts/australian-national-accounts-national-income-expenditure-and-product" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} # 获取基础页面内容 response = requests.get(base_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 定位最新子页面链接(优先找带"Latest release"标识的链接) latest_link = soup.find("a", string=lambda text: text and "Latest release" in text) if latest_link: # 拼接绝对URL subpage_url = urljoin(base_url, latest_link["href"]) else: # 备选:直接使用已知的目标子页面URL subpage_url = "https://www.abs.gov.au/statistics/economy/national-accounts/australian-national-accounts-national-income-expenditure-and-product/sep-2022#data-downloads"
步骤2:在子页面定位Table 1的下载链接
进入子页面后,直接定位到#data-downloads锚点对应的区域,筛选出Table 1的文件链接:
# 获取子页面内容 sub_response = requests.get(subpage_url, headers=headers) sub_soup = BeautifulSoup(sub_response.text, "html.parser") # 定位data-downloads区域 data_downloads_section = sub_soup.find(id="data-downloads") if data_downloads_section: # 找文本包含"Table 1"的下载链接(通常是xlsx格式) table1_link = data_downloads_section.find("a", string=lambda text: text and "Table 1" in text) if table1_link: # 拼接绝对下载URL download_url = urljoin(subpage_url, table1_link["href"]) else: print("未找到Table 1的下载链接") else: print("未找到数据下载区域")
步骤3:下载目标文件
拿到下载链接后,直接请求并保存文件:
if download_url: # 获取文件内容 file_response = requests.get(download_url, headers=headers) # 从URL提取文件名 filename = download_url.split("/")[-1] # 保存文件到本地 with open(filename, "wb") as f: f.write(file_response.content) print(f"文件 {filename} 下载完成")
注意事项
- 页面结构可能变动,需根据实际HTML结构调整选择器(比如用
class_属性定位元素,而非仅依赖文本) - 添加
User-Agent是为了模拟浏览器请求,避免被ABS服务器拦截 - 若遇到动态加载内容,可能需要改用
selenium等工具,但ABS的静态页面用requests+BeautifulSoup足够
内容的提问来源于stack exchange,提问作者Pysparker
相关产品推荐
相关产品推荐

