You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从基础URL嵌套爬取并下载指定文件?

使用BeautifulSoup定位嵌套URL并下载ABS统计文件

以下是具体实现步骤,全程用Python的requests和BeautifulSoup完成:

步骤1:从基础页面抓取最新子页面链接

ABS的基础页面会优先展示最新季度的报告入口,我们可以通过文本特征定位到目标子页面链接:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

base_url = "https://www.abs.gov.au/statistics/economy/national-accounts/australian-national-accounts-national-income-expenditure-and-product"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}

# 获取基础页面内容
response = requests.get(base_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 定位最新子页面链接(优先找带"Latest release"标识的链接)
latest_link = soup.find("a", string=lambda text: text and "Latest release" in text)
if latest_link:
    # 拼接绝对URL
    subpage_url = urljoin(base_url, latest_link["href"])
else:
    # 备选:直接使用已知的目标子页面URL
    subpage_url = "https://www.abs.gov.au/statistics/economy/national-accounts/australian-national-accounts-national-income-expenditure-and-product/sep-2022#data-downloads"

步骤2:在子页面定位Table 1的下载链接

进入子页面后,直接定位到#data-downloads锚点对应的区域,筛选出Table 1的文件链接:

# 获取子页面内容
sub_response = requests.get(subpage_url, headers=headers)
sub_soup = BeautifulSoup(sub_response.text, "html.parser")

# 定位data-downloads区域
data_downloads_section = sub_soup.find(id="data-downloads")
if data_downloads_section:
    # 找文本包含"Table 1"的下载链接(通常是xlsx格式)
    table1_link = data_downloads_section.find("a", string=lambda text: text and "Table 1" in text)
    if table1_link:
        # 拼接绝对下载URL
        download_url = urljoin(subpage_url, table1_link["href"])
    else:
        print("未找到Table 1的下载链接")
else:
    print("未找到数据下载区域")

步骤3:下载目标文件

拿到下载链接后,直接请求并保存文件:

if download_url:
    # 获取文件内容
    file_response = requests.get(download_url, headers=headers)
    # 从URL提取文件名
    filename = download_url.split("/")[-1]
    # 保存文件到本地
    with open(filename, "wb") as f:
        f.write(file_response.content)
    print(f"文件 {filename} 下载完成")

注意事项

  • 页面结构可能变动,需根据实际HTML结构调整选择器(比如用class_属性定位元素,而非仅依赖文本)
  • 添加User-Agent是为了模拟浏览器请求,避免被ABS服务器拦截
  • 若遇到动态加载内容,可能需要改用selenium等工具,但ABS的静态页面用requests+BeautifulSoup足够

内容的提问来源于stack exchange,提问作者Pysparker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 15:00:42