You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python代码仅下载CAG网站‘Monthly Key Indicators’标签下的PDF

代码修改方案:仅下载“Monthly Key Indicators”下的PDF

修改思路

  1. 定位页面中“Monthly Key Indicators”对应的专属内容区块,缩小PDF链接的查找范围
  2. 仅在该区块内筛选并下载PDF文件,避免获取其他区域的无关文件

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import os
from urllib.parse import urljoin

url = 'https://cag.gov.in/en/state-accounts-report?defuat_state_id=64'
folder_location = "./cag_monthly_indicators"  # 替换为你需要的保存路径

# 自动创建保存文件夹
if not os.path.exists(folder_location):
    os.makedirs(folder_location)

response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位目标内容区块:先匹配标题文本,再获取对应的内容容器
target_section = None
# 遍历可能的标题标签,匹配目标文本
for heading in soup.find_all(['h3', 'div'], text=lambda t: t and 'Monthly Key Indicators' in t.strip()):
    # 根据页面实际结构,获取标题对应的内容容器(若页面结构调整,需修改此处选择器)
    target_section = heading.find_next('div', class_='tab-pane')
    if target_section:
        break

# 仅在目标区块内下载PDF
if target_section:
    for link in target_section.select("a[href$='.pdf']"):
        filename = os.path.join(folder_location, link['href'].split('/')[-1])
        full_pdf_url = urljoin(url, link['href'])
        with open(filename, 'wb') as f:
            f.write(requests.get(full_pdf_url).content)
        print(f"已完成下载: {filename}")
else:
    print("未找到'Monthly Key Indicators'对应的内容区块,请检查页面结构或调整选择器")

关键修改说明

  • 区块精准定位:通过文本匹配找到目标标题,再关联到其对应的内容容器(若页面结构更新,需微调find_next中的标签或class选择器)
  • 范围限制:将PDF链接的查找范围从整个页面缩小到目标区块,确保只下载指定板块的文件
  • 容错提示:增加区块未找到时的提示信息,便于排查页面结构变化导致的问题

内容的提问来源于stack exchange,提问作者NoobCoderPy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 16:35:27