如何修改Python代码仅下载CAG网站‘Monthly Key Indicators’标签下的PDF
代码修改方案:仅下载“Monthly Key Indicators”下的PDF
修改思路
- 定位页面中“Monthly Key Indicators”对应的专属内容区块,缩小PDF链接的查找范围
- 仅在该区块内筛选并下载PDF文件,避免获取其他区域的无关文件
修改后的完整代码
import requests from bs4 import BeautifulSoup import os from urllib.parse import urljoin url = 'https://cag.gov.in/en/state-accounts-report?defuat_state_id=64' folder_location = "./cag_monthly_indicators" # 替换为你需要的保存路径 # 自动创建保存文件夹 if not os.path.exists(folder_location): os.makedirs(folder_location) response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 定位目标内容区块:先匹配标题文本,再获取对应的内容容器 target_section = None # 遍历可能的标题标签,匹配目标文本 for heading in soup.find_all(['h3', 'div'], text=lambda t: t and 'Monthly Key Indicators' in t.strip()): # 根据页面实际结构,获取标题对应的内容容器(若页面结构调整,需修改此处选择器) target_section = heading.find_next('div', class_='tab-pane') if target_section: break # 仅在目标区块内下载PDF if target_section: for link in target_section.select("a[href$='.pdf']"): filename = os.path.join(folder_location, link['href'].split('/')[-1]) full_pdf_url = urljoin(url, link['href']) with open(filename, 'wb') as f: f.write(requests.get(full_pdf_url).content) print(f"已完成下载: {filename}") else: print("未找到'Monthly Key Indicators'对应的内容区块,请检查页面结构或调整选择器")
关键修改说明
- 区块精准定位:通过文本匹配找到目标标题,再关联到其对应的内容容器(若页面结构更新,需微调
find_next中的标签或class选择器) - 范围限制:将PDF链接的查找范围从整个页面缩小到目标区块,确保只下载指定板块的文件
- 容错提示:增加区块未找到时的提示信息,便于排查页面结构变化导致的问题
内容的提问来源于stack exchange,提问作者NoobCoderPy
相关产品推荐
相关产品推荐

