使用BeautifulSoup无法提取网页中Excel文件的Href链接求助
解决方法
1. 排查核心问题
你的代码仅匹配以https://开头的绝对路径链接,但目标Excel链接可能以相对路径(如/uploads/reports/xxx.xlsx)存在于页面中,直接被正则过滤;另外部分页面内容可能通过JavaScript动态渲染,静态请求工具(urllib/requests)无法获取到动态生成的链接。
2. 静态页面适配方案
如果链接是静态存在的,修改匹配规则,同时处理相对路径转绝对路径:
from bs4 import BeautifulSoup import requests import re url = "https://ppac.gov.in/prices/international-prices-of-crude-oil" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") xlsx_links = [] # 匹配所有以.xlsx结尾的链接(兼容绝对/相对路径) for link in soup.find_all('a', attrs={'href': re.compile(r'\.xlsx$')}): href = link.get('href') # 转换相对路径为完整绝对路径 if not href.startswith('http'): href = f"https://ppac.gov.in{href}" xlsx_links.append(href) print(href)
3. 动态页面适配方案
如果页面通过JS动态渲染Excel链接,使用支持JS渲染的requests-html工具:
from requests_html import HTMLSession url = "https://ppac.gov.in/prices/international-prices-of-crude-oil" session = HTMLSession() r = session.get(url) # 渲染JS并等待页面加载完成 r.html.render(sleep=2) xlsx_links = [] # 直接筛选所有.xlsx结尾的链接 for link in r.html.find('a[href$=".xlsx"]'): href = link.attrs['href'] if not href.startswith('http'): href = f"https://ppac.gov.in{href}" xlsx_links.append(href) print(href)
内容的提问来源于stack exchange,提问作者Hunaidkhan
相关产品推荐
相关产品推荐

