You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法提取网页中Excel文件的Href链接求助

解决方法

1. 排查核心问题

你的代码仅匹配以https://开头的绝对路径链接,但目标Excel链接可能以相对路径(如/uploads/reports/xxx.xlsx)存在于页面中,直接被正则过滤;另外部分页面内容可能通过JavaScript动态渲染,静态请求工具(urllib/requests)无法获取到动态生成的链接。

2. 静态页面适配方案

如果链接是静态存在的,修改匹配规则,同时处理相对路径转绝对路径:

from bs4 import BeautifulSoup
import requests
import re

url = "https://ppac.gov.in/prices/international-prices-of-crude-oil"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

xlsx_links = []
# 匹配所有以.xlsx结尾的链接(兼容绝对/相对路径)
for link in soup.find_all('a', attrs={'href': re.compile(r'\.xlsx$')}):
    href = link.get('href')
    # 转换相对路径为完整绝对路径
    if not href.startswith('http'):
        href = f"https://ppac.gov.in{href}"
    xlsx_links.append(href)
    print(href)

3. 动态页面适配方案

如果页面通过JS动态渲染Excel链接,使用支持JS渲染的requests-html工具:

from requests_html import HTMLSession

url = "https://ppac.gov.in/prices/international-prices-of-crude-oil"
session = HTMLSession()
r = session.get(url)
# 渲染JS并等待页面加载完成
r.html.render(sleep=2)

xlsx_links = []
# 直接筛选所有.xlsx结尾的链接
for link in r.html.find('a[href$=".xlsx"]'):
    href = link.attrs['href']
    if not href.startswith('http'):
        href = f"https://ppac.gov.in{href}"
    xlsx_links.append(href)
    print(href)

内容的提问来源于stack exchange,提问作者Hunaidkhan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:55:57