You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和BeautifulSoup爬取URL页面Excel文件名返回空列表如何解决

问题解决建议

1 核心API使用错误修正

你当前代码的问题在于find_all()方法不支持直接传入CSS选择器字符串,a[class*="file-id-"]是CSS选择器语法,需要替换为select()方法,或者使用find_all()的原生属性匹配规则,修正后的两种写法如下:

# 写法1:CSS选择器写法(与你预期的匹配规则完全一致,推荐)
labour_office_web_text = requests.get("url").text
soup = BeautifulSoup(labour_office_web_text, "lxml")
file_names = soup.select('a[class*="file-id-"]')
# 写法2:find_all原生属性匹配写法
import re
labour_office_web_text = requests.get("url").text
soup = BeautifulSoup(labour_office_web_text, "lxml")
file_names = soup.find_all('a', class_=re.compile(r'file-id-'))

2 修正后仍返回空列表的排查方案

  • 确认静态页面是否包含目标元素:打印labour_office_web_text,搜索file-id-是否存在。如果不存在,说明目标元素是JS动态渲染生成的,requests只能拿到静态源码,需要改用Selenium、Playwright等动态渲染工具发起请求,或者抓包找到前端加载文件列表调用的后端接口,直接请求接口获取数据。
  • 确认属性匹配规则正确:核对页面元素属性,确认file-id-是class属性值的一部分,而非id或其他属性,同时确认大小写完全匹配。
  • 确认解析器可用:如果环境未安装lxml,可将解析器替换为Python内置的html.parser测试,避免解析偏差:soup = BeautifulSoup(labour_office_web_text, "html.parser")

3 后续批量提取Excel文件的补充实现

拿到匹配的a标签后,你可以按以下逻辑筛选Excel文件、拼接下载链接,为后续批量读取数据做准备:

excel_file_list = []
base_domain = "你爬取的站点域名(比如https://xxx.gov.cn)"
for a_tag in file_names:
    # 提取a标签显示的文件名
    show_name = a_tag.get_text(strip=True)
    # 过滤非Excel文件
    if not show_name.endswith(('.xlsx', '.xls')):
        continue
    # 提取下载链接,相对路径自动拼接域名
    download_url = a_tag.get("href")
    if not download_url.startswith("http"):
        download_url = base_domain + download_url
    excel_file_list.append({
        "file_name": show_name,
        "download_url": download_url
    })

内容的提问来源于stack exchange,提问作者Robert Soroka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 23:18:02