You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法使用requests或Selenium抓取页面href链接,求解决方案

如何提取网页中的PDF链接?

我需要提取指定页面的所有href链接,并筛选出其中的.pdf链接。尝试用requests库和Selenium工具提取都没成功,该怎么解决?谢谢。

示例:包含.pdf文件链接的情况
示例图片

我使用的requests代码:

import requests
from bs4 import BeautifulSoup

headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/113.0'}
url="https://www.bain.com/insights/topics/energy-and-natural-resources-report/"

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

for link in soup.find_all('a'):
    print(link.get('href'))

我使用的Selenium代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup

options = webdriver.ChromeOptions()
driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()), options=options)

page_source = driver.get("https://www.bain.com/insights/topics/energy-and-natural-resource-report/")
driver.implicitly_wait(10)

soup = BeautifulSoup(page_source, 'html.parser')
for link in soup.find_all('a'):
    print(link.get('href'))

driver.quit()

问题修正方案

一、requests代码的问题与修复

你的requests代码仅打印所有a标签的href,但没做PDF筛选,且未处理相对路径。如果页面内容是静态渲染的,修正后即可获取PDF链接:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/113.0'}
base_url = "https://www.bain.com/insights/topics/energy-and-natural-resources-report/"

response = requests.get(base_url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

pdf_links = []
# 只遍历带有href属性的a标签
for link in soup.find_all('a', href=True):
    href = link['href']
    # 把相对链接拼接成完整URL
    full_url = urljoin(base_url, href)
    # 筛选以.pdf结尾的链接
    if full_url.endswith('.pdf'):
        pdf_links.append(full_url)

print("找到的PDF链接:")
for link in pdf_links:
    print(link)

二、Selenium代码的问题与修复

你的Selenium代码存在关键错误:driver.get()返回None,不能直接赋值给page_source,需用driver.page_source获取页面源码。同时补充PDF筛选和路径处理逻辑:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
from urllib.parse import urljoin

options = webdriver.ChromeOptions()
# 可选:添加无头模式,后台运行浏览器
options.add_argument('--headless=new')
driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()), options=options)

base_url = "https://www.bain.com/insights/topics/energy-and-natural-resources-report/"
driver.get(base_url)
# 隐式等待需放在get之后,确保页面元素加载完成
driver.implicitly_wait(10)

# 获取完整页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'html.parser')

pdf_links = []
for link in soup.find_all('a', href=True):
    href = link['href']
    full_url = urljoin(base_url, href)
    if full_url.endswith('.pdf'):
        pdf_links.append(full_url)

print("找到的PDF链接:")
for link in pdf_links:
    print(link)

driver.quit()

额外说明

  • 如果页面存在滚动加载或延迟渲染的内容,可使用Selenium的显式等待,等待目标元素加载完成后再提取源码
  • 部分网站有反爬机制,可添加更多请求头或使用代理IP避免被封禁

内容的提问来源于stack exchange,提问作者dfcsdf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 21:54:51