无法使用requests或Selenium抓取页面href链接,求解决方案
如何提取网页中的PDF链接?
我需要提取指定页面的所有href链接,并筛选出其中的.pdf链接。尝试用requests库和Selenium工具提取都没成功,该怎么解决?谢谢。
示例:包含.pdf文件链接的情况
我使用的requests代码:
import requests from bs4 import BeautifulSoup headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/113.0'} url="https://www.bain.com/insights/topics/energy-and-natural-resources-report/" response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') for link in soup.find_all('a'): print(link.get('href'))
我使用的Selenium代码:
from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager from bs4 import BeautifulSoup options = webdriver.ChromeOptions() driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()), options=options) page_source = driver.get("https://www.bain.com/insights/topics/energy-and-natural-resource-report/") driver.implicitly_wait(10) soup = BeautifulSoup(page_source, 'html.parser') for link in soup.find_all('a'): print(link.get('href')) driver.quit()
问题修正方案
一、requests代码的问题与修复
你的requests代码仅打印所有a标签的href,但没做PDF筛选,且未处理相对路径。如果页面内容是静态渲染的,修正后即可获取PDF链接:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/113.0'} base_url = "https://www.bain.com/insights/topics/energy-and-natural-resources-report/" response = requests.get(base_url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') pdf_links = [] # 只遍历带有href属性的a标签 for link in soup.find_all('a', href=True): href = link['href'] # 把相对链接拼接成完整URL full_url = urljoin(base_url, href) # 筛选以.pdf结尾的链接 if full_url.endswith('.pdf'): pdf_links.append(full_url) print("找到的PDF链接:") for link in pdf_links: print(link)
二、Selenium代码的问题与修复
你的Selenium代码存在关键错误:driver.get()返回None,不能直接赋值给page_source,需用driver.page_source获取页面源码。同时补充PDF筛选和路径处理逻辑:
from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager from bs4 import BeautifulSoup from urllib.parse import urljoin options = webdriver.ChromeOptions() # 可选:添加无头模式,后台运行浏览器 options.add_argument('--headless=new') driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()), options=options) base_url = "https://www.bain.com/insights/topics/energy-and-natural-resources-report/" driver.get(base_url) # 隐式等待需放在get之后,确保页面元素加载完成 driver.implicitly_wait(10) # 获取完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') pdf_links = [] for link in soup.find_all('a', href=True): href = link['href'] full_url = urljoin(base_url, href) if full_url.endswith('.pdf'): pdf_links.append(full_url) print("找到的PDF链接:") for link in pdf_links: print(link) driver.quit()
额外说明
- 如果页面存在滚动加载或延迟渲染的内容,可使用Selenium的显式等待,等待目标元素加载完成后再提取源码
- 部分网站有反爬机制,可添加更多请求头或使用代理IP避免被封禁
内容的提问来源于stack exchange,提问作者dfcsdf
相关产品推荐
相关产品推荐

