You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup爬取矿业公司可行性研究PDF报告报错KeyError 'href'

Python爬虫获取Google搜索PDF链接时报KeyError 'href'的解决办法

我想写Python爬虫搜索矿业公司的可行性研究报告,优先获取PDF格式的。写了下面的代码,但最后一行报错KeyError 'href',请问问题出在哪?

import requests
from bs4 import BeautifulSoup
import pandas as pd

# Set the search query
query = 'mining companies pre-feasibility feasibility studies'

# Set the URL of the search results page
url = 'https://www.google.com.au/search?q=' + query

# Set the user-agent to avoid being detected as a scraper
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36'}

# Send a GET request to the search results page
response = requests.get(url, headers=headers)

# Parse the HTML content of the search results page with Beautiful Soup
soup = BeautifulSoup(response.content, 'html.parser')

# Find all the search result links
links = soup.find_all('a')

# Filter the links to those that point to PDF documents
pdf_links = [link['href'] for link in links if link['href'].endswith('.pdf')]

问题原因

  1. 并非所有<a>标签都包含href属性:soup.find_all('a')会抓取页面上所有超链接标签,但页面里的导航按钮、交互元素等部分<a>标签可能没有href属性,直接用link['href']取值时,就会触发KeyError。
  2. Google搜索结果的链接结构特殊:搜索结果里的<a>标签的href是Google的中转链接(格式类似/url?q=实际链接&sa=U...),并非直接指向目标PDF,直接判断href是否以.pdf结尾找不到有效结果。

解决办法

修改代码,先过滤无href的标签,再解析Google的中转链接提取真实地址,同时给搜索query加上filetype:pdf让Google直接返回PDF结果,提升效率:

import requests
from bs4 import BeautifulSoup
import urllib.parse

# 搜索query加上filetype:pdf,让Google直接筛选PDF结果
query = 'mining companies pre-feasibility feasibility studies filetype:pdf'
# 对query进行URL编码,避免特殊字符干扰
url = 'https://www.google.com.au/search?q=' + urllib.parse.quote(query)

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36'}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')

pdf_links = []
for link in soup.find_all('a'):
    # 用get方法获取href,不存在则跳过,避免KeyError
    href = link.get('href')
    if not href:
        continue
    # 识别Google中转链接,提取真实目标URL
    if href.startswith('/url?q='):
        # 拆分出实际链接部分并解码URL编码
        actual_url = urllib.parse.unquote(href.split('/url?q=')[1].split('&')[0])
        # 判断是否为PDF格式链接
        if actual_url.endswith('.pdf'):
            pdf_links.append(actual_url)

# 输出有效PDF链接
for idx, link in enumerate(pdf_links, 1):
    print(f"{idx}. {link}")

关键修改说明

  • 添加filetype:pdf搜索限定:让Google直接返回PDF格式结果,减少无效数据处理。
  • 使用link.get('href')替代link['href']:get方法在属性不存在时返回None,不会抛出异常,代码更稳健。
  • 解析中转链接:通过拆分/url?q=后的内容,提取真实的PDF目标链接,确保筛选结果准确。

内容的提问来源于stack exchange,提问作者Spooked

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 22:17:36