如何用Python高效获取搜索引擎特定查询的结果数量?
需求
- 处理包含约40000条CVE已知企业软件漏洞实体的CSV数据集(实体示例:Lenovo EZ Media、Sylius、Oracle Corporation Hyperion Infrastructure Technology)
- 通过获取搜索引擎查询结果数量,衡量对应产品/企业的公众关注度与曝光度
- 核心要求:快速且可靠,无需极高精度
已尝试的方案及问题
方案1:BeautifulSoup + Requests
通过请求Bing搜索页面解析结果数量,但存在动态加载问题:搜索冷门内容时程序返回“结果数量不存在”,但手动在浏览器搜索后再运行程序却能获取结果。代码如下:
import requests import urllib3 from bs4 import BeautifulSoup urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning) def get_search_results_count(query): search_url = f"https://www.bing.com/search?q={query}" response = requests.get(search_url, verify=False) soup = BeautifulSoup(response.content, "html.parser") count_element = soup.find("span", class_="sb_count") if count_element: count_text = count_element.text.strip() return count_text else: return "Search results count not found" search_query = "cheese" result_count = get_search_results_count(search_query) print(result_count) numeric_part = ''.join(filter(str.isdigit, result_count)) print(numeric_part)
方案2:Selenium
通过模拟浏览器获取结果,但运行效率极低,且会打开本地浏览器窗口,完全不适合批量处理4万条数据。代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException def get_search_results_count(query): driver = webdriver.Edge() try: # Construct the search URL search_url = f"https://www.bing.com/search?q={query}" driver.get(search_url) count_element = driver.find_element(By.CLASS_NAME, "sb_count") count_text = count_element.text.strip() return count_text except NoSuchElementException: return "Search results count not found" finally: driver.quit() search_query = input("Enter the Company and Product name as one...:\n") result_count = get_search_results_count(search_query) print(result_count) numeric_part = ''.join(filter(str.isdigit, result_count)) search_results = numeric_part
期望解决方案
仅需获取搜索结果数量,无需处理实际搜索结果;认为Google搜索API过于复杂,不匹配当前需求,寻求可行的替代方案。
内容的提问来源于stack exchange,提问作者hrithik8555
相关产品推荐
相关产品推荐

