Bing爬虫获取URL不稳定问题排查:为何部分关键词无结果?
问题分析与解决办法
你的代码爬取Bing搜索结果时出现时灵时不灵的情况,主要是以下几个原因导致的:
1. 关键词未做URL编码
直接把关键词拼接到URL里,虽然像"doctor"这种纯英文单词不会有格式问题,但未编码的请求可能被Bing的反爬系统标记;如果关键词包含空格、特殊字符(比如中文、&、=等),还会直接破坏URL结构,导致请求异常。
解决办法:
使用urllib.parse.quote()对关键词进行URL编码,确保请求URL格式完全合规:
from urllib.parse import quote def getBingLink(query): l=[] count = 0 num = random.uniform(1, 3) # 顺便延长休眠时间 time.sleep(num) # 对query进行URL编码 target_url="https://www.bing.com/search?q=" + quote(query) headers = {'User-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36'} resp=requests.get(target_url, headers ) print(resp.status_code) soup = BeautifulSoup(resp.text, 'html.parser') completeData = soup.find_all("li",{"class":"b_algo"}) if (len(completeData) == 0): # 调试用:打印响应内容预览,排查是否是验证页 print("无结果,响应预览:", resp.text[:500]) return l for item in completeData: o=item.find("a").get("href") l.append(o) return l
2. 反爬机制触发
你当前的休眠时间是0-1秒,太短,很容易被Bing识别为爬虫;同时请求头过于单一,缺少浏览器请求的必要字段,也会被标记。当触发反爬时,Bing会返回没有b_algo类的页面(比如人机验证页或空白结果页),导致你拿到空列表。
解决办法:
- 延长随机休眠时间到1-3秒,避免请求过于密集;
- 补充完整的请求头,模拟真实浏览器行为;
- 使用
requests.Session()维持会话,保持cookie一致性:
import requests from bs4 import BeautifulSoup import random import time from urllib.parse import quote def getBingLink(query): l=[] session = requests.Session() # 使用会话保持cookie num = random.uniform(1, 3) time.sleep(num) target_url="https://www.bing.com/search?q=" + quote(query) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.bing.com/' } resp=session.get(target_url, headers=headers) print(resp.status_code) soup = BeautifulSoup(resp.text, 'lxml') # 用lxml解析器更健壮 completeData = soup.find_all("li",{"class":"b_algo"}) if (len(completeData) == 0): print("无结果,响应预览:", resp.text[:500]) return l for item in completeData: o=item.find("a").get("href") l.append(o) return l
3. 页面解析器不够健壮
你使用的html.parser是Python内置的解析器,处理复杂HTML时容易出现解析错误,导致找不到目标元素。
解决办法:
安装并使用lxml解析器(需要先执行pip install lxml),它的解析效率和容错性都更强,代码里已经在上例中替换。
内容的提问来源于stack exchange,提问作者Michael Walker
相关产品推荐
相关产品推荐

