You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Bing爬虫获取URL不稳定问题排查:为何部分关键词无结果?

问题分析与解决办法

你的代码爬取Bing搜索结果时出现时灵时不灵的情况,主要是以下几个原因导致的:

1. 关键词未做URL编码

直接把关键词拼接到URL里,虽然像"doctor"这种纯英文单词不会有格式问题,但未编码的请求可能被Bing的反爬系统标记;如果关键词包含空格、特殊字符(比如中文、&、=等),还会直接破坏URL结构,导致请求异常。

解决办法:
使用urllib.parse.quote()对关键词进行URL编码,确保请求URL格式完全合规:

from urllib.parse import quote

def getBingLink(query):
    l=[]
    count = 0
    num = random.uniform(1, 3)  # 顺便延长休眠时间
    time.sleep(num)
    # 对query进行URL编码
    target_url="https://www.bing.com/search?q=" + quote(query) 
    headers = {'User-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36'}
    resp=requests.get(target_url, headers )
    print(resp.status_code)
    soup = BeautifulSoup(resp.text, 'html.parser')
    completeData = soup.find_all("li",{"class":"b_algo"})
    if (len(completeData) == 0):
        # 调试用:打印响应内容预览,排查是否是验证页
        print("无结果,响应预览:", resp.text[:500])
        return l
    for item in completeData:       
        o=item.find("a").get("href")    
        l.append(o)
    return l

2. 反爬机制触发

你当前的休眠时间是0-1秒,太短,很容易被Bing识别为爬虫;同时请求头过于单一,缺少浏览器请求的必要字段,也会被标记。当触发反爬时,Bing会返回没有b_algo类的页面(比如人机验证页或空白结果页),导致你拿到空列表。

解决办法:

  • 延长随机休眠时间到1-3秒,避免请求过于密集;
  • 补充完整的请求头,模拟真实浏览器行为;
  • 使用requests.Session()维持会话,保持cookie一致性:
import requests
from bs4 import BeautifulSoup
import random
import time
from urllib.parse import quote

def getBingLink(query):
    l=[]
    session = requests.Session()  # 使用会话保持cookie
    num = random.uniform(1, 3)
    time.sleep(num)
    target_url="https://www.bing.com/search?q=" + quote(query) 
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
        'Accept-Language': 'en-US,en;q=0.5',
        'Referer': 'https://www.bing.com/'
    }
    resp=session.get(target_url, headers=headers)
    print(resp.status_code)
    soup = BeautifulSoup(resp.text, 'lxml')  # 用lxml解析器更健壮
    completeData = soup.find_all("li",{"class":"b_algo"})
    if (len(completeData) == 0):
        print("无结果,响应预览:", resp.text[:500])
        return l
    for item in completeData:       
        o=item.find("a").get("href")    
        l.append(o)
    return l

3. 页面解析器不够健壮

你使用的html.parser是Python内置的解析器,处理复杂HTML时容易出现解析错误,导致找不到目标元素。

解决办法:
安装并使用lxml解析器(需要先执行pip install lxml),它的解析效率和容错性都更强,代码里已经在上例中替换。


内容的提问来源于stack exchange,提问作者Michael Walker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 08:03:18