使用BeautifulSoup爬取谷歌搜索时出现字符串拼接错误求助
问题排查:Django中使用BeautifulSoup实现谷歌搜索时的字符串拼接错误
问题描述
在Django项目中使用BeautifulSoup实现谷歌搜索功能,尝试搜索时触发报错:
can only concatenate str (not "NoneType") to str
相关代码
search.py
from django.shortcuts import render, redirect import requests from bs4 import BeautifulSoup def google(s): USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.83 Safari/537.36' headers = {"user-agent": USER_AGENT} r=None links = [] text = [] r = requests.get("https://www.google.com/search?q=" + s, headers=headers) soup = BeautifulSoup(r.content, "html.parser") for g in soup.find_all('div', class_='yuRUbf'): a = g.find('a') t = g.find('h3') links.append(a.get('href')) text.append(t.text) return links, text
views.py
from django.shortcuts import render, redirect from netsurfers.search import google from bs4 import BeautifulSoup def home(request): return render(request,'home.html') def results(request): if request.method == "POST": result = request.POST.get('search') google_link,google_text = google(result) google_data = zip(google_link,google_text) if result == '': return redirect('home') else: return render(request,'results.html',{'google': google_data})
urls.py
from django.contrib import admin from django.urls import path,include from . import views urlpatterns = [ path('admin/', admin.site.urls), path('', views.home,name='home'), path('results/',views.results,name='Result') ]
home模板
<form method='post' action="{% url 'Result' %}" class="d-flex" role="search"> {% csrf_token %} <input class="form-control me-2 " type="search" placeholder="ابحث وشارك بحثك مع الاخرين" aria-label="Search" style="width:22rem;"> <input type="submit" class="btn btn-outline-success" value="ابحث" > </form>
results模板
{% for i,j in google %} <a href="{{ i }}" class="btn mt-3 w-100">{{ j }}</a><br> {% endfor %}
错误原因及修复方案
1. 输入框缺少name属性(核心报错原因)
home模板中的搜索输入框未设置name="search",导致request.POST.get('search')返回None。将None传入google函数拼接URL时,"https://www.google.com/search?q=" + None触发字符串拼接错误。
修复后的home模板:
<form method='post' action="{% url 'Result' %}" class="d-flex" role="search"> {% csrf_token %} <input class="form-control me-2 " type="search" name="search" placeholder="ابحث وشارك بحثك مع الاخرين" aria-label="Search" style="width:22rem;"> <input type="submit" class="btn btn-outline-success" value="ابحث" > </form>
2. views中判断逻辑顺序错误
当前results函数先调用google(result),再检查输入是否为空。若result为空或None,会先触发拼接错误,无法执行后续跳转逻辑。需先验证输入合法性,再调用搜索函数。
修复后的views.py:
from django.shortcuts import render, redirect from netsurfers.search import google def home(request): return render(request,'home.html') def results(request): if request.method == "POST": result = request.POST.get('search') # 先判断输入是否为空或None if not result: return redirect('home') google_link, google_text = google(result) google_data = zip(google_link, google_text) return render(request,'results.html',{'google': google_data}) # 处理GET请求的情况 return redirect('home')
3. google函数中的循环逻辑错误
google函数的return语句放在for循环内部,导致循环仅执行一次就返回,只能获取第一个搜索结果;同时未处理a或t为None的情况,可能触发属性访问错误。
修复后的search.py:
import requests from bs4 import BeautifulSoup # 无需导入Django的render和redirect,此函数仅处理搜索逻辑 def google(s): USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.83 Safari/537.36' headers = {"user-agent": USER_AGENT} links = [] text = [] # 确保s是字符串,避免拼接错误 if not isinstance(s, str): return links, text # 使用f-string拼接URL更安全 r = requests.get(f"https://www.google.com/search?q={s}", headers=headers) soup = BeautifulSoup(r.content, "html.parser") for g in soup.find_all('div', class_='yuRUbf'): a = g.find('a') t = g.find('h3') # 检查a和t是否存在,再访问属性 if a and t: links.append(a.get('href')) text.append(t.text) # 将return移到循环外,返回所有结果 return links, text
补充说明
- 谷歌可能拦截无授权的搜索请求,建议添加
Accept-Language等请求头,或使用官方Custom Search API避免被封禁。 - 处理网络请求时建议添加异常捕获(如
requests.exceptions.RequestException),避免请求失败导致页面崩溃。
内容的提问来源于stack exchange,提问作者ashraf sadik
相关产品推荐
相关产品推荐

