You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取谷歌搜索时出现字符串拼接错误求助

问题排查:Django中使用BeautifulSoup实现谷歌搜索时的字符串拼接错误

问题描述

在Django项目中使用BeautifulSoup实现谷歌搜索功能,尝试搜索时触发报错:

can only concatenate str (not "NoneType") to str

相关代码

search.py

from django.shortcuts import render, redirect
import requests
from bs4 import BeautifulSoup


def google(s):
    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.83 Safari/537.36'
    headers = {"user-agent": USER_AGENT}
    r=None
    links = []
    text = []
    r = requests.get("https://www.google.com/search?q=" + s, headers=headers)
    soup = BeautifulSoup(r.content, "html.parser")
    for g in soup.find_all('div', class_='yuRUbf'):
        a = g.find('a')
        t = g.find('h3')
        links.append(a.get('href'))
        text.append(t.text)   
        
        return links, text

views.py

from django.shortcuts import render, redirect
from netsurfers.search import google
from bs4 import BeautifulSoup
def home(request):
    return render(request,'home.html')

def results(request):
      if request.method == "POST":
        result = request.POST.get('search')
        google_link,google_text = google(result)
        google_data = zip(google_link,google_text)
        if result == '':
            return redirect('home')
        else:
            return render(request,'results.html',{'google': google_data})

urls.py

from django.contrib import admin
from django.urls import path,include
from . import views

urlpatterns = [
    path('admin/', admin.site.urls),
    path('', views.home,name='home'),
    path('results/',views.results,name='Result')
]

home模板

<form method='post' action="{% url 'Result' %}" class="d-flex" role="search">
                {% csrf_token %}
                 <input class="form-control me-2 " type="search" placeholder="ابحث وشارك بحثك مع الاخرين" aria-label="Search" style="width:22rem;">
                <input type="submit" class="btn btn-outline-success" value="ابحث" >
              </form>

results模板

{% for i,j in google %}
         <a href="{{ i }}" class="btn mt-3 w-100">{{ j }}</a><br>
        {% endfor %}

错误原因及修复方案

1. 输入框缺少name属性(核心报错原因)

home模板中的搜索输入框未设置name="search",导致request.POST.get('search')返回None。将None传入google函数拼接URL时,"https://www.google.com/search?q=" + None触发字符串拼接错误。

修复后的home模板:

<form method='post' action="{% url 'Result' %}" class="d-flex" role="search">
    {% csrf_token %}
    <input class="form-control me-2 " type="search" name="search" placeholder="ابحث وشارك بحثك مع الاخرين" aria-label="Search" style="width:22rem;">
    <input type="submit" class="btn btn-outline-success" value="ابحث" >
</form>

2. views中判断逻辑顺序错误

当前results函数先调用google(result),再检查输入是否为空。若result为空或None,会先触发拼接错误,无法执行后续跳转逻辑。需先验证输入合法性,再调用搜索函数。

修复后的views.py:

from django.shortcuts import render, redirect
from netsurfers.search import google

def home(request):
    return render(request,'home.html')

def results(request):
    if request.method == "POST":
        result = request.POST.get('search')
        # 先判断输入是否为空或None
        if not result:
            return redirect('home')
        google_link, google_text = google(result)
        google_data = zip(google_link, google_text)
        return render(request,'results.html',{'google': google_data})
    # 处理GET请求的情况
    return redirect('home')

3. google函数中的循环逻辑错误

google函数的return语句放在for循环内部,导致循环仅执行一次就返回,只能获取第一个搜索结果;同时未处理a或t为None的情况,可能触发属性访问错误。

修复后的search.py:

import requests
from bs4 import BeautifulSoup

# 无需导入Django的render和redirect,此函数仅处理搜索逻辑
def google(s):
    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.83 Safari/537.36'
    headers = {"user-agent": USER_AGENT}
    links = []
    text = []
    # 确保s是字符串,避免拼接错误
    if not isinstance(s, str):
        return links, text
    # 使用f-string拼接URL更安全
    r = requests.get(f"https://www.google.com/search?q={s}", headers=headers)
    soup = BeautifulSoup(r.content, "html.parser")
    for g in soup.find_all('div', class_='yuRUbf'):
        a = g.find('a')
        t = g.find('h3')
        # 检查a和t是否存在,再访问属性
        if a and t:
            links.append(a.get('href'))
            text.append(t.text)
    # 将return移到循环外,返回所有结果
    return links, text

补充说明

  • 谷歌可能拦截无授权的搜索请求,建议添加Accept-Language等请求头,或使用官方Custom Search API避免被封禁。
  • 处理网络请求时建议添加异常捕获(如requests.exceptions.RequestException),避免请求失败导致页面崩溃。

内容的提问来源于stack exchange,提问作者ashraf sadik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 00:25:31