You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取Google统计指定词次数遇404错误求解决方案

解决Scrapy爬取Google搜索出现404的问题

先看你的错误日志,一眼就能发现核心问题:你请求的URL是https://www.google.co.in/searchq=newyear——这里的searchq明显不对,正确的Google搜索地址格式应该是search?q=关键词,少了个关键的问号,这就是返回404的直接原因。

问题分析

虽然你代码里写的是'https://www.google.co.in/search?q='+self.word,但可能是字符串拼接时的隐性问题(比如参数含特殊字符干扰),或者运行时的意外情况导致问号丢失。另外,Google的反爬机制也会给爬虫带来麻烦,就算URL正确,也大概率会被拦截。

修复方案

我给你调整了代码,主要做了这几点优化:

  • 使用urllib.parse.urlencode安全构建查询参数,彻底避免URL拼接错误
  • 更换了更贴近真实浏览器的User-Agent标识
  • 增加了基础的反爬应对(设置下载延迟、处理特殊关键词的正则匹配)
  • 加入了拦截检测逻辑,方便排查反爬问题

修改后的代码:

import scrapy
import re
from urllib.parse import urlencode
from scrapy.crawler import CrawlerProcess
import sys

class GoogleSpider(scrapy.Spider):
    name = 'Google'
    allowed_domains = ['www.google.co.in']
    
    def __init__(self, word=None):
        super().__init__()
        self.word = word.lower()  # 统一转小写,避免大小写匹配偏差
        # 使用urlencode构建查询参数,自动处理特殊字符,避免拼接错误
        params = urlencode({'q': self.word})
        self.start_urls = [f'https://www.google.co.in/search?{params}']

    def parse(self, response):
        print('当前请求URL:', response.url)
        # 检查是否被Google的反爬机制拦截
        if 'captcha' in response.url or 'sorry' in response.url:
            print('警告:请求被Google拦截,可能需要人机验证、更换代理或调整User-Agent')
            return
        
        # 提取搜索结果中的文本内容
        text = response.xpath('//div[@class="g"]//text()').extract()
        text = ''.join(text).lower()
        # 使用re.escape处理关键词中的正则特殊字符(比如.、*等)
        count = len(re.findall(re.escape(self.word), text))
        print(f'单词"{self.word}"在搜索结果中的出现次数:{count}')

if __name__ == "__main__":
    # 检查是否传入了搜索关键词参数
    if len(sys.argv) < 2:
        print('请传入要搜索的单词,示例:python your_script.py newyear')
        sys.exit(1)
    
    process = CrawlerProcess({
        'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'DOWNLOAD_DELAY': 2,  # 设置下载延迟,降低被封IP的风险
    })
    process.crawl(GoogleSpider, word=sys.argv[1])
    process.start()

额外注意事项

  1. 反爬应对:Google的反爬非常严格,就算用了上述代码,也可能遇到人机验证或IP封禁。如果需要稳定爬取,建议考虑使用Google自定义搜索API(需申请密钥),或者搭建代理IP池。
  2. 运行方式:确保运行时传入搜索关键词参数,比如在终端执行:
python your_script.py newyear

内容的提问来源于stack exchange,提问作者khaja mohiddin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:52:57