You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助修正Google搜索爬虫代码:search()无num_results参数

问题修正:解决TypeError及Google搜索结果抓取问题

错误分析

你遇到的TypeError: search() got an unexpected keyword argument 'num_results'是因为googlesearch库的search函数不支持num_results参数,正确参数名是num(部分版本可用stop控制返回数量)。另外,search函数返回的是URL字符串列表,而非带link属性的对象,直接调用result.link会触发属性错误。

修正后的代码

# 导入所需库
from googlesearch import search
from bs4 import BeautifulSoup
import requests

# 执行Google搜索,获取前10个结果URL
results = search("Python", num=10)

# 打开文件准备写入结果
with open('resultados.txt', 'w', encoding='utf-8') as file:
    # 遍历每个搜索结果URL
    for url in results:
        try:
            # 发送HTTP请求获取页面内容
            response = requests.get(url, timeout=10)
            response.raise_for_status()  # 检查请求是否成功
            
            # 解析页面HTML
            soup = BeautifulSoup(response.text, 'html.parser')
            
            # 提取h1、h2、h3标签文本
            h1_texts = [tag.get_text(strip=True) for tag in soup.find_all('h1')]
            h2_texts = [tag.get_text(strip=True) for tag in soup.find_all('h2')]
            h3_texts = [tag.get_text(strip=True) for tag in soup.find_all('h3')]
            
            # 写入文件,添加分隔线区分不同页面
            file.write(f"=== 页面URL: {url} ===\n")
            file.write(f"h1: {h1_texts}\n")
            file.write(f"h2: {h2_texts}\n")
            file.write(f"h3: {h3_texts}\n\n")
            
        except Exception as e:
            # 捕获异常并记录错误信息
            file.write(f"=== 页面URL: {url} 抓取失败 ===\n")
            file.write(f"错误信息: {str(e)}\n\n")

关键修改说明

  • 将num_results=10替换为num=10,匹配googlesearch库的参数要求
  • 直接使用遍历得到的url字符串作为请求地址,替代错误的result.link
  • 添加encoding='utf-8'确保文件写入时不会出现中文乱码
  • 增加异常处理,避免单个页面请求失败导致整个程序终止
  • 使用get_text(strip=True)清理标签文本中的多余空格和换行
  • 添加页面URL和分隔线,让输出文件的结构更清晰

内容的提问来源于stack exchange,提问作者Tony Darko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 02:50:33