You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup解析谷歌学术时遇TypeError错误求助

错误原因与修正方案

核心错误

result.select('.gs_a a')返回的是一个BeautifulSoup标签列表,你直接用['href']访问列表元素,违反了列表只能用整数/切片做索引的规则,因此抛出TypeError。

额外语法问题

params字典里的'q': 'Machine learning,缺少闭合单引号,会导致代码直接报错无法运行。

修正后的代码

import requests
from bs4 import BeautifulSoup

headers = {
    'User-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36'
}

params = {
    'q': 'Machine learning',  # 补上缺失的闭合单引号
    'hl': 'en'
}
html = requests.get('https://scholar.google.com/scholar', headers=headers, params=params).text
soup = BeautifulSoup(html, 'lxml')
for result in soup.select('.gs_r.gs_or.gs_scl'):
    # 遍历所有匹配的a标签,提取href属性存入列表
    profiles = [a['href'] for a in result.select('.gs_a a')]
    print(profiles)
    
    # 若仅需第一个作者的链接,可添加非空判断避免索引越界
    # if result.select('.gs_a a'):
    #     first_profile = result.select('.gs_a a')[0]['href']
    #     print(first_profile)

修正说明

  1. 补全params中q参数的闭合引号,修复基础语法错误。
  2. 用列表推导式遍历select返回的标签列表,逐个提取href,确保符合列表操作规则。
  3. 增加了单链接提取的备选方案,同时加入非空判断防止索引越界。

内容的提问来源于stack exchange,提问作者Saad Hamim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:15:32