You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup选择器报NotImplementedError,求Colab适配方案

修复Colab中BeautifulSoup选择器的报错问题

原代码在Google Colab运行时触发NotImplementedError,核心原因是BeautifulSoup内置的CSS选择器仅支持nth-of-type伪类,不兼容:last-of-type和:has()这类伪类语法。

原代码

import requests
import pandas as pd
from bs4 import BeautifulSoup

html = requests.get("https://www.tce.sp.gov.br/jurisprudencia/exibir?proc=18955/989/20&offset=0")

soup = BeautifulSoup(html.content)

data = []

for e in soup.select('table:last-of-type tr:has(td)'):
    it = iter(soup.table.stripped_strings)
    d = dict(zip(it,it))
    d.update({
        'link': e.a.get('href'),
        'date': e.select('td')[-2].text,
        'type': e.select('td')[-1].text
    })
    data.append(d)

报错信息

NotImplementedError                       Traceback (most recent call last)
<ipython-input-14-c9c2af04191b> in <module>
      9 data = []
     10 
---> 11 for e in soup.select('table:last-of-type tr:has(td)'):
     12     it = iter(soup.table.stripped_strings)
     13     d = dict(zip(it,it))

/usr/local/lib/python3.7/dist-packages/bs4/element.py in select(self, selector, _candidate_generator, limit)
   1526                 else:
   1527                     raise NotImplementedError(
-> 1528                         'Only the following pseudo-classes are implemented: nth-of-type.')
   1529 
   1530             elif token == '*':

NotImplementedError: Only the following pseudo-classes are implemented: nth-of-type.

修复后的代码

import requests
import pandas as pd
from bs4 import BeautifulSoup

html = requests.get("https://www.tce.sp.gov.br/jurisprudencia/exibir?proc=18955/989/20&offset=0")

soup = BeautifulSoup(html.content, 'html.parser')

data = []

# 获取页面所有表格,取最后一个目标表格
tables = soup.find_all('table')
if tables:
    target_table = tables[-1]
    # 遍历目标表格下的所有行
    for row in target_table.find_all('tr'):
        tds = row.find_all('td')
        # 跳过无内容的行
        if len(tds) < 3:
            continue
        # 提取页面顶部的主信息
        main_info = {}
        main_table = soup.find('table')
        if main_table:
            string_iter = iter(main_table.stripped_strings)
            main_info = dict(zip(string_iter, string_iter))
        # 拼接当前行的详情信息
        row_data = main_info.copy()
        row_data.update({
            'link': tds[0].a.get('href') if tds[0].a else None,
            'date': tds[-2].get_text(strip=True),
            'type': tds[-1].get_text(strip=True)
        })
        data.append(row_data)

# 转换为DataFrame查看结果
df = pd.DataFrame(data)
print(df)

修改说明

  1. 替换不兼容选择器:用find_all('table')获取所有表格后取最后一个tables[-1],替代:last-of-type;通过判断行内td数量筛选有效行,替代:has(td)。
  2. 增加容错判断:避免表格不存在、行元素不足等场景触发报错。
  3. 优化文本提取:用get_text(strip=True)自动清理文本中的多余空格和换行符。

内容的提问来源于stack exchange,提问作者João Pedro Rodrigues Oliveira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 18:55:23