BeautifulSoup选择器报NotImplementedError,求Colab适配方案
修复Colab中BeautifulSoup选择器的报错问题
原代码在Google Colab运行时触发NotImplementedError,核心原因是BeautifulSoup内置的CSS选择器仅支持nth-of-type伪类,不兼容:last-of-type和:has()这类伪类语法。
原代码
import requests import pandas as pd from bs4 import BeautifulSoup html = requests.get("https://www.tce.sp.gov.br/jurisprudencia/exibir?proc=18955/989/20&offset=0") soup = BeautifulSoup(html.content) data = [] for e in soup.select('table:last-of-type tr:has(td)'): it = iter(soup.table.stripped_strings) d = dict(zip(it,it)) d.update({ 'link': e.a.get('href'), 'date': e.select('td')[-2].text, 'type': e.select('td')[-1].text }) data.append(d)
报错信息
NotImplementedError Traceback (most recent call last) <ipython-input-14-c9c2af04191b> in <module> 9 data = [] 10 ---> 11 for e in soup.select('table:last-of-type tr:has(td)'): 12 it = iter(soup.table.stripped_strings) 13 d = dict(zip(it,it)) /usr/local/lib/python3.7/dist-packages/bs4/element.py in select(self, selector, _candidate_generator, limit) 1526 else: 1527 raise NotImplementedError( -> 1528 'Only the following pseudo-classes are implemented: nth-of-type.') 1529 1530 elif token == '*': NotImplementedError: Only the following pseudo-classes are implemented: nth-of-type.
修复后的代码
import requests import pandas as pd from bs4 import BeautifulSoup html = requests.get("https://www.tce.sp.gov.br/jurisprudencia/exibir?proc=18955/989/20&offset=0") soup = BeautifulSoup(html.content, 'html.parser') data = [] # 获取页面所有表格,取最后一个目标表格 tables = soup.find_all('table') if tables: target_table = tables[-1] # 遍历目标表格下的所有行 for row in target_table.find_all('tr'): tds = row.find_all('td') # 跳过无内容的行 if len(tds) < 3: continue # 提取页面顶部的主信息 main_info = {} main_table = soup.find('table') if main_table: string_iter = iter(main_table.stripped_strings) main_info = dict(zip(string_iter, string_iter)) # 拼接当前行的详情信息 row_data = main_info.copy() row_data.update({ 'link': tds[0].a.get('href') if tds[0].a else None, 'date': tds[-2].get_text(strip=True), 'type': tds[-1].get_text(strip=True) }) data.append(row_data) # 转换为DataFrame查看结果 df = pd.DataFrame(data) print(df)
修改说明
- 替换不兼容选择器:用
find_all('table')获取所有表格后取最后一个tables[-1],替代:last-of-type;通过判断行内td数量筛选有效行,替代:has(td)。 - 增加容错判断:避免表格不存在、行元素不足等场景触发报错。
- 优化文本提取:用
get_text(strip=True)自动清理文本中的多余空格和换行符。
内容的提问来源于stack exchange,提问作者João Pedro Rodrigues Oliveira
相关产品推荐
相关产品推荐

