使用BeautifulSoup爬虫调用CSS选择器时报AttributeError如何解决
错误本质
AttributeError: 'NoneType' object has no attribute 'text'的触发原因是select_one方法未匹配到符合规则的元素,返回了空值None,此时对None调用text属性就会抛出异常。
常见触发原因
- 目标内容为JavaScript动态渲染:你在浏览器开发者工具中看到的DOM结构是页面加载完成后JS渲染生成的,
requests库仅能获取初始静态HTML源码,源码中不存在对应元素自然无法匹配。 - CSS选择器稳定性差:你使用的选择器层级过长,且依赖
nth-child这种位置匹配规则,只要页面DOM结构有细微调整就会匹配失败,同时BeautifulSoup对部分CSS语法的支持和浏览器存在差异也会导致匹配失效。 - 解析器兼容问题:默认的
html.parser对部分不规范HTML页面的解析效果较差,会导致解析生成的DOM结构和实际不一致,匹配失败。
解决方案
第一步:验证目标内容是否存在于静态HTML
运行如下代码确认内容加载方式:
import requests url = "https://www.cryptocompare.com/coins/bnb/influence/USDT" resp = requests.get(url) resp.encoding = 'utf-8' if "We don't have any code repository data yet" in resp.text: print("静态页面存在目标内容,仅需调整选择器即可") else: print("目标内容为动态渲染,需要使用浏览器渲染工具获取页面")
情况1:静态页面存在目标内容
优化选择器、调整解析器即可:
- 更换解析器为兼容性更强的
lxml,需先执行安装命令pip install lxml - 简化选择器规则,优先用稳定的属性匹配,避免过长层级和位置匹配规则
- 匹配完成先判空再调用
text属性,规避报错
修正后代码示例:
from bs4 import BeautifulSoup import requests html = requests.get("https://www.cryptocompare.com/coins/bnb/influence/USDT").text # 更换为lxml解析器 soup = BeautifulSoup(html, 'lxml') # 先匹配元素再判空 ele = soup.select_one("#col-body div.social-influence div.panel-inactive div:nth-child(3) h4") total_commit = ele.text.strip() if ele else "未匹配到对应元素" print(total_commit)
情况2:目标内容为动态渲染
需要使用Selenium、Playwright等工具模拟浏览器运行,获取JS渲染完成后的DOM结构再解析,Selenium示例如下:
首先安装依赖:pip install selenium webdriver-manager
代码示例:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager # 启动无头浏览器 options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get("https://www.cryptocompare.com/coins/bnb/influence/USDT") # 等待页面渲染完成 driver.implicitly_wait(10) # 获取渲染后的完整HTML html = driver.page_source driver.quit() soup = BeautifulSoup(html, 'lxml') ele = soup.select_one("#col-body > div > social-influence > div.row.row-zero.influence-others.panel-inactive > div:nth-child(3) > h4") total_commit = ele.text.strip() if ele else "未匹配到对应元素" print(total_commit)
内容的提问来源于stack exchange,提问作者Siddharth Tiwari
相关产品推荐
相关产品推荐

