Python中如何解决BeautifulSoup的GuessedAtParserWarning警告问题
问题修复方法
- 首先处理警告:按照提示给
BeautifulSoup构造函数补充features参数指定解析器即可,推荐直接用Python内置的html.parser,无需额外安装依赖。 - 其次修正网页抓取逻辑:你现在直接把网址字符串
"google.com"传给BeautifulSoup是无效的,它只能解析HTML源码,你需要先通过网络请求拿到对应网页的内容再传入解析。
完整修正代码
import pyttsx3 import speech_recognition as sr import speech_recognition from bs4 import BeautifulSoup import requests # 新增网络请求依赖 r = sr.Recognizer() def speak_text(command): engine = pyttsx3.init() engine.say(command) engine.runAndWait() source2: speech_recognition.Microphone with sr.Microphone() as source2: r.adjust_for_ambient_noise(source2, duration=0.2) print("Listening...") audio2 = r.listen(source2) assert isinstance(audio2, object) my_text = r.recognize_google(audio2) print("Recognizing...") if my_text == "test": print("SECRET") else: # 先请求网页获取HTML内容,注意网址要写完整带http/https resp = requests.get("https://google.com") # 补充features参数指定解析器,消除警告 soup = BeautifulSoup(resp.text, features="html.parser") for link in soup.find_all("a"): print(link.get("href"))
如果你更习惯用lxml或者html5lib作为解析器,只需要先通过pip install lxml或者pip install html5lib安装对应依赖,再把features参数的值改成对应的解析器名称即可。
内容的提问来源于stack exchange,提问作者Blobby YT
相关产品推荐
相关产品推荐

