网页爬虫新手求助:BeautifulSoup无返回结果,更换解析器仍无效
解决BeautifulSoup无法获取结果的问题
嘿,作为web scraping新手碰到这种获取不到结果的情况确实挺闹心的,换了解析器还没解决?咱们来捋捋可能的问题点:
显式指定解析器,同时解决警告问题
你当前的代码里bsObj = BeautifulSoup(html.read())没有明确指定解析器,虽然你说换过解析器,但这里显式声明会更稳妥,还能消除那个警告。修改后的代码可以这样写:from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen("http://www.pythonscraping.com/pages/page1.html") bsObj = BeautifulSoup(html.read(), 'html.parser') # 明确指定解析器试试更规范的资源管理方式
用with语句处理urlopen的返回对象,既能避免资源泄漏,代码也更健壮:from urllib.request import urlopen from bs4 import BeautifulSoup with urlopen("http://www.pythonscraping.com/pages/page1.html") as html: bsObj = BeautifulSoup(html, 'html.parser') # 直接传入html对象,无需调用read()别忘了提取具体内容
你现在的代码只是创建了BeautifulSoup对象,但没尝试提取页面元素哦!比如想要获取页面的h1标签内容,可以加一行:print(bsObj.h1) # 打印页面中的h1标签内容这样就能看到具体的结果了。
如果还是有问题,可以检查下网络连接是否正常,或者页面能否正常打开——毕竟有时候网络波动也会导致获取不到页面内容~
内容的提问来源于stack exchange,提问作者Ethic
相关产品推荐
相关产品推荐

