Python中BeautifulSoup的find_all返回空列表问题咨询
问题原因及解决办法
可能的原因1:文本节点不匹配精确字符串
目标页面里的"Marino"可能被HTML标签拆分,或者包含空格、换行、特殊字符(比如Marino 、\nMarino),而find_all(text="Marino")是精确匹配完整文本节点内容,所以找不到。
解决办法:
- 使用正则表达式模糊匹配:
import re marino = doc.find_all(text=re.compile(r"Marino", re.IGNORECASE)) # 可选忽略大小写 - 使用lambda表达式判断文本是否包含目标字符串:
marino = doc.find_all(text=lambda x: x and "Marino" in x.strip())
可能的原因2:内容是动态加载的
requests.get()只能获取服务器返回的静态HTML,如果"Marino"是通过JavaScript动态渲染(比如从接口拉取数据后生成),静态页面里根本没有这个文本,自然找不到。
解决办法:
- 使用Selenium/Playwright这类工具模拟浏览器加载页面,获取渲染后的完整HTML:
示例(Selenium):from selenium import webdriver from bs4 import BeautifulSoup driver = webdriver.Chrome() driver.get("https://recreation.northeastern.edu/") doc = BeautifulSoup(driver.page_source, "html.parser") marino = doc.find_all(text=lambda x: x and "Marino" in x.strip()) print(marino) driver.quit()
可能的原因3:网站反爬导致返回内容异常
目标网站可能检测到请求是爬虫,返回了伪装的页面或者403响应,此时result.text里没有目标内容。
解决办法:
- 检查响应状态码:
print(result.status_code),如果是403/401,需要添加请求头模拟浏览器:headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } result = requests.get(url, headers=headers)
内容的提问来源于stack exchange,提问作者akanes
相关产品推荐
相关产品推荐

