使用BeautifulSoup提取动态页面高亮文本失败及工具适配咨询
网页抓取问题:提取动态加载的高亮词失败
我需要抓取目标页面中的红色高亮词,这些高亮内容对应class为image-overlay hit-rect ng-star-inserted的div元素,目标是提取它们的title属性。预期得到一个长度为17的列表,示例如下:
EXPECTED_RESULT = ["Katri", "Katrina", "Katri", "Katri", "Katri", "Katri", "Katri", "Katri", "Ikonen.", "Katrina", "Katri", "Ikonen.", "Katri", "Katrina", "Katri", "Katri", "Katri"]
尝试的代码片段
from bs4 import BeautifulSoup pg_snippet_highlighted_words = soup.find_all("div", class_="image-overlay hit-rect ng-star-inserted") print(pg_snippet_highlighted_words) # 返回空列表: [] # 如果用soup.find()会触发错误: AttributeError: "'NoneType' object has no attribute 'get'" # print(pg_snippet_highlighted_words.get("title"))
遇到的问题
运行代码后,pg_snippet_highlighted_words返回空列表[],无法获取到目标元素。
核心疑问
BeautifulSoup是否适合处理动态内容?
内容的提问来源于stack exchange,提问作者farid
相关产品推荐
相关产品推荐

