You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取动态页面高亮文本失败及工具适配咨询

网页抓取问题:提取动态加载的高亮词失败

我需要抓取目标页面中的红色高亮词,这些高亮内容对应class为image-overlay hit-rect ng-star-inserted的div元素,目标是提取它们的title属性。预期得到一个长度为17的列表,示例如下:

EXPECTED_RESULT = ["Katri", "Katrina", "Katri", "Katri", "Katri", "Katri", "Katri", "Katri", "Ikonen.", "Katrina", "Katri", "Ikonen.", "Katri", "Katrina", "Katri", "Katri", "Katri"]

尝试的代码片段

from bs4 import BeautifulSoup
pg_snippet_highlighted_words = soup.find_all("div", class_="image-overlay hit-rect ng-star-inserted")
print(pg_snippet_highlighted_words) # 返回空列表: []
# 如果用soup.find()会触发错误: AttributeError: "'NoneType' object has no attribute 'get'"
# print(pg_snippet_highlighted_words.get("title"))

遇到的问题

运行代码后,pg_snippet_highlighted_words返回空列表[],无法获取到目标元素。

核心疑问

BeautifulSoup是否适合处理动态内容?

内容的提问来源于stack exchange,提问作者farid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 04:15:39