如何在Playwright中获取特定div类(msg char-msg)的加载中文本
问题解决:Playwright获取嵌套DOM文本报错及优化方案
错误原因
你用了json_value()方法获取普通DOM节点的内容,这个方法是用来提取元素的JSON格式属性值(比如表单元素的value或dataset中的JSON数据),对普通文本节点使用就会抛出ref: <Node>错误。
正确解决方案
1. 替换文本获取方法
用text_content()(获取元素及其所有子元素的全部文本,包括隐藏内容)或inner_text()(仅获取可见文本)替代json_value()。
2. 用Playwright原生等待替代time.sleep()
固定等待time.sleep()既不高效也不可靠,改用Playwright的自动等待机制,确保元素渲染完成后再获取文本。
修改后的完整代码
from playwright.sync_api import Playwright, sync_playwright, expect def run(playwright: Playwright) -> None: browser = playwright.firefox.launch(headless=False) context = browser.new_context() page = context.new_page() page.goto("https://url") # 处理初始同意按钮 page.get_by_role("button", name="I understand").click() # 输入并发送消息 page.get_by_placeholder("Type a message").fill("hello") page.locator("form").get_by_role("button").first.click() # 等待目标元素可见,确保文本已渲染完成 char_msg_locator = page.locator('.char-msg') expect(char_msg_locator).to_be_visible(timeout=10000) # 获取文本(两种方式二选一) # 方式1:获取所有嵌套文本(包括隐藏内容) output_text = char_msg_locator.text_content() # 方式2:仅获取页面可见文本 # output_text = char_msg_locator.inner_text() print(output_text.strip()) # 去除多余空格和换行 context.close() browser.close() with sync_playwright() as playwright: run(playwright)
更精准的文本定位(可选)
如果目标文本明确在<p>标签内,可以直接定位该元素,避免无关文本干扰:
# 等待<p>元素出现并获取文本 output_text = page.locator('.char-msg p').text_content() print(output_text.strip())
关键说明
- Playwright的
expect断言会自动等待元素满足条件(比如可见、存在),无需手动设置固定等待时间。 text_content()和inner_text()的区别:前者获取元素的所有文本内容(包括CSS隐藏的部分),后者只返回页面上可见的文本。
内容的提问来源于stack exchange,提问作者Nizam
相关产品推荐
相关产品推荐

