You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Playwright中获取特定div类(msg char-msg)的加载中文本

问题解决:Playwright获取嵌套DOM文本报错及优化方案

错误原因

你用了json_value()方法获取普通DOM节点的内容,这个方法是用来提取元素的JSON格式属性值(比如表单元素的value或dataset中的JSON数据),对普通文本节点使用就会抛出ref: <Node>错误。

正确解决方案

1. 替换文本获取方法

用text_content()(获取元素及其所有子元素的全部文本,包括隐藏内容)或inner_text()(仅获取可见文本)替代json_value()。

2. 用Playwright原生等待替代time.sleep()

固定等待time.sleep()既不高效也不可靠,改用Playwright的自动等待机制,确保元素渲染完成后再获取文本。

修改后的完整代码

from playwright.sync_api import Playwright, sync_playwright, expect

def run(playwright: Playwright) -> None:
    browser = playwright.firefox.launch(headless=False)
    context = browser.new_context()
    page = context.new_page()
    page.goto("https://url")
    
    # 处理初始同意按钮
    page.get_by_role("button", name="I understand").click()
    
    # 输入并发送消息
    page.get_by_placeholder("Type a message").fill("hello")
    page.locator("form").get_by_role("button").first.click()
    
    # 等待目标元素可见,确保文本已渲染完成
    char_msg_locator = page.locator('.char-msg')
    expect(char_msg_locator).to_be_visible(timeout=10000)
    
    # 获取文本(两种方式二选一)
    # 方式1:获取所有嵌套文本(包括隐藏内容)
    output_text = char_msg_locator.text_content()
    # 方式2:仅获取页面可见文本
    # output_text = char_msg_locator.inner_text()
    
    print(output_text.strip())  # 去除多余空格和换行
    
    context.close()
    browser.close()

with sync_playwright() as playwright:
    run(playwright)

更精准的文本定位(可选)

如果目标文本明确在<p>标签内,可以直接定位该元素,避免无关文本干扰:

# 等待<p>元素出现并获取文本
output_text = page.locator('.char-msg p').text_content()
print(output_text.strip())

关键说明

  • Playwright的expect断言会自动等待元素满足条件(比如可见、存在),无需手动设置固定等待时间。
  • text_content()和inner_text()的区别:前者获取元素的所有文本内容(包括CSS隐藏的部分),后者只返回页面上可见的文本。

内容的提问来源于stack exchange,提问作者Nizam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 09:34:51