如何用SoapUI和Selenium捕获框架中指定类的HTML代码并保存为TXT?
当然有可行的办法!根据你的需求,这里有几个不同场景下的方案,你可以按需选择:
方案1:浏览器开发者工具快速导出(适合临时手动操作)
如果只是偶尔需要抓取一次,用浏览器自带的工具最方便,步骤如下:
- 打开目标网页,按下
F12(或者右键页面空白处选择「检查」),打开开发者工具 - 在「Elements」面板的搜索框(按
Ctrl+F/Cmd+F调出)输入.productos-mant,定位到目标div标签 - 右键这个div标签,选择「Copy」→「Copy outerHTML」,这样就能复制该标签及其所有子元素的完整HTML代码
- 打开记事本(或任意文本编辑器),粘贴复制的内容,最后保存为
.txt文件即可
方案2:Python脚本自动化捕获(适合重复/批量操作)
如果需要定期抓取或者批量处理,用Python脚本更高效,分两种情况处理:
情况A:网页是静态渲染的(内容直接在HTML源码里)
需要用到requests和BeautifulSoup库,先安装依赖:
pip install requests beautifulsoup4
然后运行下面的脚本:
import requests from bs4 import BeautifulSoup # 替换成你要抓取的目标网页URL target_url = "https://www.supermercado.com/your-target-page" # 模拟浏览器请求头,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 获取网页内容 response = requests.get(target_url, headers=headers) response.encoding = response.apparent_encoding # 自动识别编码,避免乱码 # 解析HTML soup = BeautifulSoup(response.text, "html.parser") # 定位目标div productos_container = soup.find("div", class_="productos-mant") if productos_container: # 格式化HTML代码,更易读 formatted_html = productos_container.prettify() # 保存到txt文件 with open("productos_html.txt", "w", encoding="utf-8") as file: file.write(formatted_html) print("✅ HTML内容已成功保存到productos_html.txt") else: print("❌ 未找到目标<div class='productos-mant'>标签")
情况B:网页是动态渲染的(内容通过JavaScript加载)
如果目标内容是页面加载后通过JS动态生成的,requests无法获取到,这时候需要用Selenium模拟浏览器加载页面:
先安装依赖:
pip install selenium
还要下载对应浏览器的驱动(比如Chrome的chromedriver,版本要和你的浏览器一致),然后运行脚本:
from selenium import webdriver from selenium.webdriver.common.by import By # 初始化Chrome浏览器驱动(如果用其他浏览器,替换成对应的驱动,比如Firefox的geckodriver) driver = webdriver.Chrome() # 打开目标网页 driver.get("https://www.supermercado.com/your-target-page") # 等待页面加载完成(可根据实际情况调整等待时间,单位:秒) driver.implicitly_wait(10) try: # 定位目标div productos_container = driver.find_element(By.CLASS_NAME, "productos-mant") # 获取完整的HTML代码 html_content = productos_container.get_attribute("outerHTML") # 保存到文件 with open("productos_html.txt", "w", encoding="utf-8") as file: file.write(html_content) print("✅ HTML内容已成功保存到productos_html.txt") except Exception as e: print(f"❌ 抓取失败:{str(e)}") finally: # 关闭浏览器 driver.quit()
注意事项
- 用脚本抓取时,要遵守网站的
robots.txt规则,不要频繁发送请求,避免触发反爬机制 - 如果遇到登录后才能访问的页面,脚本需要额外处理登录逻辑(比如模拟输入账号密码)
内容的提问来源于stack exchange,提问作者FC_J
相关产品推荐
相关产品推荐

