You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在爬虫脚本执行前运行指定单元测试并获取结果标识

嘿,我来帮你实现这个需求!你想要在爬虫执行前先跑指定单元测试,只判断通过/失败而不需要详细输出,对吧?咱们可以借助Python的unittest框架来搞定——直接调用test_1()方法是行不通的,得用测试运行器来执行并捕获结果,具体实现如下:

1. 完善你的单元测试(tests.py)

首先确保你的测试用例能正确验证XPath有效性,比如模拟目标页面的HTML结构(或者直接请求真实页面),用断言来检查XPath是否能找到预期元素:

import unittest
from lxml import etree  # 或者用BeautifulSoup,根据你的爬虫解析方式选择

class CrwTst(unittest.TestCase):
    def test_1(self):
        # 这里可以替换成真实请求目标页面的代码,比如用requests.get()获取HTML
        mock_html = """
        <div class="target-container">
            <h2 class="title">Expected Title</h2>
            <ul class="items">
                <li>Item 1</li>
                <li>Item 2</li>
            </ul>
        </div>
        """
        tree = etree.HTML(mock_html)
        
        # 验证你的核心XPath是否可用
        title_result = tree.xpath('//h2[@class="title"]/text()')
        self.assertEqual(len(title_result), 1, "XPath failed to locate title element")
        self.assertEqual(title_result[0], "Expected Title", "Title text does not match expected value")
        
        items_result = tree.xpath('//ul[@class="items"]/li/text()')
        self.assertEqual(len(items_result), 2, "XPath failed to locate all list items")

2. 在爬虫脚本中添加测试执行逻辑(crawler.py)

编写一个辅助函数来运行指定的测试用例,捕获通过/失败状态,然后根据结果决定是否执行爬虫:

import unittest
from tests import CrwTst  # 导入你的测试类

class Crawler(object):
    @staticmethod
    def action_1():
        # 这里是你的核心爬虫逻辑
        print("Starting crawler action_1...")
        # 示例:爬取页面、解析数据、存储结果等操作
        # ...

def run_target_test(test_class, test_method):
    # 创建测试套件,添加指定的测试方法
    test_suite = unittest.TestSuite()
    test_suite.addTest(test_class(test_method))
    
    # 创建测试结果对象,用于捕获测试状态
    test_result = unittest.TestResult()
    # 运行测试
    test_suite.run(test_result)
    
    # 检查是否有失败或错误,返回布尔值:True=通过,False=失败/出错
    return len(test_result.failures) == 0 and len(test_result.errors) == 0

if __name__ == "__main__":
    # 运行指定的test_1测试
    is_test_passed = run_target_test(CrwTst, "test_1")
    
    if is_test_passed:
        print("✅ Tests passed, executing crawler...")
        Crawler.action_1()
    else:
        print("❌ Tests failed, aborting crawler execution.")

关键说明

  • run_target_test函数负责加载并运行指定的测试方法,通过检查TestResult中的failures和errors列表长度,判断测试是否全部通过。
  • 这种方式不会输出unittest默认的测试报告,只会返回一个布尔值的通过/失败标识,完全符合你的需求。
  • 如果需要运行多个测试方法,可以扩展run_target_test函数,支持传入方法名列表,批量运行后再判断是否全部通过。
  • 若测试需要真实页面数据,只需把mock_html替换为requests.get(target_url).text即可,让测试更贴近实际场景。

内容的提问来源于stack exchange,提问作者Panagiotis Simakis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:18:50