You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python单元测试:如何优雅地Mock网页爬虫函数?

Great question! Your generator approach works perfectly fine, but there are definitely more concise and Pythonic ways to handle this scenario where you need a mock to return different values on successive calls. Let’s break down a few cleaner options:

1. Directly use side_effect with a list (simplest solution)

The side_effect parameter in unittest.mock.patch accepts iterables like lists. Each time your mocked function is called, it will return the next element from the list automatically. This eliminates the need for separate generator and wrapper functions entirely:

def test_mockscrapper(self):
    # Preload your two response contents into a list
    mock_responses = [
        open("search_content.dat", "rb").read(),
        open("feed_content.dat", "rb").read()
    ]
    
    with patch("http.client.HTTPResponse.read", side_effect=mock_responses):
        # Run your test logic here
        # First call to read() returns search_content, second returns feed_content
        self.your_test_workflow()

This is the most straightforward approach—clean, readable, and exactly tailored to your two-call use case. If you ever need more sequential responses, just add more items to the list.

2. Use a generator expression for compactness

If you prefer a more concise one-liner for the responses, you can pass a generator expression directly to side_effect:

def test_mockscrapper(self):
    # Create an iterator that yields your two content files
    response_generator = (
        open(f"{type}_content.dat", "rb").read() 
        for type in ["search", "feed"]
    )
    
    with patch("http.client.HTTPResponse.read", side_effect=response_generator):
        self.your_test_workflow()

This keeps things tight and avoids explicitly building a list, which is nice if you want to generate responses dynamically (like looping through a list of filenames).

3. Closures for more complex logic (if needed later)

If you ever need to add extra logic to your mock (like validating arguments, logging calls, or dynamic response generation), a closure is cleaner than a standalone generator function:

def test_mockscrapper(self):
    def create_mock_reader():
        # Maintain state inside the closure
        responses = iter([
            open("search_content.dat", "rb").read(),
            open("feed_content.dat", "rb").read()
        ])
        
        def mock_read(*args, **kwargs):
            # Add any extra checks here if needed
            # e.g., assert that args match expected values
            return next(responses)
        
        return mock_read

    with patch("http.client.HTTPResponse.read", side_effect=create_mock_reader()):
        self.your_test_workflow()

This gives you flexibility without the clunk of a separate generator function, and keeps all mock-related logic contained within the test method.

For your current use case, option 1 is the most Pythonic choice—it’s simple, readable, and gets the job done with minimal boilerplate.

内容的提问来源于stack exchange,提问作者Learning is a mess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:41:28