You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多层循环中参数回传实现方案:将内层循环参数传回外层及初始循环

解决方案:多层嵌套循环的参数回传与断点续跑

Alright, let's tackle this problem. You need to update the outer loop's starting parameters (page_number_start and id_number) from the innermost exception block so your crawler can resume from the broken point. Here are two practical approaches tailored to your code:

核心思路

Python's nested loops can't directly modify outer loop iterators, but we can achieve breakpoint resumption by either using flag variables to pass state or encapsulating loop logic into a function that returns updated parameters. Both methods work for your crawler scenario—let's break them down:


方案1:使用标志位逐层跳出循环

This approach uses a global flag to mark when we need to restart the loop. When the breakpoint is triggered, we set the flag and break out of each nested loop layer, then reset the starting parameters in the outer loop to resume execution.

Modified Code Example

from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.keys import Keys
import time

id_number = 19
page_number_start = 4
auto_proxy_verification = 1
last_proxy_not_enough = 0
my_proxy = ['***************', '*****************', '**************', '**********', '****************', ]
proxy_number_start = 0

# Replace outer for-loop with while-loop for manual parameter control
while page_number_start < 999 and proxy_number_start < len(my_proxy):
    restart_flag = False  # Initialize restart flag
    page_number = page_number_start
    login_url = f'https://***********&amp;pageNumber={page_number}&amp;**********'
    
    # Proxy test loop
    for proxy_number in range(proxy_number_start, len(my_proxy)):
        try:
            driver.get(login_url)
            all_divs = driver.find_elements_by_class_name('****classname1****')
            locked_divs = driver.find_elements_by_xpath("****xpath1****")
            unlocked_divs = driver.find_elements_by_xpath("****xpath2****")
            
            if len(unlocked_divs) == 0: # Bad proxy condition 1
                print('proxy not active')
                proxy_number_start += 1
            elif len(all_divs) == 10: # Bad proxy condition 2
                print('proxy not enough')
                proxy_number_start += 1
                if proxy_number_start == len(my_proxy):
                    last_proxy_not_enough = 1
                    break
            else: # Good proxy found
                print('proxy ok')
                proxy_number_start += 1
                break
        except:
            print('proxy bad')
            proxy_number_start += 1
    
    main_window = driver.current_window_handle
    id_number_in = 0
    
    # Inner loop for page elements
    for div in all_divs:
        if last_proxy_not_enough == 1:
            break
        id_number_in += 1
        
        if id_number_in < id_number:
            continue
        else:
            id_number += 1
            if id_number >= len(all_divs):
                break
            
            detail_body_number = 'detail-body-%d' % id_number
            first_div_class = driver.find_element_by_xpath('//*[@id="{}"]/***/div[1]'.format(detail_body_number)).get_attribute('class')
            
            if first_div_class == '****certainclassname****': ## Valid element condition
                action = ActionChains(driver)
                action.key_down(Keys.CONTROL).click(title_href).key_up(Keys.CONTROL).perform()
                driver.switch_to.window(driver.window_handles[-1])
                
                try:
                    download_button = driver.find_element_by_xpath('****downloadbutton*******').click()
                    print('%d is unlocked, and downloaded' % id_number)
                    time.sleep(0.5)
                except:
                    try:
                        purchase_button_text = driver.find_element_by_xpath('****purchasebutton*******').text
                        if purchase_button_text == 'Purchase':
                            print('here purchase button!!!' % id_number)
                    except:
                        print('sometimes the proxy turns to be not active')
                        # Update parameters and trigger restart
                        page_number_start = page_number  # Resume from current page
                        id_number = id_number_in  # Resume from current element
                        restart_flag = True
                        break  # Exit element loop
                
                try:
                    driver.find_element_by_xpath('****downloadbutton*******').click()
                    time.sleep(2)
                except:
                    pass
            else:
                print('%d is locked, pass' % id_number)
                id_number = 0
        
        if restart_flag:
            break  # Exit element loop early
    
    if restart_flag:
        continue  # Go back to outer while-loop with updated parameters
    else:
        page_number_start += 1  # Move to next page if current page is fully processed

Key Changes

  • Replaced the outer for loop with a while loop to manually control starting parameters
  • Added a restart_flag to mark when we need to resume from the breakpoint
  • Each loop layer checks the flag and breaks out early if needed, returning to the outer loop to restart with updated parameters

方案2:封装为函数返回更新后的参数(推荐)

This method encapsulates the single-page processing logic into a function that returns updated parameters when a breakpoint is triggered. The main loop calls this function repeatedly, making the code more modular and readable.

Modified Code Example

from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.keys import Keys
import time

def process_page(page_number_start, id_number, proxy_number_start, my_proxy, driver):
    last_proxy_not_enough = 0
    page_number = page_number_start
    login_url = f'https://***********&amp;pageNumber={page_number}&amp;**********'
    
    # Proxy test loop
    for proxy_number in range(proxy_number_start, len(my_proxy)):
        try:
            driver.get(login_url)
            all_divs = driver.find_elements_by_class_name('****classname1****')
            locked_divs = driver.find_elements_by_xpath("****xpath1****")
            unlocked_divs = driver.find_elements_by_xpath("****xpath2****")
            
            if len(unlocked_divs) == 0: # Bad proxy condition 1
                print('proxy not active')
                proxy_number_start += 1
            elif len(all_divs) == 10: # Bad proxy condition 2
                print('proxy not enough')
                proxy_number_start += 1
                if proxy_number_start == len(my_proxy):
                    last_proxy_not_enough = 1
                    break
            else: # Good proxy found
                print('proxy ok')
                proxy_number_start += 1
                break
        except:
            print('proxy bad')
            proxy_number_start += 1
    
    if last_proxy_not_enough == 1:
        # All proxies exhausted, move to next page
        return page_number_start + 1, id_number, proxy_number_start
    
    main_window = driver.current_window_handle
    id_number_in = 0
    
    # Inner loop for page elements
    for div in all_divs:
        id_number_in += 1
        
        if id_number_in < id_number:
            continue
        else:
            id_number += 1
            if id_number >= len(all_divs):
                break
            
            detail_body_number = 'detail-body-%d' % id_number
            first_div_class = driver.find_element_by_xpath('//*[@id="{}"]/***/div[1]'.format(detail_body_number)).get_attribute('class')
            
            if first_div_class == '****certainclassname****': ## Valid element condition
                action = ActionChains(driver)
                action.key_down(Keys.CONTROL).click(title_href).key_up(Keys.CONTROL).perform()
                driver.switch_to.window(driver.window_handles[-1])
                
                try:
                    download_button = driver.find_element_by_xpath('****downloadbutton*******').click()
                    print('%d is unlocked, and downloaded' % id_number)
                    time.sleep(0.5)
                except:
                    try:
                        purchase_button_text = driver.find_element_by_xpath('****purchasebutton*******').text
                        if purchase_button_text == 'Purchase':
                            print('here purchase button!!!' % id_number)
                    except:
                        print('sometimes the proxy turns to be not active')
                        # Return updated parameters to resume from current page/element
                        return page_number_start, id_number_in, proxy_number_start
                
                try:
                    driver.find_element_by_xpath('****downloadbutton*******').click()
                    time.sleep(2)
                except:
                    pass
            else:
                print('%d is locked, pass' % id_number)
                id_number = 0
    
    # Current page fully processed, return parameters for next page
    return page_number_start + 1, 1, proxy_number_start

# Main execution logic
if __name__ == "__main__":
    id_number = 19
    page_number_start = 4
    proxy_number_start = 0
    my_proxy = ['***************', '*****************', '**************', '**********', '****************', ]
    
    # Initialize your driver here (add your own driver setup code)
    # driver = webdriver.Chrome(...)
    
    while page_number_start < 999 and proxy_number_start < len(my_proxy):
        page_number_start, id_number, proxy_number_start = process_page(
            page_number_start, id_number, proxy_number_start, my_proxy, driver
        )
    
    driver.quit()

Key Changes

  • Encapsulated single-page logic into process_page function, which receives current parameters and returns updated ones
  • When a breakpoint is triggered, the function returns the current page and element parameters, so the main loop can restart from that point
  • When a page is fully processed, the function returns parameters to move to the next page

Additional Tips

  • To persist breakpoints across program restarts, save page_number_start, id_number, and proxy_number_start to a local file (like JSON) and load them on startup
  • Ensure your driver object is properly passed between functions (in方案2, it's passed as a parameter)
  • Add error handling for edge cases like empty proxy pools or unexpected page structures

内容的提问来源于stack exchange,提问作者sampan0423

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:50:32