Python多层循环中参数回传实现方案:将内层循环参数传回外层及初始循环
Alright, let's tackle this problem. You need to update the outer loop's starting parameters (page_number_start and id_number) from the innermost exception block so your crawler can resume from the broken point. Here are two practical approaches tailored to your code:
核心思路
Python's nested loops can't directly modify outer loop iterators, but we can achieve breakpoint resumption by either using flag variables to pass state or encapsulating loop logic into a function that returns updated parameters. Both methods work for your crawler scenario—let's break them down:
方案1:使用标志位逐层跳出循环
This approach uses a global flag to mark when we need to restart the loop. When the breakpoint is triggered, we set the flag and break out of each nested loop layer, then reset the starting parameters in the outer loop to resume execution.
Modified Code Example
from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.common.keys import Keys import time id_number = 19 page_number_start = 4 auto_proxy_verification = 1 last_proxy_not_enough = 0 my_proxy = ['***************', '*****************', '**************', '**********', '****************', ] proxy_number_start = 0 # Replace outer for-loop with while-loop for manual parameter control while page_number_start < 999 and proxy_number_start < len(my_proxy): restart_flag = False # Initialize restart flag page_number = page_number_start login_url = f'https://***********&pageNumber={page_number}&**********' # Proxy test loop for proxy_number in range(proxy_number_start, len(my_proxy)): try: driver.get(login_url) all_divs = driver.find_elements_by_class_name('****classname1****') locked_divs = driver.find_elements_by_xpath("****xpath1****") unlocked_divs = driver.find_elements_by_xpath("****xpath2****") if len(unlocked_divs) == 0: # Bad proxy condition 1 print('proxy not active') proxy_number_start += 1 elif len(all_divs) == 10: # Bad proxy condition 2 print('proxy not enough') proxy_number_start += 1 if proxy_number_start == len(my_proxy): last_proxy_not_enough = 1 break else: # Good proxy found print('proxy ok') proxy_number_start += 1 break except: print('proxy bad') proxy_number_start += 1 main_window = driver.current_window_handle id_number_in = 0 # Inner loop for page elements for div in all_divs: if last_proxy_not_enough == 1: break id_number_in += 1 if id_number_in < id_number: continue else: id_number += 1 if id_number >= len(all_divs): break detail_body_number = 'detail-body-%d' % id_number first_div_class = driver.find_element_by_xpath('//*[@id="{}"]/***/div[1]'.format(detail_body_number)).get_attribute('class') if first_div_class == '****certainclassname****': ## Valid element condition action = ActionChains(driver) action.key_down(Keys.CONTROL).click(title_href).key_up(Keys.CONTROL).perform() driver.switch_to.window(driver.window_handles[-1]) try: download_button = driver.find_element_by_xpath('****downloadbutton*******').click() print('%d is unlocked, and downloaded' % id_number) time.sleep(0.5) except: try: purchase_button_text = driver.find_element_by_xpath('****purchasebutton*******').text if purchase_button_text == 'Purchase': print('here purchase button!!!' % id_number) except: print('sometimes the proxy turns to be not active') # Update parameters and trigger restart page_number_start = page_number # Resume from current page id_number = id_number_in # Resume from current element restart_flag = True break # Exit element loop try: driver.find_element_by_xpath('****downloadbutton*******').click() time.sleep(2) except: pass else: print('%d is locked, pass' % id_number) id_number = 0 if restart_flag: break # Exit element loop early if restart_flag: continue # Go back to outer while-loop with updated parameters else: page_number_start += 1 # Move to next page if current page is fully processed
Key Changes
- Replaced the outer
forloop with awhileloop to manually control starting parameters - Added a
restart_flagto mark when we need to resume from the breakpoint - Each loop layer checks the flag and breaks out early if needed, returning to the outer loop to restart with updated parameters
方案2:封装为函数返回更新后的参数(推荐)
This method encapsulates the single-page processing logic into a function that returns updated parameters when a breakpoint is triggered. The main loop calls this function repeatedly, making the code more modular and readable.
Modified Code Example
from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.common.keys import Keys import time def process_page(page_number_start, id_number, proxy_number_start, my_proxy, driver): last_proxy_not_enough = 0 page_number = page_number_start login_url = f'https://***********&pageNumber={page_number}&**********' # Proxy test loop for proxy_number in range(proxy_number_start, len(my_proxy)): try: driver.get(login_url) all_divs = driver.find_elements_by_class_name('****classname1****') locked_divs = driver.find_elements_by_xpath("****xpath1****") unlocked_divs = driver.find_elements_by_xpath("****xpath2****") if len(unlocked_divs) == 0: # Bad proxy condition 1 print('proxy not active') proxy_number_start += 1 elif len(all_divs) == 10: # Bad proxy condition 2 print('proxy not enough') proxy_number_start += 1 if proxy_number_start == len(my_proxy): last_proxy_not_enough = 1 break else: # Good proxy found print('proxy ok') proxy_number_start += 1 break except: print('proxy bad') proxy_number_start += 1 if last_proxy_not_enough == 1: # All proxies exhausted, move to next page return page_number_start + 1, id_number, proxy_number_start main_window = driver.current_window_handle id_number_in = 0 # Inner loop for page elements for div in all_divs: id_number_in += 1 if id_number_in < id_number: continue else: id_number += 1 if id_number >= len(all_divs): break detail_body_number = 'detail-body-%d' % id_number first_div_class = driver.find_element_by_xpath('//*[@id="{}"]/***/div[1]'.format(detail_body_number)).get_attribute('class') if first_div_class == '****certainclassname****': ## Valid element condition action = ActionChains(driver) action.key_down(Keys.CONTROL).click(title_href).key_up(Keys.CONTROL).perform() driver.switch_to.window(driver.window_handles[-1]) try: download_button = driver.find_element_by_xpath('****downloadbutton*******').click() print('%d is unlocked, and downloaded' % id_number) time.sleep(0.5) except: try: purchase_button_text = driver.find_element_by_xpath('****purchasebutton*******').text if purchase_button_text == 'Purchase': print('here purchase button!!!' % id_number) except: print('sometimes the proxy turns to be not active') # Return updated parameters to resume from current page/element return page_number_start, id_number_in, proxy_number_start try: driver.find_element_by_xpath('****downloadbutton*******').click() time.sleep(2) except: pass else: print('%d is locked, pass' % id_number) id_number = 0 # Current page fully processed, return parameters for next page return page_number_start + 1, 1, proxy_number_start # Main execution logic if __name__ == "__main__": id_number = 19 page_number_start = 4 proxy_number_start = 0 my_proxy = ['***************', '*****************', '**************', '**********', '****************', ] # Initialize your driver here (add your own driver setup code) # driver = webdriver.Chrome(...) while page_number_start < 999 and proxy_number_start < len(my_proxy): page_number_start, id_number, proxy_number_start = process_page( page_number_start, id_number, proxy_number_start, my_proxy, driver ) driver.quit()
Key Changes
- Encapsulated single-page logic into
process_pagefunction, which receives current parameters and returns updated ones - When a breakpoint is triggered, the function returns the current page and element parameters, so the main loop can restart from that point
- When a page is fully processed, the function returns parameters to move to the next page
Additional Tips
- To persist breakpoints across program restarts, save
page_number_start,id_number, andproxy_number_startto a local file (like JSON) and load them on startup - Ensure your
driverobject is properly passed between functions (in方案2, it's passed as a parameter) - Add error handling for edge cases like empty proxy pools or unexpected page structures
内容的提问来源于stack exchange,提问作者sampan0423

