BeautifulSoup find_all函数在main函数内无法正常工作的问题排查求助
soup.find_all() Returns Empty in Main Function but Works Line-by-Line This is a super common issue with dynamic sites when using Selenium—let's break down what's happening and fix it.
The Root Cause
When you run code line-by-line, you're giving the browser enough time to fully load the page and render all the dynamic content (like Conforama's product search results). But in your main function, you're immediately grabbing driver.page_source right after driver.get(url.format(page))—the page hasn't finished loading yet, so the ais-Hits-item box-product fragItem elements don't exist in the HTML you're parsing with BeautifulSoup.
Fixes to Try (Ordered by Recommendation)
1. Use Explicit Waits (Best Practice)
Explicit waits tell Selenium to wait until a specific condition is met (like the presence of your target elements) before proceeding. This is way more reliable than arbitrary sleep times.
Modify your main function to add this:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def main(search_term): driver = webdriver.Chrome(ChromeDriverManager().install()) # Set up explicit wait (wait up to 10 seconds for elements to appear) wait = WebDriverWait(driver, 10) records = [] url = get_url(search_term) somme = 0 for page in range (1,4): driver.get(url.format(page)) # Wait until at least one product item is present on the page wait.until( EC.presence_of_element_located((By.CLASS_NAME, "ais-Hits-item")) ) # Now parse the fully loaded page source soup = BeautifulSoup(driver.page_source, 'html.parser') results = soup.find_all('li', {'class' : 'ais-Hits-item box-product fragItem'}) print(len(results)) somme+=len(results) for result in results: record = extract_record(result) if record: print(record) records.append(record) driver.close() print('somme',somme)
2. Add Implicit Wait
If you prefer a simpler (though less precise) approach, set an implicit wait once when initializing the driver. This tells Selenium to wait up to X seconds when trying to locate any element:
def main(search_term): driver = webdriver.Chrome(ChromeDriverManager().install()) # Wait up to 10 seconds for elements to load whenever searching driver.implicitly_wait(10) records = [] url = get_url(search_term) somme = 0 for page in range (1,4): driver.get(url.format(page)) soup = BeautifulSoup(driver.page_source, 'html.parser') results = soup.find_all('li', {'class' : 'ais-Hits-item box-product fragItem'}) # ... rest of your existing code
3. Temporary Fix: Force Sleep (Not Recommended Long-Term)
As a quick test, you can add a time.sleep() to give the page time to load. This is brittle though—if the page takes longer than your sleep time to load, it'll still fail:
import time def main(search_term): driver = webdriver.Chrome(ChromeDriverManager().install()) records = [] url = get_url(search_term) somme = 0 for page in range (1,4): driver.get(url.format(page)) # Wait 3 seconds for dynamic content to render time.sleep(3) soup = BeautifulSoup(driver.page_source, 'html.parser') results = soup.find_all('li', {'class' : 'ais-Hits-item box-product fragItem'}) # ... rest of your existing code
Quick Additional Checks
- Double-check that the pagination URL is correct: print
url.format(page)to verify it generates valid page 1/2/3 links. - Confirm the class name
ais-Hits-item box-product fragItemdoesn't change between pages (your line-by-line test for page 1 works, so this is unlikely).
内容的提问来源于stack exchange,提问作者shelltief

