You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup与Selenium进行网页抓取时find_all返回空列表问题求助

Troubleshooting Your Empty Beautiful Soup Result

Hey there! Let's dig into why your code is returning an empty list when trying to scrape those elements. I've gone through your snippet and thought through common issues with dynamic sites like this—here's how to fix it:

1. Double-check your class names first

First things first: make sure the class names you're targeting (companyName, presentation, etc.) are exactly what's on the page. Sometimes web devs use compound classes (with spaces) or dynamic class names that change slightly.

Open the target page in your browser, hit F12 to open DevTools, use the element picker to select a company name, and copy the exact class attribute from the HTML. It's easy to miss a typo or a hidden prefix here!

2. Replace fixed sleep with explicit waits

Your sleep(randint(2,10)) is a good start to avoid blocking, but fixed waits aren't reliable—sometimes the page takes longer to load dynamic content, or AJAX calls finish after your sleep ends.

Instead, use Selenium's explicit waits to wait until the elements you need are actually present on the page before grabbing the source. Here's how to adjust your code:

from bs4 import BeautifulSoup
import numpy as np
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

page = "https://www.acheteralasource.com/producteurs-en-france/all/departement/75/page/1"
driver = webdriver.Chrome()
driver.get(page)

# Wait up to 10 seconds for at least one company name element to appear
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CLASS_NAME, "companyName")))

# Now grab the page source—your target elements should be loaded
soup = BeautifulSoup(driver.page_source, 'html.parser')
my_table = soup.find_all(class_=['companyName', 'presentation','addressCity','addressPostalCode'])

# Test if it works
print([elem.text.strip() for elem in my_table])

3. Try targeting parent containers first

If the direct class search still fails, try narrowing down to the parent elements that hold all the producer info first. For example, if each producer is in a card div, you can scrape each card individually and then extract the details from inside:

# After getting the soup with explicit wait
producer_cards = soup.find_all('div', class_='producer-card') # Replace with actual parent class from DevTools

for card in producer_cards:
    # Extract each piece of info with fallbacks in case an element is missing
    company = card.find(class_='companyName').text.strip() if card.find(class_='companyName') else "No name"
    bio = card.find(class_='presentation').text.strip() if card.find(class_='presentation') else "No bio"
    city = card.find(class_='addressCity').text.strip() if card.find(class_='addressCity') else "No city"
    postal = card.find(class_='addressPostalCode').text.strip() if card.find(class_='addressPostalCode') else "No postal code"
    
    print(f"Company: {company}\nBio: {bio[:50]}...\nLocation: {postal}, {city}\n---")

4. Check for iframes (a common gotcha)

Occasionally, dynamic content lives inside an <iframe> tag. If that's the case, your driver won't see the elements until you switch to the iframe first. You can check for this in DevTools, and if found, add this line before waiting for elements:

# Replace with the actual iframe ID or index if needed
driver.switch_to.frame(driver.find_element(By.TAG_NAME, "iframe"))

Start with steps 1 and 2—those are the most likely fixes for your empty list issue. Once you confirm the elements are loading correctly, you can refine how you extract and organize the data.

内容的提问来源于stack exchange,提问作者user13998985

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 15:12:52