You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup4爬虫findAll()解析失败问题求助

Troubleshooting Your BeautifulSoup findAll() Issue

Hey there, let's walk through why your web scraper isn't picking up the content you want. Here are the most common fixes for your scenario:

1. Fix the Incomplete Class Selector

First off, your code has {'class': 'c...'} — that truncated class name is definitely not matching anything! You need to use the full, exact class attribute from the page.

To get it:

  • Open the target URL in your browser, hit F12 to open DevTools
  • Use the element picker (the arrow icon) to select the <main> tag you're targeting
  • Copy the full value of its class attribute (it might look something like c-main-content c-layout)

Then update your code to use the full class. Note that in BeautifulSoup, class is a reserved keyword, so you can use the class_ parameter for clarity:

# Using class_ parameter
for eachItem in soup.findAll('main', class_='full-class-name-here'):
    # Your code to extract product names here

Or use a CSS selector, which is often more flexible:

# Using CSS selector syntax
for eachItem in soup.select('main.full-class-name-here'):
    # Your code to extract product names here

2. Check if Content is Dynamically Loaded

Many retail sites use JavaScript to load products after the initial page load. The requests.get() method only fetches the static HTML source, so if the products are added via JS, they won't show up in your plain_text variable.

If this is the case, you'll need to use a tool that simulates a browser to render the JS. Selenium is a popular option:

from selenium import webdriver
from bs4 import BeautifulSoup

def louis_spider(max_pages):
    page = 0
    # Initialize the browser (make sure you have the ChromeDriver installed)
    driver = webdriver.Chrome()
    while page <= max_pages:
        url = f'https://us.testcompany.com/eng-us/women/hanbags/_/N-r4xtxc/to-{page}'
        driver.get(url)
        # Wait for JS to load (you might need explicit waits here for reliability)
        plain_text = driver.page_source
        soup = BeautifulSoup(plain_text, 'html.parser')
        # Now search for your product elements
        for product in soup.select('your-product-selector-here'):
            product_name = product.get_text(strip=True)
            print(product_name)
        page += 1
    driver.quit()

3. Add Request Headers to Avoid Blocking

Some sites block requests that don't have a proper User-Agent header (since requests uses a default one that's easy to spot as a bot). Add headers to mimic a real browser:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

source_code = requests.get(url, headers=headers)

4. Verify You're Targeting the Right Element

Wait a second — are you sure the product names are inside the <main> tag? Often, product lists are nested in <div>s, <ul>s, or <article>s with classes like product-grid or item-list. Double-check in DevTools: find a product name, then trace up its parent elements to find the correct container to target with BeautifulSoup.

For example, if each product name is in an <h2> with class product-title, you'd skip the <main> tag and go straight for:

for product_title in soup.findAll('h2', class_='product-title'):
    print(product_title.get_text(strip=True))

内容的提问来源于stack exchange,提问作者OrangeOwner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:28:51