Python BeautifulSoup4爬虫findAll()解析失败问题求助
Hey there, let's walk through why your web scraper isn't picking up the content you want. Here are the most common fixes for your scenario:
1. Fix the Incomplete Class Selector
First off, your code has {'class': 'c...'} — that truncated class name is definitely not matching anything! You need to use the full, exact class attribute from the page.
To get it:
- Open the target URL in your browser, hit F12 to open DevTools
- Use the element picker (the arrow icon) to select the
<main>tag you're targeting - Copy the full value of its
classattribute (it might look something likec-main-content c-layout)
Then update your code to use the full class. Note that in BeautifulSoup, class is a reserved keyword, so you can use the class_ parameter for clarity:
# Using class_ parameter for eachItem in soup.findAll('main', class_='full-class-name-here'): # Your code to extract product names here
Or use a CSS selector, which is often more flexible:
# Using CSS selector syntax for eachItem in soup.select('main.full-class-name-here'): # Your code to extract product names here
2. Check if Content is Dynamically Loaded
Many retail sites use JavaScript to load products after the initial page load. The requests.get() method only fetches the static HTML source, so if the products are added via JS, they won't show up in your plain_text variable.
If this is the case, you'll need to use a tool that simulates a browser to render the JS. Selenium is a popular option:
from selenium import webdriver from bs4 import BeautifulSoup def louis_spider(max_pages): page = 0 # Initialize the browser (make sure you have the ChromeDriver installed) driver = webdriver.Chrome() while page <= max_pages: url = f'https://us.testcompany.com/eng-us/women/hanbags/_/N-r4xtxc/to-{page}' driver.get(url) # Wait for JS to load (you might need explicit waits here for reliability) plain_text = driver.page_source soup = BeautifulSoup(plain_text, 'html.parser') # Now search for your product elements for product in soup.select('your-product-selector-here'): product_name = product.get_text(strip=True) print(product_name) page += 1 driver.quit()
3. Add Request Headers to Avoid Blocking
Some sites block requests that don't have a proper User-Agent header (since requests uses a default one that's easy to spot as a bot). Add headers to mimic a real browser:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } source_code = requests.get(url, headers=headers)
4. Verify You're Targeting the Right Element
Wait a second — are you sure the product names are inside the <main> tag? Often, product lists are nested in <div>s, <ul>s, or <article>s with classes like product-grid or item-list. Double-check in DevTools: find a product name, then trace up its parent elements to find the correct container to target with BeautifulSoup.
For example, if each product name is in an <h2> with class product-title, you'd skip the <main> tag and go straight for:
for product_title in soup.findAll('h2', class_='product-title'): print(product_title.get_text(strip=True))
内容的提问来源于stack exchange,提问作者OrangeOwner

