使用BeautifulSoup爬虫时Python循环报错问题求助
Hey there! Let's break down why your loop is throwing that 'NoneType' object is not subscriptable error and fix it up.
What's Causing the Error?
The problem happens because not every product container has the exact same HTML structure. When you run con.findAll("div",{"class":"item-branding"}), if a product doesn't have that item-branding div (or the div doesn't contain an img with a title), titleco becomes an empty list. Trying to access titleco[0] gives you None, and then trying to get .img["title"] from None triggers the error.
How to Fix It
You need to add checks to make sure elements exist before trying to access their properties. Here's updated code with safety checks:
containers = pagesoup.findAll("div",{"class":"item-container"}) for con in containers: # Get product title with safety check title = con.img["title"] if (con.img and "title" in con.img.attrs) else "No Title Available" # Get brand with safety check titleco = con.findAll("div",{"class":"item-branding"}) if titleco: brand_elem = titleco[0].img brand = brand_elem["title"] if (brand_elem and "title" in brand_elem.attrs) else "No Brand Listed" else: brand = "No Brand Listed" # Get price with safety check priceco = con.findAll("li",{"class":"price-current"}) price = priceco[0].text.strip() if priceco else "No Price Found" # Print or store your data print(f"Title: {title}\nBrand: {brand}\nPrice: {price}\n---")
Key Takeaways
Web pages often have inconsistent structures—some products might skip certain elements (like brand logos) for promotions or special listings. Always add existence checks for elements and their attributes before accessing them to avoid breaking your loop.
内容的提问来源于stack exchange,提问作者hamza mhadhbi

