Python网页爬虫报错求助:爬取新蛋网GPU信息时遇两类错误
Hey Ethan, let's tackle that frustrating error you're hitting. The TypeError: 'NoneType' object is not subscriptable on this line:
brand = container.div.div.a.img["title"]
means one of the elements in that chain (container.div, container.div.div, container.div.div.a, or container.div.div.a.img) is returning None—so when you try to access ["title"] on a None value, Python throws an error. This usually happens because Newegg's page structure has changed, or some GPU listings don't follow the exact DOM hierarchy you're expecting.
Here are a few actionable fixes to make your scraper more robust:
1. Add Defensive Checks for Each Element
Instead of chaining all the div/a/img calls together, break them down and verify each step exists before moving on:
# Step through each element with safety checks top_div = container.find('div') if top_div: inner_div = top_div.find('div') if inner_div: brand_link = inner_div.find('a') if brand_link: brand_img = brand_link.find('img') if brand_img and 'title' in brand_img.attrs: brand = brand_img['title'] else: brand = 'Unknown Brand' else: brand = 'Unknown Brand' else: brand = 'Unknown Brand' else: brand = 'Unknown Brand'
This way, if any layer of the DOM is missing, your program will just default to a placeholder value instead of crashing.
2. Wrap the Call in a Try-Except Block
For a cleaner approach, use exception handling to catch errors when elements are missing:
try: brand = container.div.div.a.img["title"] except (AttributeError, TypeError): # Catch cases where a div/a/img doesn't exist, or we hit a None value brand = 'Unknown Brand'
This is concise and handles all scenarios where the element chain breaks.
3. Use Precise CSS Selectors (More Reliable Than Chaining Divs)
Newegg's product listings likely use consistent class names for elements like brand logos. Instead of relying on nested divs, target the specific element directly with a CSS selector. For example, if the brand image has a class like product-brand-img, you could do:
brand_img = container.select_one('img.product-brand-img') brand = brand_img.get('title', 'Unknown Brand') if brand_img else 'Unknown Brand'
select_one() is more flexible and less prone to breaking if the site adds/removes a nested div.
Bonus Tips for Web Scraping
- Always double-check the page structure using your browser's DevTools (right-click > Inspect) to confirm your selectors match the current HTML.
- Add small delays between requests (using
time.sleep()) to avoid getting blocked by Newegg's anti-scraping measures. - Once you fix this error, implementing spreadsheet functionality will be straightforward—libraries like
pandasoropenpyxlmake it easy to export your GPU data to Excel files.
内容的提问来源于stack exchange,提问作者Ethan Price

