You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫报错求助:爬取新蛋网GPU信息时遇两类错误

Fixing the 'NoneType' Error in Your Newegg GPU Scraper

Hey Ethan, let's tackle that frustrating error you're hitting. The TypeError: 'NoneType' object is not subscriptable on this line:

brand = container.div.div.a.img["title"]

means one of the elements in that chain (container.div, container.div.div, container.div.div.a, or container.div.div.a.img) is returning None—so when you try to access ["title"] on a None value, Python throws an error. This usually happens because Newegg's page structure has changed, or some GPU listings don't follow the exact DOM hierarchy you're expecting.

Here are a few actionable fixes to make your scraper more robust:

1. Add Defensive Checks for Each Element

Instead of chaining all the div/a/img calls together, break them down and verify each step exists before moving on:

# Step through each element with safety checks
top_div = container.find('div')
if top_div:
    inner_div = top_div.find('div')
    if inner_div:
        brand_link = inner_div.find('a')
        if brand_link:
            brand_img = brand_link.find('img')
            if brand_img and 'title' in brand_img.attrs:
                brand = brand_img['title']
            else:
                brand = 'Unknown Brand'
        else:
            brand = 'Unknown Brand'
    else:
        brand = 'Unknown Brand'
else:
    brand = 'Unknown Brand'

This way, if any layer of the DOM is missing, your program will just default to a placeholder value instead of crashing.

2. Wrap the Call in a Try-Except Block

For a cleaner approach, use exception handling to catch errors when elements are missing:

try:
    brand = container.div.div.a.img["title"]
except (AttributeError, TypeError):
    # Catch cases where a div/a/img doesn't exist, or we hit a None value
    brand = 'Unknown Brand'

This is concise and handles all scenarios where the element chain breaks.

3. Use Precise CSS Selectors (More Reliable Than Chaining Divs)

Newegg's product listings likely use consistent class names for elements like brand logos. Instead of relying on nested divs, target the specific element directly with a CSS selector. For example, if the brand image has a class like product-brand-img, you could do:

brand_img = container.select_one('img.product-brand-img')
brand = brand_img.get('title', 'Unknown Brand') if brand_img else 'Unknown Brand'

select_one() is more flexible and less prone to breaking if the site adds/removes a nested div.

Bonus Tips for Web Scraping

  • Always double-check the page structure using your browser's DevTools (right-click > Inspect) to confirm your selectors match the current HTML.
  • Add small delays between requests (using time.sleep()) to avoid getting blocked by Newegg's anti-scraping measures.
  • Once you fix this error, implementing spreadsheet functionality will be straightforward—libraries like pandas or openpyxl make it easy to export your GPU data to Excel files.

内容的提问来源于stack exchange,提问作者Ethan Price

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:27:30