You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页标题爬取故障排查:如何获取目标网站商家标题?

Hey there, let's fix your issue with scraping those business titles from the Official Black Wall Street directory! The main problem in your current code is how you're using findAll — that syntax ["a","title data-original-title"] isn't how you select elements with specific attributes in BeautifulSoup. Let's break down the solution step by step:

1. Fix the Element Selection

First, inspect the page's HTML (right-click > "Inspect") to confirm where the business titles live. For this directory, titles like McClean Photography are typically wrapped in <a> tags inside listing containers. Replace your incorrect findAll call with one of these targeted approaches:

  • Option 1: Grab anchor tags with visible title text
    If titles are directly visible in <a> tags within listing titles:

    # Target the listing containers first, then extract titles
    listing_containers = bsObj.find_all("div", class_="listing-item")
    for container in listing_containers:
        title_tag = container.find("h3", class_="listing-title").find("a")
        if title_tag:
            bws_titles_bags.append(title_tag.get_text(strip=True))
    
  • Option 2: Extract from title attributes
    If titles are stored in the title attribute of anchor tags:

    title_links = bsObj.find_all("a", attrs={"title": True})
    for link in title_links:
        # Use link["title"] to get the attribute value, or link.get_text() for visible text
        bws_titles_bags.append(link.get_text(strip=True))
    

2. Add Request Headers to Avoid Blocking

Many websites block default urllib requests because they lack a proper user-agent. Mimic a browser request by adding headers:

from urllib.request import Request, urlopen

# Replace your original urlopen line with this
req = Request("https://officialblackwallstreet.com/directory/", headers={
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
})
html = urlopen(req)

3. Full Corrected Code

Putting it all together, here's the updated script:

from bs4 import BeautifulSoup as soup
from urllib.request import Request, urlopen
import pandas as pd
import ssl
from time import sleep
from random import randint

# Handle SSL verification
try:
    _create_unverified_https_context = ssl._create_unverified_context
except AttributeError:
    pass
else:
    ssl._create_default_https_context = _create_unverified_context

# Request page with proper headers
req = Request("https://officialblackwallstreet.com/directory/", headers={
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
})
html = urlopen(req)
bsObj = soup(html.read(), "html.parser")  # Explicitly specify parser for reliability

bws_titles_bags = []

# Extract titles from listing containers
listing_items = bsObj.find_all("div", class_="listing-item")
for item in listing_items:
    title_elem = item.find("h3", class_="listing-title").find("a")
    if title_elem:
        bws_titles_bags.append(title_elem.get_text(strip=True))

# Check your results
print(bws_titles_bags)

Quick Notes

  • Always double-check the page's HTML structure to adjust selectors if needed — site layouts can change!
  • Adding small random delays (sleep(randint(1,3))) between requests can help avoid triggering anti-scraping measures if you expand to multiple pages.

内容的提问来源于stack exchange,提问作者Yo'python

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:02:42