网页标题爬取故障排查:如何获取目标网站商家标题?
Hey there, let's fix your issue with scraping those business titles from the Official Black Wall Street directory! The main problem in your current code is how you're using findAll — that syntax ["a","title data-original-title"] isn't how you select elements with specific attributes in BeautifulSoup. Let's break down the solution step by step:
1. Fix the Element Selection
First, inspect the page's HTML (right-click > "Inspect") to confirm where the business titles live. For this directory, titles like McClean Photography are typically wrapped in <a> tags inside listing containers. Replace your incorrect findAll call with one of these targeted approaches:
Option 1: Grab anchor tags with visible title text
If titles are directly visible in<a>tags within listing titles:# Target the listing containers first, then extract titles listing_containers = bsObj.find_all("div", class_="listing-item") for container in listing_containers: title_tag = container.find("h3", class_="listing-title").find("a") if title_tag: bws_titles_bags.append(title_tag.get_text(strip=True))Option 2: Extract from
titleattributes
If titles are stored in thetitleattribute of anchor tags:title_links = bsObj.find_all("a", attrs={"title": True}) for link in title_links: # Use link["title"] to get the attribute value, or link.get_text() for visible text bws_titles_bags.append(link.get_text(strip=True))
2. Add Request Headers to Avoid Blocking
Many websites block default urllib requests because they lack a proper user-agent. Mimic a browser request by adding headers:
from urllib.request import Request, urlopen # Replace your original urlopen line with this req = Request("https://officialblackwallstreet.com/directory/", headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' }) html = urlopen(req)
3. Full Corrected Code
Putting it all together, here's the updated script:
from bs4 import BeautifulSoup as soup from urllib.request import Request, urlopen import pandas as pd import ssl from time import sleep from random import randint # Handle SSL verification try: _create_unverified_https_context = ssl._create_unverified_context except AttributeError: pass else: ssl._create_default_https_context = _create_unverified_context # Request page with proper headers req = Request("https://officialblackwallstreet.com/directory/", headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' }) html = urlopen(req) bsObj = soup(html.read(), "html.parser") # Explicitly specify parser for reliability bws_titles_bags = [] # Extract titles from listing containers listing_items = bsObj.find_all("div", class_="listing-item") for item in listing_items: title_elem = item.find("h3", class_="listing-title").find("a") if title_elem: bws_titles_bags.append(title_elem.get_text(strip=True)) # Check your results print(bws_titles_bags)
Quick Notes
- Always double-check the page's HTML structure to adjust selectors if needed — site layouts can change!
- Adding small random delays (
sleep(randint(1,3))) between requests can help avoid triggering anti-scraping measures if you expand to multiple pages.
内容的提问来源于stack exchange,提问作者Yo'python

