网页爬取:如何正确指定HTML标签与类?解决新闻链接爬取返回None问题
Let's walk through the issues in your code and fix them step by step—you're close, just a few small mistakes are causing that None output.
1. Incorrect requests.get() Syntax
You passed "html.parser" as the second argument to requests.get(), but that's not how the function works. The second parameter is for URL parameters, not the HTML parser. The parser belongs to the BeautifulSoup initialization instead.
2. Wrong Class Name
You used col_sm_5 in your attrs dictionary, but the actual class name is col-sm-5 (hyphens, not underscores). Class names are case and character-sensitive, so this mismatch is why your find call returns None.
3. Targeting the Right Elements
Since your target links live inside 4 li tags containing h5.col-sm-5 elements, you'll want to use find_all() instead of find() to capture all matching elements (not just the first one).
Corrected Code
Here's the fixed version with explanations:
import requests from bs4 import BeautifulSoup # Optional but critical: Add a user-agent to mimic a browser (avoids being blocked) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Correctly fetch the page page = requests.get("http://www3.asiainsurancereview.com/News", headers=headers) # Check if the request succeeded (raises an error for 4xx/5xx status codes) page.raise_for_status() # Parse the HTML with BeautifulSoup soup = BeautifulSoup(page.text, "html.parser") # Find all h5 tags with the correct class name target_h5s = soup.find_all("h5", class_="col-sm-5") # Extract links from each h5 for h5 in target_h5s: # Get the <a> tag inside the h5 link_tag = h5.find("a") if link_tag: # Extract the href attribute news_link = link_tag.get("href") print(news_link)
Key Notes:
- Use
class_instead ofattrs={'class': ...}(it's a cleaner syntax for BeautifulSoup, sinceclassis a reserved keyword in Python). find_all()returns a list of all matching elements, which is perfect for your 4 target li/h5 pairs.- Adding a
User-Agentheader helps avoid being flagged as a bot by the website's server. page.raise_for_status()helps catch issues like broken URLs or blocked requests early.
If It Still Doesn't Work:
If you still get no results, the content might be dynamically loaded with JavaScript. In that case, you'd need to use a tool like Selenium to render the page fully before parsing. But start with the code above—most news sites serve static content for their listings.
内容的提问来源于stack exchange,提问作者Kristada673

