You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取:如何正确指定HTML标签与类?解决新闻链接爬取返回None问题

Let's walk through the issues in your code and fix them step by step—you're close, just a few small mistakes are causing that None output.

1. Incorrect requests.get() Syntax

You passed "html.parser" as the second argument to requests.get(), but that's not how the function works. The second parameter is for URL parameters, not the HTML parser. The parser belongs to the BeautifulSoup initialization instead.

2. Wrong Class Name

You used col_sm_5 in your attrs dictionary, but the actual class name is col-sm-5 (hyphens, not underscores). Class names are case and character-sensitive, so this mismatch is why your find call returns None.

3. Targeting the Right Elements

Since your target links live inside 4 li tags containing h5.col-sm-5 elements, you'll want to use find_all() instead of find() to capture all matching elements (not just the first one).


Corrected Code

Here's the fixed version with explanations:

import requests
from bs4 import BeautifulSoup

# Optional but critical: Add a user-agent to mimic a browser (avoids being blocked)
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# Correctly fetch the page
page = requests.get("http://www3.asiainsurancereview.com/News", headers=headers)
# Check if the request succeeded (raises an error for 4xx/5xx status codes)
page.raise_for_status()

# Parse the HTML with BeautifulSoup
soup = BeautifulSoup(page.text, "html.parser")

# Find all h5 tags with the correct class name
target_h5s = soup.find_all("h5", class_="col-sm-5")

# Extract links from each h5
for h5 in target_h5s:
    # Get the <a> tag inside the h5
    link_tag = h5.find("a")
    if link_tag:
        # Extract the href attribute
        news_link = link_tag.get("href")
        print(news_link)

Key Notes:

  • Use class_ instead of attrs={'class': ...} (it's a cleaner syntax for BeautifulSoup, since class is a reserved keyword in Python).
  • find_all() returns a list of all matching elements, which is perfect for your 4 target li/h5 pairs.
  • Adding a User-Agent header helps avoid being flagged as a bot by the website's server.
  • page.raise_for_status() helps catch issues like broken URLs or blocked requests early.

If It Still Doesn't Work:

If you still get no results, the content might be dynamically loaded with JavaScript. In that case, you'd need to use a tool like Selenium to render the page fully before parsing. But start with the code above—most news sites serve static content for their listings.

内容的提问来源于stack exchange,提问作者Kristada673

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:38:26