You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Plain Table中提取href与title?Python爬虫求助

Hey there! Let's work through why you're getting None when trying to pull href and title attributes from that plain class table, and get you a working solution to generate a list or dictionary of those links.

Common Reasons for None Returns

First, let's quickly cover why you might be hitting this issue:

  • You're not drilling down far enough into the table's structure (e.g., trying to pull links directly from the <table> tag instead of its child <tr>/<td> elements)
  • The class name for the table is misspelled (double-check if it's exactly plain—capitalization matters in BeautifulSoup!)
  • The links are nested inside other elements (like <span>) within table cells, so your selector isn't reaching them

Working Solution Code

Let's start with a step-by-step implementation. First, assuming you already have your table object, here's how to extract the links properly:

1. Import Required Modules (if you haven't already)

from bs4 import BeautifulSoup
from urllib.request import urlopen

2. Fetch the Table & Extract Links

This example will generate a list of dictionaries, each containing the href and title for every link in the table:

# Fetch your page and parse it
url = "your-target-website-url"
html_content = urlopen(url).read()
soup = BeautifulSoup(html_content, "html.parser")

# Get the plain-class table (make sure the class name matches exactly!)
plain_table = soup.find("table", class_="plain")

# Initialize an empty list to store our links
extracted_links = []

# Option 1: Traverse table rows and cells explicitly
for row in plain_table.find_all("tr"):
    # Check both data cells (<td>) and header cells (<th>) in case links are in headers
    for cell in row.find_all(["td", "th"]):
        # Find all <a> tags inside the cell
        link_tags = cell.find_all("a")
        for tag in link_tags:
            # Extract href and title—use a default empty string if title doesn't exist
            link_data = {
                "href": tag.get("href"),
                "title": tag.get("title", "")
            }
            extracted_links.append(link_data)

# Option 2: Use CSS selectors for a more concise approach
# extracted_links = []
# for link in plain_table.select("td a, th a"):
#     extracted_links.append({
#         "href": link.get("href"),
#         "title": link.get("title", "")
#     })

# Print or use your extracted links
for link in extracted_links:
    print(f"Title: {link['title']} | URL: {link['href']}")

Troubleshooting Tips

  • Verify your table object: Print plain_table.prettify() to make sure you're actually selecting the correct table. If this returns None, your find() call for the table is wrong (double-check the class name or try using soup.find_all("table") to list all tables on the page).
  • Check for nested elements: If links are inside <span> or other tags, adjust your selector to cell.find_all("a", recursive=True) or use plain_table.select("table.plain * a") to catch all nested links.
  • Dynamic content?: If the table loads after JavaScript runs, urlopen won't capture it. In that case, you'll need a tool like Selenium to render the page first—but this is only if you confirm the links aren't present in the static HTML.

That should get you sorted! Let me know if you hit any snags with specific edge cases.

内容的提问来源于stack exchange,提问作者Reinaldo Chaves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:49:32