如何从Plain Table中提取href与title?Python爬虫求助
Hey there! Let's work through why you're getting None when trying to pull href and title attributes from that plain class table, and get you a working solution to generate a list or dictionary of those links.
Common Reasons for None Returns
First, let's quickly cover why you might be hitting this issue:
- You're not drilling down far enough into the table's structure (e.g., trying to pull links directly from the
<table>tag instead of its child<tr>/<td>elements) - The class name for the table is misspelled (double-check if it's exactly
plain—capitalization matters in BeautifulSoup!) - The links are nested inside other elements (like
<span>) within table cells, so your selector isn't reaching them
Working Solution Code
Let's start with a step-by-step implementation. First, assuming you already have your table object, here's how to extract the links properly:
1. Import Required Modules (if you haven't already)
from bs4 import BeautifulSoup from urllib.request import urlopen
2. Fetch the Table & Extract Links
This example will generate a list of dictionaries, each containing the href and title for every link in the table:
# Fetch your page and parse it url = "your-target-website-url" html_content = urlopen(url).read() soup = BeautifulSoup(html_content, "html.parser") # Get the plain-class table (make sure the class name matches exactly!) plain_table = soup.find("table", class_="plain") # Initialize an empty list to store our links extracted_links = [] # Option 1: Traverse table rows and cells explicitly for row in plain_table.find_all("tr"): # Check both data cells (<td>) and header cells (<th>) in case links are in headers for cell in row.find_all(["td", "th"]): # Find all <a> tags inside the cell link_tags = cell.find_all("a") for tag in link_tags: # Extract href and title—use a default empty string if title doesn't exist link_data = { "href": tag.get("href"), "title": tag.get("title", "") } extracted_links.append(link_data) # Option 2: Use CSS selectors for a more concise approach # extracted_links = [] # for link in plain_table.select("td a, th a"): # extracted_links.append({ # "href": link.get("href"), # "title": link.get("title", "") # }) # Print or use your extracted links for link in extracted_links: print(f"Title: {link['title']} | URL: {link['href']}")
Troubleshooting Tips
- Verify your table object: Print
plain_table.prettify()to make sure you're actually selecting the correct table. If this returnsNone, yourfind()call for the table is wrong (double-check the class name or try usingsoup.find_all("table")to list all tables on the page). - Check for nested elements: If links are inside
<span>or other tags, adjust your selector tocell.find_all("a", recursive=True)or useplain_table.select("table.plain * a")to catch all nested links. - Dynamic content?: If the table loads after JavaScript runs,
urlopenwon't capture it. In that case, you'll need a tool like Selenium to render the page first—but this is only if you confirm the links aren't present in the static HTML.
That should get you sorted! Let me know if you hit any snags with specific edge cases.
内容的提问来源于stack exchange,提问作者Reinaldo Chaves
相关产品推荐
相关产品推荐

