如何用Python爬取无ID/Class标签的网页表格并提取指定地址
Hey Joseph, let's figure out how to extract that Location address from the table using BeautifulSoup and Requests. Your idea of using next_sibling is solid, but we need to account for common HTML quirks like whitespace text nodes. Here's how to make it work, plus some more reliable approaches depending on the page structure:
First, let's start with a complete version of your base code to ensure the request and parsing setup is correct:
import requests from bs4 import BeautifulSoup # Target URL url = "http://wakefield.patriotproperties.com/Summary.asp?AccountNumber=6867" response = requests.get(url) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') else: print(f"Request failed with status code: {response.status_code}")
Approach 1: Using next_sibling (with whitespace handling)
Directly targeting the "Location" label and grabbing its next sibling works, but you'll often run into empty text nodes (from line breaks/spaces in HTML). We'll filter those out:
# Find the <td> containing "Location" (adjust tag if your page uses <th> or another element) location_label = soup.find('td', string=lambda text: text and 'Location' in text.strip()) if location_label: # Start with the immediate next sibling address_node = location_label.next_sibling # Skip empty/whitespace-only nodes while address_node and address_node.strip() == "": address_node = address_node.next_sibling # Extract the clean address if address_node: print("Extracted Location:", address_node.strip()) else: print("No address found next to Location label") else: print("Couldn't find the Location label")
Approach 2: Target the Parent Row (More Reliable)
If the Location label and address are in the same table row (<tr>), grabbing the row first then accessing the address cell is often more robust:
# Find the table row containing "Location" location_row = soup.find('tr', string=lambda text: text and 'Location' in text.strip()) if location_row: # Get all cells in the row cells = location_row.find_all('td') # Assuming address is the second cell (adjust index if needed) if len(cells) >= 2: clean_address = cells[1].get_text(strip=True) print("Extracted Location:", clean_address) else: print("Not enough cells in the Location row") else: print("Couldn't find the row with Location")
Approach 3: Using find_next for Simplicity
A shorter alternative is to find the Location label, then directly grab the next relevant cell:
location_label = soup.find('td', string=lambda text: text and 'Location' in text.strip()) if location_label: # Grab the next <td> (adjust tag if your address is in a different element) address_cell = location_label.find_next('td') print("Extracted Location:", address_cell.get_text(strip=True)) else: print("Couldn't find the Location label")
Pro Tip: Always inspect the page's HTML structure first (using your browser's dev tools) to confirm exactly what tags wrap the Location label and address. You can also print soup.prettify() to see the parsed structure and adjust the selectors accordingly.
内容的提问来源于stack exchange,提问作者Joseph DeCola

