You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬取无ID/Class标签的网页表格并提取指定地址

Hey Joseph, let's figure out how to extract that Location address from the table using BeautifulSoup and Requests. Your idea of using next_sibling is solid, but we need to account for common HTML quirks like whitespace text nodes. Here's how to make it work, plus some more reliable approaches depending on the page structure:

Extracting Location Address from Web Tables

First, let's start with a complete version of your base code to ensure the request and parsing setup is correct:

import requests
from bs4 import BeautifulSoup

# Target URL
url = "http://wakefield.patriotproperties.com/Summary.asp?AccountNumber=6867"
response = requests.get(url)

if response.status_code == 200:
    soup = BeautifulSoup(response.text, 'html.parser')
else:
    print(f"Request failed with status code: {response.status_code}")

Approach 1: Using next_sibling (with whitespace handling)

Directly targeting the "Location" label and grabbing its next sibling works, but you'll often run into empty text nodes (from line breaks/spaces in HTML). We'll filter those out:

# Find the <td> containing "Location" (adjust tag if your page uses <th> or another element)
location_label = soup.find('td', string=lambda text: text and 'Location' in text.strip())

if location_label:
    # Start with the immediate next sibling
    address_node = location_label.next_sibling
    # Skip empty/whitespace-only nodes
    while address_node and address_node.strip() == "":
        address_node = address_node.next_sibling
    # Extract the clean address
    if address_node:
        print("Extracted Location:", address_node.strip())
    else:
        print("No address found next to Location label")
else:
    print("Couldn't find the Location label")

Approach 2: Target the Parent Row (More Reliable)

If the Location label and address are in the same table row (<tr>), grabbing the row first then accessing the address cell is often more robust:

# Find the table row containing "Location"
location_row = soup.find('tr', string=lambda text: text and 'Location' in text.strip())

if location_row:
    # Get all cells in the row
    cells = location_row.find_all('td')
    # Assuming address is the second cell (adjust index if needed)
    if len(cells) >= 2:
        clean_address = cells[1].get_text(strip=True)
        print("Extracted Location:", clean_address)
    else:
        print("Not enough cells in the Location row")
else:
    print("Couldn't find the row with Location")

Approach 3: Using find_next for Simplicity

A shorter alternative is to find the Location label, then directly grab the next relevant cell:

location_label = soup.find('td', string=lambda text: text and 'Location' in text.strip())

if location_label:
    # Grab the next <td> (adjust tag if your address is in a different element)
    address_cell = location_label.find_next('td')
    print("Extracted Location:", address_cell.get_text(strip=True))
else:
    print("Couldn't find the Location label")

Pro Tip: Always inspect the page's HTML structure first (using your browser's dev tools) to confirm exactly what tags wrap the Location label and address. You can also print soup.prettify() to see the parsed structure and adjust the selectors accordingly.


内容的提问来源于stack exchange,提问作者Joseph DeCola

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:15:20