You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python & BS4抓取网页中含Product A的表格TD元素

Solution for Scraping TD Elements Containing "Product A"

Hey there! Since you're new to Python and web scraping, let's adjust your code to meet your requirement of finding all <td> elements that contain the text "Product A" across the entire webpage. Here's how to do it step by step:

Step 1: Core Approach

Instead of targeting a specific table upfront, we'll:

  • Scan the entire parsed HTML for text nodes that include "Product A"
  • For each matching text node, grab its parent <td> element
  • Collect all valid <td> elements into an array for later use

Step 2: Modified Working Code

Here's the updated code that implements this logic, with notes on key changes from your original script:

import urllib2
from bs4 import BeautifulSoup

# Specify your target webpage/local file path
quote_page = 'Website.html'

# Fetch and parse the HTML content
page = urllib2.urlopen(quote_page)
soup = BeautifulSoup(page, 'html.parser')

# Initialize an empty list to store our target TD elements
target_tds = []

# Find all text nodes that contain "Product A"
# The lambda filters out None values and checks for the target text
matching_texts = soup.find_all(text=lambda text: text and "Product A" in text.strip())

# Iterate over each matching text node to get its parent TD
for text_node in matching_texts:
    parent_tag = text_node.parent
    # Add a safety check to ensure we only collect TD elements
    if parent_tag.name == 'td':
        target_tds.append(parent_tag)

# Output the results
print(f"Found {len(target_tds)} TD elements containing 'Product A':")
for td in target_tds:
    # Print the full HTML of the TD element
    print(td.prettify())
    # Uncomment below to print just the text content instead:
    # print(td.get_text(strip=True))

Key Details to Understand:

  • soup.find_all(text=lambda...): This searches the entire document for text nodes that include "Product A". Using strip() ensures we don't miss matches hidden behind extra whitespace.
  • text_node.parent: Grabs the direct parent element of the matching text node—this is the <td> element you need, based on your HTML structure.
  • Validation Check: The if parent_tag.name == 'td' line adds a safety net in case some matching text ends up inside an unexpected tag (though your provided HTML suggests this won't be an issue).

Quick Note for Python 3 Users:

If you're using Python 3, urllib2 is deprecated. Replace the import and page-fetching lines with:

import urllib.request
page = urllib.request.urlopen(quote_page)

内容的提问来源于stack exchange,提问作者Yanick L.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:03:59