You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和BeautifulSoup解析含Java的HTML页面遇阻求助

How to Extract Data from Your HTML Snippet with BeautifulSoup

Got it, let's break down how to parse this HTML with BeautifulSoup step by step. I've looked at your snippet and put together concrete examples to extract all the key data points you're probably after.

Step 1: Setup and Parse the HTML

First, make sure you've imported BeautifulSoup and loaded your HTML content correctly. I've completed the truncated part of your snippet for testing:

from bs4 import BeautifulSoup

# Your HTML snippet (with truncated section fixed)
html = """
<table class="customer-info-table"> 
    <tr> 
        <td> <label>Nome</label> XXXXXX XXXX </td> 
        <td> <label>Cognome</label> XXXXXXX </td> 
    </tr> 
    <tr> 
        <td> <label>Codice cliente</label> N/A </td> 
        <td> <label>Indirizzo e-mail</label> XXXXXXXXX@gmail.com </td> 
    </tr> 
</table> 
<span class="panel-headline">Commento articolo</span> 
<hr/> 
<i class="rating r1"></i><br/> BLACKLIST 
<textarea name="text" id="review-text" rows="12" readonly="readonly"> Mi avete mandato e-mail Ke il prodotto era disponibile; invece è esaurito........... Mah.</textarea> 
<span class="panel-headline">Commenti sulla vestibilità</span> 
<hr/> 
<table class="size-info-table"> 
    <tr> 
        <td> <label>Lunghezza</label> Viel zu kurz </td> 
        <td> <label>Larghezza</label> Viel zu eng </td> 
        <td> <label>Taglia</label> &nbsp </td> 
        <td> <label>Varianti</label> &nbsp </td> 
        <td> <label>Statura</label> &nbsp </td> 
    </tr> 
</table> 
<p> 
<table class="table"> 
    <tr> 
        <td> <b>Rezensions-ID:</b> <span id="review-id">11166707</span> </td> 
        <td> <b>Creata:</b> <span class="utc-date"> 23.03.2018 12:23:15 </span> </td> 
    </tr> 
    <tr> 
        <td> <b>Letzte Änderung</b> <span class="utc-date"> 24.03.2018 09:12:00 </span> </td>
    </tr>
</table>
"""

# Parse the HTML
soup = BeautifulSoup(html, 'html.parser')

Step 2: Extract Customer Information

The customer data lives in the customer-info-table table. We can loop through rows and cells to pair labels with their values:

customer_info = {}
customer_table = soup.find('table', class_='customer-info-table')

for row in customer_table.find_all('tr'):
    cells = row.find_all('td')
    for cell in cells:
        label = cell.find('label').text.strip()
        # Grab the text after the label, clean up whitespace
        value = cell.text.replace(label, '').strip()
        customer_info[label] = value

print("Customer Info:", customer_info)
# Output: {'Nome': 'XXXXXX XXXX', 'Cognome': 'XXXXXXX', 'Codice cliente': 'N/A', 'Indirizzo e-mail': 'XXXXXXXXXX@gmail.com'}

Step 3: Extract Article Comment & Rating

The main comment is in the textarea with ID review-text. You can grab its content directly. We can also pull the rating from the i tag's class (here r1 likely means 1 star):

# Get the article comment
article_comment = soup.find('textarea', id='review-text').text.strip()
print("Article Comment:", article_comment)
# Output: "Mi avete mandato e-mail Ke il prodotto era disponibile; invece è esaurito........... Mah."

# Get the rating
rating = soup.find('i', class_='rating')['class'][-1]
print("Rating:", rating)
# Output: 'r1'

Step 4: Extract Size Fit Information

Similar to the customer table, parse the size-info-table to get fit details. We'll handle non-breaking spaces (&nbsp;) by replacing them with "N/A":

size_info = {}
size_table = soup.find('table', class_='size-info-table')

for row in size_table.find_all('tr'):
    cells = row.find_all('td')
    for cell in cells:
        label = cell.find('label').text.strip()
        value = cell.text.replace(label, '').strip()
        # Replace non-breaking space with N/A for clarity
        if value == '\xa0':
            value = 'N/A'
        size_info[label] = value

print("Size Info:", size_info)
# Output: {'Lunghezza': 'Viel zu kurz', 'Larghezza': 'Viel zu eng', 'Taglia': 'N/A', 'Varianti': 'N/A', 'Statura': 'N/A'}

Step 5: Extract Review Metadata (ID & Dates)

The review ID and timestamps are in the generic table class. We'll target specific spans to pull these values:

review_metadata = {}
meta_table = soup.find('table', class_='table')
rows = meta_table.find_all('tr')

# Extract review ID
review_id = rows[0].find('span', id='review-id').text.strip()
review_metadata['Rezensions-ID'] = review_id

# Extract creation date
creation_date = rows[0].find('span', class_='utc-date').text.strip()
review_metadata['Creata'] = creation_date

# Extract last modification date
last_mod_date = rows[1].find('span', class_='utc-date').text.strip()
review_metadata['Letzte Änderung'] = last_mod_date

print("Review Metadata:", review_metadata)
# Output: {'Rezensions-ID': '11166707', 'Creata': '23.03.2018 12:23:15', 'Letzte Änderung': '24.03.2018 09:12:00'}

Quick Tips to Avoid Common Issues:

  • Use class_ instead of class when searching for elements by class name (since class is a reserved keyword in Python).
  • If the label-value structure was more complex, you could use cell.find('label').next_sibling.strip() instead of replace() to get the value after the label.
  • Always clean up whitespace with .strip() to avoid extra spaces or newlines in your extracted data.

内容的提问来源于stack exchange,提问作者Greenfox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:02:44