使用Python和BeautifulSoup解析含Java的HTML页面遇阻求助
Got it, let's break down how to parse this HTML with BeautifulSoup step by step. I've looked at your snippet and put together concrete examples to extract all the key data points you're probably after.
Step 1: Setup and Parse the HTML
First, make sure you've imported BeautifulSoup and loaded your HTML content correctly. I've completed the truncated part of your snippet for testing:
from bs4 import BeautifulSoup # Your HTML snippet (with truncated section fixed) html = """ <table class="customer-info-table"> <tr> <td> <label>Nome</label> XXXXXX XXXX </td> <td> <label>Cognome</label> XXXXXXX </td> </tr> <tr> <td> <label>Codice cliente</label> N/A </td> <td> <label>Indirizzo e-mail</label> XXXXXXXXX@gmail.com </td> </tr> </table> <span class="panel-headline">Commento articolo</span> <hr/> <i class="rating r1"></i><br/> BLACKLIST <textarea name="text" id="review-text" rows="12" readonly="readonly"> Mi avete mandato e-mail Ke il prodotto era disponibile; invece è esaurito........... Mah.</textarea> <span class="panel-headline">Commenti sulla vestibilità</span> <hr/> <table class="size-info-table"> <tr> <td> <label>Lunghezza</label> Viel zu kurz </td> <td> <label>Larghezza</label> Viel zu eng </td> <td> <label>Taglia</label>   </td> <td> <label>Varianti</label>   </td> <td> <label>Statura</label>   </td> </tr> </table> <p> <table class="table"> <tr> <td> <b>Rezensions-ID:</b> <span id="review-id">11166707</span> </td> <td> <b>Creata:</b> <span class="utc-date"> 23.03.2018 12:23:15 </span> </td> </tr> <tr> <td> <b>Letzte Änderung</b> <span class="utc-date"> 24.03.2018 09:12:00 </span> </td> </tr> </table> """ # Parse the HTML soup = BeautifulSoup(html, 'html.parser')
Step 2: Extract Customer Information
The customer data lives in the customer-info-table table. We can loop through rows and cells to pair labels with their values:
customer_info = {} customer_table = soup.find('table', class_='customer-info-table') for row in customer_table.find_all('tr'): cells = row.find_all('td') for cell in cells: label = cell.find('label').text.strip() # Grab the text after the label, clean up whitespace value = cell.text.replace(label, '').strip() customer_info[label] = value print("Customer Info:", customer_info) # Output: {'Nome': 'XXXXXX XXXX', 'Cognome': 'XXXXXXX', 'Codice cliente': 'N/A', 'Indirizzo e-mail': 'XXXXXXXXXX@gmail.com'}
Step 3: Extract Article Comment & Rating
The main comment is in the textarea with ID review-text. You can grab its content directly. We can also pull the rating from the i tag's class (here r1 likely means 1 star):
# Get the article comment article_comment = soup.find('textarea', id='review-text').text.strip() print("Article Comment:", article_comment) # Output: "Mi avete mandato e-mail Ke il prodotto era disponibile; invece è esaurito........... Mah." # Get the rating rating = soup.find('i', class_='rating')['class'][-1] print("Rating:", rating) # Output: 'r1'
Step 4: Extract Size Fit Information
Similar to the customer table, parse the size-info-table to get fit details. We'll handle non-breaking spaces ( ) by replacing them with "N/A":
size_info = {} size_table = soup.find('table', class_='size-info-table') for row in size_table.find_all('tr'): cells = row.find_all('td') for cell in cells: label = cell.find('label').text.strip() value = cell.text.replace(label, '').strip() # Replace non-breaking space with N/A for clarity if value == '\xa0': value = 'N/A' size_info[label] = value print("Size Info:", size_info) # Output: {'Lunghezza': 'Viel zu kurz', 'Larghezza': 'Viel zu eng', 'Taglia': 'N/A', 'Varianti': 'N/A', 'Statura': 'N/A'}
Step 5: Extract Review Metadata (ID & Dates)
The review ID and timestamps are in the generic table class. We'll target specific spans to pull these values:
review_metadata = {} meta_table = soup.find('table', class_='table') rows = meta_table.find_all('tr') # Extract review ID review_id = rows[0].find('span', id='review-id').text.strip() review_metadata['Rezensions-ID'] = review_id # Extract creation date creation_date = rows[0].find('span', class_='utc-date').text.strip() review_metadata['Creata'] = creation_date # Extract last modification date last_mod_date = rows[1].find('span', class_='utc-date').text.strip() review_metadata['Letzte Änderung'] = last_mod_date print("Review Metadata:", review_metadata) # Output: {'Rezensions-ID': '11166707', 'Creata': '23.03.2018 12:23:15', 'Letzte Änderung': '24.03.2018 09:12:00'}
Quick Tips to Avoid Common Issues:
- Use
class_instead ofclasswhen searching for elements by class name (sinceclassis a reserved keyword in Python). - If the label-value structure was more complex, you could use
cell.find('label').next_sibling.strip()instead ofreplace()to get the value after the label. - Always clean up whitespace with
.strip()to avoid extra spaces or newlines in your extracted data.
内容的提问来源于stack exchange,提问作者Greenfox

