使用Python的BeautifulSoup无法提取网页文本求助
Hey there! Let's get your review extraction working properly—your empty Review field is almost certainly because you're targeting the wrong HTML element for the comment content.
What's Going Wrong?
If you inspect the page structure (using your browser's F12 developer tools), you'll see the actual review text ("I have open my account....") isn't stored in a <p> tag directly under the row review-article div. That's why row.p.text returns nothing.
Solution: Target the Correct Element
The review content is wrapped in a div with a class like review-description (double-check the exact class name on the live page if needed). Here's your adjusted code:
import requests from bs4 import BeautifulSoup import csv URL = "https://www.mouthshut.com/product-reviews/HDFC-Bank-reviews-925004501" r = requests.get(URL) soup = BeautifulSoup(r.content, 'html5lib') reviews = [] # a list to store reviews # Confirm the outer container selector is valid (tweak if the site updates its classes) review_container = soup.find('div', attrs={'class': 'container box-module ofinherit'}) if review_container: for review_row in review_container.findAll('div', attrs={'class': 'row review-article'}): # Extract title and URL with fallbacks for missing data review_title = review_row.a.text.strip() if review_row.a else "No Title" review_url = review_row.a['href'] if review_row.a else "" # Target the actual review content element review_content_tag = review_row.find('div', class_='review-description') review_content = review_content_tag.text.strip() if review_content_tag else "No Review Content" # Build the review dictionary review = { 'Review_Title': review_title, 'URL': review_url, 'Review': review_content } reviews.append(review) print(reviews)
Key Improvements:
- Correct Element Targeting: We now pull text from the
div.review-descriptiontag (adjust the class name if it changes on the live site) instead of a non-existent<p>tag. - Error Resilience: Added checks for missing titles/URLs/content to prevent your code from crashing if some reviews have incomplete data.
- Cleaner Text: Used
.strip()to remove extra whitespace and newlines from extracted text.
Quick Tip:
Always double-check the page's HTML structure with developer tools if your scraping fails—websites often update their class names or element hierarchy without warning!
内容的提问来源于stack exchange,提问作者Aayush Kaushal

