You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的BeautifulSoup无法提取网页文本求助

Fixing Empty Review Field in Your Web Scraping Code

Hey there! Let's get your review extraction working properly—your empty Review field is almost certainly because you're targeting the wrong HTML element for the comment content.

What's Going Wrong?

If you inspect the page structure (using your browser's F12 developer tools), you'll see the actual review text ("I have open my account....") isn't stored in a <p> tag directly under the row review-article div. That's why row.p.text returns nothing.

Solution: Target the Correct Element

The review content is wrapped in a div with a class like review-description (double-check the exact class name on the live page if needed). Here's your adjusted code:

import requests
from bs4 import BeautifulSoup
import csv

URL = "https://www.mouthshut.com/product-reviews/HDFC-Bank-reviews-925004501"
r = requests.get(URL)
soup = BeautifulSoup(r.content, 'html5lib')
reviews = [] # a list to store reviews

# Confirm the outer container selector is valid (tweak if the site updates its classes)
review_container = soup.find('div', attrs={'class': 'container box-module ofinherit'})

if review_container:
    for review_row in review_container.findAll('div', attrs={'class': 'row review-article'}):
        # Extract title and URL with fallbacks for missing data
        review_title = review_row.a.text.strip() if review_row.a else "No Title"
        review_url = review_row.a['href'] if review_row.a else ""
        
        # Target the actual review content element
        review_content_tag = review_row.find('div', class_='review-description')
        review_content = review_content_tag.text.strip() if review_content_tag else "No Review Content"
        
        # Build the review dictionary
        review = {
            'Review_Title': review_title,
            'URL': review_url,
            'Review': review_content
        }
        reviews.append(review)

print(reviews)

Key Improvements:

  • Correct Element Targeting: We now pull text from the div.review-description tag (adjust the class name if it changes on the live site) instead of a non-existent <p> tag.
  • Error Resilience: Added checks for missing titles/URLs/content to prevent your code from crashing if some reviews have incomplete data.
  • Cleaner Text: Used .strip() to remove extra whitespace and newlines from extracted text.

Quick Tip:

Always double-check the page's HTML structure with developer tools if your scraping fails—websites often update their class names or element hierarchy without warning!

内容的提问来源于stack exchange,提问作者Aayush Kaushal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 08:42:32