You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式提取含子div的Product div全部内容?Python爬虫问题

Fixing Your Regex for Product Div Extraction

Hey there! The issue with your current regex is that the non-greedy [\s\S]*? stops at the first </div> it encounters—which happens to be the closing tag of the first child div (product-image-and-name-container), not the parent Product div.

To fix this, we need to adjust the regex to only stop at the </div> that closes the Product div. Here's how you can modify your root_pattern:

root_pattern = r'<div class="Product">([\s\S]*?)</div>(?=\s*<div class="Product">|$)'

How This Works:

  • The (?=\s*<div class="Product">|$) part is a positive lookahead that ensures the </div> we're matching is either:
    1. Followed immediately (with optional whitespace) by another Product div, or
    2. At the end of the HTML string.

This tells the regex to keep matching content inside the Product div until it hits the correct closing tag, ignoring the inner child divs' closing tags.

Updated Full Code:

from urllib.request import Request, urlopen
import re

class Shopping_Spider():
    url = 'http://www....com/Shop-Online/587'
    # Modified root pattern
    root_pattern = r'<div class="Product">([\s\S]*?)</div>(?=\s*<div class="Product">|$)'
    name_pattern = '<div class="product-name">([\s\S]*?)</div>'
    price_pattern = '<span class="Price">([\s\S]*?)</span>'

    def __fetch_content(self):
        r = Request(Shopping_Spider.url, headers={'User-Agent': 'Mozilla/5.0'})
        html_s = urlopen(r).read()
        html_s = str(html_s, encoding='utf-8')
        return html_s

    def __analysis(self, html_s):
        root_html = re.findall(Shopping_Spider.root_pattern, html_s)
        anchors = []
        for html in root_html:
            name = re.findall(Shopping_Spider.name_pattern, html)
            price = re.findall(Shopping_Spider.price_pattern, html)
            # Clean up results to handle empty matches
            anchor = {'name': name[0].strip() if name else 'No name', 
                      'price': price[0].strip() if price else 'No price'}
            anchors.append(anchor)
        return anchors

    def go(self):
        html_s = self.__fetch_content()
        results = self.__analysis(html_s)
        # Print results to verify extraction
        for item in results:
            print(f"Product: {item['name']}, Price: {item['price']}")

shopping_spider = Shopping_Spider()
shopping_spider.go()

A Quick Note:

While this regex fix works for your specific HTML structure, it's important to remember that regex is not designed for parsing HTML—small changes to the website's structure (like extra whitespace, nested divs, or updated class names) can break it. You mentioned you know BeautifulSoup is better, and that's absolutely true for robust web scraping. But since you wanted to stick with regex, this should solve your immediate problem.

内容的提问来源于stack exchange,提问作者Jimmy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:23:44