如何用正则表达式提取含子div的Product div全部内容?Python爬虫问题
Hey there! The issue with your current regex is that the non-greedy [\s\S]*? stops at the first </div> it encounters—which happens to be the closing tag of the first child div (product-image-and-name-container), not the parent Product div.
To fix this, we need to adjust the regex to only stop at the </div> that closes the Product div. Here's how you can modify your root_pattern:
root_pattern = r'<div class="Product">([\s\S]*?)</div>(?=\s*<div class="Product">|$)'
How This Works:
- The
(?=\s*<div class="Product">|$)part is a positive lookahead that ensures the</div>we're matching is either:- Followed immediately (with optional whitespace) by another
Productdiv, or - At the end of the HTML string.
- Followed immediately (with optional whitespace) by another
This tells the regex to keep matching content inside the Product div until it hits the correct closing tag, ignoring the inner child divs' closing tags.
Updated Full Code:
from urllib.request import Request, urlopen import re class Shopping_Spider(): url = 'http://www....com/Shop-Online/587' # Modified root pattern root_pattern = r'<div class="Product">([\s\S]*?)</div>(?=\s*<div class="Product">|$)' name_pattern = '<div class="product-name">([\s\S]*?)</div>' price_pattern = '<span class="Price">([\s\S]*?)</span>' def __fetch_content(self): r = Request(Shopping_Spider.url, headers={'User-Agent': 'Mozilla/5.0'}) html_s = urlopen(r).read() html_s = str(html_s, encoding='utf-8') return html_s def __analysis(self, html_s): root_html = re.findall(Shopping_Spider.root_pattern, html_s) anchors = [] for html in root_html: name = re.findall(Shopping_Spider.name_pattern, html) price = re.findall(Shopping_Spider.price_pattern, html) # Clean up results to handle empty matches anchor = {'name': name[0].strip() if name else 'No name', 'price': price[0].strip() if price else 'No price'} anchors.append(anchor) return anchors def go(self): html_s = self.__fetch_content() results = self.__analysis(html_s) # Print results to verify extraction for item in results: print(f"Product: {item['name']}, Price: {item['price']}") shopping_spider = Shopping_Spider() shopping_spider.go()
A Quick Note:
While this regex fix works for your specific HTML structure, it's important to remember that regex is not designed for parsing HTML—small changes to the website's structure (like extra whitespace, nested divs, or updated class names) can break it. You mentioned you know BeautifulSoup is better, and that's absolutely true for robust web scraping. But since you wanted to stick with regex, this should solve your immediate problem.
内容的提问来源于stack exchange,提问作者Jimmy

