You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup文本中匹配正则?Angular ng-template页面爬取求助

Scraping Angular ng-template Content with Selenium & BeautifulSoup

Hey there! Sounds like you’re already off to a solid start using Selenium to handle Angular’s dynamic content and BeautifulSoup for parsing. Let’s break down how to effectively work with ng-template content, whether you stick with BeautifulSoup or bring in regex for specific use cases.

First, a quick reality check: ng-template is Angular’s way of defining reusable templates that don’t render until activated (via directives like *ngIf, *ngFor, or component logic). Since you’re using Selenium, you’re already covering the critical step of waiting for the page to fully render those templates—smart move, because static scrapers would completely miss that dynamic content.

Option 1: Stick with BeautifulSoup (More Reliable for HTML)

Since you already grabbed <address> tag content, extending this to ng-template content is straightforward. BeautifulSoup can target ng-template tags directly, then parse their inner HTML just like any other element:

from bs4 import BeautifulSoup

# Assuming you've already fetched the fully rendered page source via Selenium
soup = BeautifulSoup(driver.page_source, 'html.parser')

# Find all ng-template elements (narrow down with attributes like template references or IDs)
target_templates = soup.find_all('ng-template', attrs={'#addressTemplate': True})  # Example with template reference

for template in target_templates:
    # Extract inner elements just like normal HTML
    address = template.find('address')
    if address:
        clean_address = address.get_text(strip=True)
        print(f"Extracted address: {clean_address}")
    
    # Grab other nested elements too
    phone_number = template.find('p', class_='contact-phone')
    if phone_number:
        print(f"Extracted phone: {phone_number.get_text(strip=True)}")

Option 2: Using Regex for Specific Patterns

If you need to extract something tricky that’s hard to target with tag selectors (like dynamic text patterns), regex can work—but use it sparingly. HTML structure changes can break regex easily, so only reach for it when BeautifulSoup can’t handle the job. Here’s how to combine it with your existing setup:

import re

# Get the full HTML of your target ng-template as a string
template_html = str(soup.find('ng-template', id='userProfileTemplate'))

# Example: Extract an email address pattern from the template
email_match = re.search(r'Email:\s*([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})', template_html)
if email_match:
    extracted_email = email_match.group(1)
    print(f"Found email: {extracted_email}")

# Another example: Pull a zip code from address text
address_text = soup.find('address').get_text(strip=True)
zip_match = re.search(r'\b\d{5}(-\d{4})?\b', address_text)
if zip_match:
    print(f"Found zip code: {zip_match.group(0)}")

Pro Tips for Angular Scraping

  • Wait for Full Rendering: Use Selenium’s WebDriverWait to wait for specific elements inside the ng-template to load before parsing. This avoids trying to scrape content that hasn’t been rendered yet:
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    # Wait up to 10 seconds for the address inside the template to appear
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "ng-template #addressTemplate address"))
    )
    
  • Prioritize BeautifulSoup Over Regex: Whenever possible, use tag/attribute selectors—they’re way more resilient to minor HTML tweaks than regex.
  • Inspect Rendered HTML: Always check the browser’s "Elements" tab (not the raw "Page Source") to see the fully rendered ng-template content. The raw source will only show the unprocessed template, not the dynamic content you need.

内容的提问来源于stack exchange,提问作者TheDude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:27:11