使用simple_html_dom.php爬取Houzz网站手机号失败求助
Hey there! I see you're stuck getting "Click to call" instead of actual phone numbers when scraping Houzz with simple_html_dom.php—let's break down why this happens and how to fix it.
Why You're Getting "Click to call"
Houzz loads phone numbers dynamically with JavaScript. When you first load the page, the HTML only shows the "Click to call" placeholder. The real number is fetched and inserted into the DOM only after you click the button. Since simple_html_dom only grabs the initial static HTML (no JavaScript execution), it never sees the real phone number.
Solutions to Get Real Phone Numbers
1. Check for Hidden Phone Numbers in Attributes
First, quick win: Some sites hide the real number in an HTML attribute (like data-phone or data-contact) even before you click. Let's check:
- Open the page in Chrome/Firefox, right-click "Click to call" → Inspect.
- Look at the parent
<span class="hz-pro-search-result__contact-info">or the cover span for any data attributes holding the phone number.
If you find something like <span class="hz-pro-search-result__contact-info__cover" data-phone="(954) 399-0454">Click to call</span>, modify your code to extract that attribute instead of plaintext:
// Replace this line: $videoTitle1 = $videoDetails1->plaintext; // With this (adjust the attribute name to match what you find): $videoTitle1 = $videoDetails1->find('span.hz-pro-search-result__contact-info__cover', 0)->getAttribute('data-phone');
2. Use a Tool That Executes JavaScript
If there's no hidden attribute, you need a scraper that mimics a real browser (loads JS, clicks buttons, waits for dynamic content). For PHP, great options are:
- Spatie Browsershot: Wraps Puppeteer to render pages with Chrome.
- Symfony Panther: A browser automation library for PHP.
Here's a quick example with Spatie Browsershot:
Step 1: Install the Package
composer require spatie/browsershot
(Note: You'll need Node.js and Puppeteer installed too—follow the official setup docs for Browsershot to get this working.)
Step 2: Modified Scraper Code
use Spatie\Browsershot\Browsershot; function scrape(){ // Render the page with Chrome, wait for all JS to load and dynamic content to populate $renderedHtml = Browsershot::url('https://www.houzz.com/professionals/handyman/c/US') ->waitUntilNetworkIdle() // Wait for all background requests to finish ->body(); // Parse the fully rendered HTML with simple_html_dom require_once APPPATH .'third_party/simple_html_dom.php'; $html = str_get_html($renderedHtml); $videos = []; $i = 1; foreach ($html->find('li.hz-pro-search-results__item') as $video) { if ($i > 10) { break; } // Now this will grab the actual phone number, not the placeholder $contactElement = $video->find('span.hz-pro-search-result__contact-info', 0); $phoneNumber = trim($contactElement->plaintext); $locationElement = $video->find('span.hz-pro-search-result__location-info__text', 0); $location = $locationElement->plaintext; $videos[] = [ 'phone' => $phoneNumber, 'location' => $location ]; $i++; } var_dump($videos); }
3. Reverse-Engineer the AJAX Request
When you click "Click to call", Houzz sends an AJAX request to fetch the phone number. You can find this request in your browser's DevTools (Network tab, filter by XHR/Fetch):
- Click "Click to call" and note the request URL (e.g.,
https://www.houzz.com/api/professional/contact?proId=12345). - Extract the professional ID from each listing (look for
data-pro-idor similar attributes on the list item). - Directly call this API endpoint with the pro ID to get the phone number, skipping the HTML scraping entirely.
This is often the most efficient method, but you'll need to handle any authentication or required headers the API expects.
Important Notes
- Always check Houzz's Terms of Service before scraping—make sure you're allowed to collect this data.
- Add small delays between requests to avoid getting your IP blocked.
- Some sites may use anti-scraping measures (like CAPTCHAs) if you make too many requests quickly.
内容的提问来源于stack exchange,提问作者shashi kumar

