You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:基于Robobrowser爬取ReferenceUSA的技术问题

Guide to Completing Your ReferenceUSA Scraping with Robobrowser

Hey there! No worries at all about being a beginner—we all start somewhere, and it’s totally okay if your initial question isn’t perfectly polished. Let’s break down the next steps to get your scraping task up and running with Robobrowser:

1. Complete the Login Process

First, you’ll need to fill and submit the login form on the page you opened. Robobrowser makes this straightforward, but you’ll need to identify the correct form fields (like username and password inputs) from the page’s HTML.

Here’s a sample implementation:

# Locate the login form (adjust the selector if needed—use browser.select('form') to list all forms)
login_form = browser.get_form()

# Fill in your credentials (replace 'username' and 'password' with the actual field names from the form)
login_form['username'].value = "your_columbia_username"
login_form['password'].value = "your_columbia_password"

# Submit the form to log in
browser.submit_form(login_form)

# Verify login success (check for a logged-in element, like a welcome message or user profile link)
if browser.find(text="Welcome"):
    print("Login successful!")
else:
    print("Login failed—double-check your form field names and credentials.")

Pro tip: If you’re unsure about the field names, print the form with print(login_form) to see all available fields.

2. Navigate to Your Hardcoded Search Results Page

Once logged in, use browser.open() to load your pre-defined search results URL:

# Replace with your actual hardcoded search results URL
search_results_url = "https://www.referenceusa.com/your-search-results-link"
browser.open(search_results_url)

Make sure this URL is accessible only after login—if you get redirected back to the login page, your session might not be persisting correctly (Robobrowser usually handles cookies automatically, but double-check the login step).

3. Extract Data from Search Results

Now it’s time to scrape the data you need. Use Robobrowser’s built-in BeautifulSoup methods to locate elements on the page. For example, if your results are in a list with class result-item, here’s how to extract key details:

# Locate all result items (adjust the CSS selector to match the page's structure)
result_items = browser.select(".result-item")

for item in result_items:
    # Extract company name (update the tag and class to match your page)
    company_name = item.find("h2", class_="company-title").get_text(strip=True)
    
    # Extract address (handle cases where the element might be missing)
    address = item.find("div", class_="company-address").get_text(strip=True) if item.find("div", class_="company-address") else "N/A"
    
    # Extract phone number
    phone = item.find("span", class_="company-phone").get_text(strip=True) if item.find("span", class_="company-phone") else "N/A"
    
    # Print or store the data
    print(f"Company: {company_name}\nAddress: {address}\nPhone: {phone}\n---")

4. Handle Pagination (If Needed)

If your search results span multiple pages, you’ll need to navigate through them. Look for a "Next Page" link and use browser.follow_link() to click it:

import time

while True:
    # Extract data from current page (use the code from step 3)
    result_items = browser.select(".result-item")
    # ... process items ...
    
    # Find the next page link
    next_page_link = browser.find("a", class_="next-page")
    if not next_page_link:
        break  # No more pages
    
    # Follow the link and add a small delay to avoid overwhelming the server
    browser.follow_link(next_page_link)
    time.sleep(2)  # Adjust delay as needed to comply with site rules

5. Save Data to CSV

You already imported the csv module, so let’s put it to use saving your scraped data:

with open("referenceusa_data.csv", "w", newline="", encoding="utf-8") as csv_file:
    # Define your CSV columns
    fieldnames = ["Company Name", "Address", "Phone"]
    writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
    
    writer.writeheader()  # Write column headers
    
    # Loop through results and write rows
    result_items = browser.select(".result-item")
    for item in result_items:
        company_name = item.find("h2", class_="company-title").get_text(strip=True)
        address = item.find("div", class_="company-address").get_text(strip=True) if item.find("div", class_="company-address") else "N/A"
        phone = item.find("span", class_="company-phone").get_text(strip=True) if item.find("span", class_="company-phone") else "N/A"
        
        writer.writerow({
            "Company Name": company_name,
            "Address": address,
            "Phone": phone
        })

Important Notes to Keep in Mind

  • Respect Site Policies: ReferenceUSA (and Columbia’s access) likely has terms of service—avoid scraping at high speeds, and don’t scrape more data than you need.
  • Dynamic Content: If parts of the page load with JavaScript, Robobrowser won’t execute it. In that case, you might need to switch to a tool like Selenium that can render JS.
  • Error Handling: Add try-except blocks around your scraping code to handle missing elements or network errors, so your script doesn’t crash unexpectedly.

内容的提问来源于stack exchange,提问作者user2970395

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:04:20