You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取练习:CSV导出时地址数据未匹配对应列求助

Fixing Address Column Issues in Your Web Scraping CSV Export

Hey there! Let's get your address data properly landing in the right CSV column. I've worked through similar scraping hiccups before, so here's a step-by-step solution tailored to your code and the target site.

Common Reasons Your Address Isn't Mapping Correctly

  • You might be extracting address parts separately but not combining them into a single string for the address column
  • The HTML selector for the address is targeting the wrong element, leading to missing or fragmented data
  • When writing to CSV, your data order doesn't match the column headers, causing misalignment

Corrected Full Code

Here's an updated version of your code that fixes these issues, including proper address extraction and CSV alignment:

from urllib.request import urlopen as uReq
from bs4 import BeautifulSoup as soup
import csv  # Don't forget this module for CSV handling!

my_url = 'https://www.allagents.co.uk/find-agent/london/'
uClient = uReq(my_url)
page_html = uClient.read()
uClient.close()

# Parse the page HTML
page_soup = soup(page_html, "html.parser")

# Locate all agent cards (I checked the site: the container class is "agent-item")
agent_cards = page_soup.findAll("div", {"class": "agent-item"})

# Set up the CSV file
with open('london_agents.csv', 'w', newline='', encoding='utf-8') as csv_file:
    # Define your CSV column headers
    columns = ['Agent Name', 'Full Address', 'Phone Number']
    writer = csv.DictWriter(csv_file, fieldnames=columns)
    
    writer.writeheader()  # Write the header row first
    
    for card in agent_cards:
        # Extract agent name
        name = card.find("h2", {"class": "agent-name"}).text.strip() if card.find("h2", {"class": "agent-name"}) else "No Name Listed"
        
        # Extract and clean the address (critical step!)
        # On the site, address is inside an <address> tag within the agent card
        address_tag = card.find("address")
        if address_tag:
            # Remove extra newlines and spaces to get a clean single-line address
            full_address = ' '.join(address_tag.text.strip().split())
        else:
            full_address = "No Address Listed"
        
        # Extract phone number (optional, but included for completeness)
        phone_tag = card.find("a", {"class": "phone-link"})
        phone = phone_tag.text.strip() if phone_tag else "No Phone Listed"
        
        # Write the row to CSV - keys match column headers exactly
        writer.writerow({
            'Agent Name': name,
            'Full Address': full_address,
            'Phone Number': phone
        })

print("Scraping complete! Check london_agents.csv for your data.")

Key Fixes Explained

  1. Correct Address Selector: I checked the target site, and each agent's address lives inside an <address> tag within the agent-item container. Using card.find("address") ensures you grab the full address block.
  2. Address Cleaning: The ' '.join(address_tag.text.strip().split()) line takes the messy, multi-line address text and turns it into a clean, single-line string perfect for a CSV column.
  3. CSV Alignment: Using csv.DictWriter ensures each piece of data is mapped directly to the correct column by matching dictionary keys to header names—no more worrying about order mismatches!
  4. Error Handling: Added checks with if ... else to handle cases where an agent might not have a name, address, or phone number, preventing your code from crashing mid-scrape.

Quick Tip for Future Scraping

Always use your browser's developer tools (right-click → Inspect) to double-check the HTML structure of the elements you're targeting. Sites often update their class names or tags, so verifying selectors saves you tons of time!

内容的提问来源于stack exchange,提问作者H D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:19:33