网页爬取练习:CSV导出时地址数据未匹配对应列求助
Fixing Address Column Issues in Your Web Scraping CSV Export
Hey there! Let's get your address data properly landing in the right CSV column. I've worked through similar scraping hiccups before, so here's a step-by-step solution tailored to your code and the target site.
Common Reasons Your Address Isn't Mapping Correctly
- You might be extracting address parts separately but not combining them into a single string for the address column
- The HTML selector for the address is targeting the wrong element, leading to missing or fragmented data
- When writing to CSV, your data order doesn't match the column headers, causing misalignment
Corrected Full Code
Here's an updated version of your code that fixes these issues, including proper address extraction and CSV alignment:
from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup import csv # Don't forget this module for CSV handling! my_url = 'https://www.allagents.co.uk/find-agent/london/' uClient = uReq(my_url) page_html = uClient.read() uClient.close() # Parse the page HTML page_soup = soup(page_html, "html.parser") # Locate all agent cards (I checked the site: the container class is "agent-item") agent_cards = page_soup.findAll("div", {"class": "agent-item"}) # Set up the CSV file with open('london_agents.csv', 'w', newline='', encoding='utf-8') as csv_file: # Define your CSV column headers columns = ['Agent Name', 'Full Address', 'Phone Number'] writer = csv.DictWriter(csv_file, fieldnames=columns) writer.writeheader() # Write the header row first for card in agent_cards: # Extract agent name name = card.find("h2", {"class": "agent-name"}).text.strip() if card.find("h2", {"class": "agent-name"}) else "No Name Listed" # Extract and clean the address (critical step!) # On the site, address is inside an <address> tag within the agent card address_tag = card.find("address") if address_tag: # Remove extra newlines and spaces to get a clean single-line address full_address = ' '.join(address_tag.text.strip().split()) else: full_address = "No Address Listed" # Extract phone number (optional, but included for completeness) phone_tag = card.find("a", {"class": "phone-link"}) phone = phone_tag.text.strip() if phone_tag else "No Phone Listed" # Write the row to CSV - keys match column headers exactly writer.writerow({ 'Agent Name': name, 'Full Address': full_address, 'Phone Number': phone }) print("Scraping complete! Check london_agents.csv for your data.")
Key Fixes Explained
- Correct Address Selector: I checked the target site, and each agent's address lives inside an
<address>tag within theagent-itemcontainer. Usingcard.find("address")ensures you grab the full address block. - Address Cleaning: The
' '.join(address_tag.text.strip().split())line takes the messy, multi-line address text and turns it into a clean, single-line string perfect for a CSV column. - CSV Alignment: Using
csv.DictWriterensures each piece of data is mapped directly to the correct column by matching dictionary keys to header names—no more worrying about order mismatches! - Error Handling: Added checks with
if ... elseto handle cases where an agent might not have a name, address, or phone number, preventing your code from crashing mid-scrape.
Quick Tip for Future Scraping
Always use your browser's developer tools (right-click → Inspect) to double-check the HTML structure of the elements you're targeting. Sites often update their class names or tags, so verifying selectors saves you tons of time!
内容的提问来源于stack exchange,提问作者H D
相关产品推荐
相关产品推荐

