Python爬取Rightmove房产数据时关联原始邮编的解决方法
关联爬取数据与原始邮编的实现方案
核心逻辑
现有代码无法关联邮编的原因是遍历爬取任务时,只单独提取了URL列做循环,没有把URL对应的邮编在整个爬取流程中传递,导致解析存储数据时拿不到对应的邮编值。只需要调整参数传递逻辑,把邮编随爬取链路传到存储步骤即可,不需要修改原有房产字段提取的核心逻辑。
具体修改步骤
- 调整
run方法的遍历逻辑:不再单独遍历url_name列,改为逐行遍历数据集,每次循环同时拿到当前行的目标URL和对应的postcode值 - 调整
parse方法的入参:在原有html参数的基础上,新增postcode参数接收当前爬取页面对应的邮编 - 调整结果存储逻辑:在
parse方法拼接房产数据字典时,新增postcode字段,把传入的邮编值写入每条爬取结果中
修改后的核心代码
import requests import csv from bs4 import BeautifulSoup import pandas as pd class RightmoveScraper: results = [] def fetch(self, url): print('HTTP GET request to URL: %s' % url, end ='') response = requests.get(url) print(' | Status code: %s' % response.status_code) return response # 新增postcode入参 def parse(self, html, postcode): content = BeautifulSoup(html, 'html.parser') titles = [title.text.strip() for title in content.findAll('h2', {'class': 'propertyCard-title'})] bedrooms = [title.text.split('bedroom')[0].strip() for title in content.findAll('h2', {'class': 'propertyCard-title'})] addresses = [address['content'] for address in content.findAll('meta', {'itemprop': 'streetAddress'})] descriptions = [description.text for description in content.findAll('span', {'data-test': 'property-description'})] prices = [price.text.strip() for price in content.findAll('div', {'class': 'propertyCard-priceValue'})] under_over = [underover.text.strip() for underover in content.findAll('div', {'class': 'propertyCard-priceQualifier'})] dates = [date.text for date in content.findAll('span', {'class': 'propertyCard-branchSummary-addedOrReduced'})] sellers = [seller.text.split('by')[-1].strip() for seller in content.findAll('span',{'class': 'propertyCard-branchSummary-branchName'})] for index in range(0, len(titles)): self.results.append({ # 新增邮编字段 'postcode': postcode, 'title': titles[index], 'no_of_bedrooms' : bedrooms[index], 'address': addresses[index], 'description': descriptions[index], 'price': prices[index], 'under_over': under_over[index], 'date': dates[index], 'seller': sellers[index]}) def to_csv(self): with open('rightmove_data.csv','w', encoding='utf-8-sig', newline='') as csv_file: writer = csv.DictWriter(csv_file,fieldnames=self.results[0].keys()) writer.writeheader() for row in self.results: writer.writerow(row) print('Stored results to "rightmove_data.csv"') def run(self): # 逐行遍历同时取URL和对应邮编,itertuples遍历效率高于iterrows for row in data.itertuples(index=False): current_url = row.url_name current_postcode = row.postcode response = self.fetch(current_url) # 爬取成功再解析,避免报错 if response.status_code == 200: self.parse(response.text, current_postcode) self.to_csv() if __name__ == '__main__': # 替换成你自己的数据集读取逻辑即可 # data = pd.read_csv('your_source_data_path.csv') scraper = RightmoveScraper() scraper.run()
补充说明
- 给文件写入加了
encoding='utf-8-sig'和newline=''参数,避免导出的csv出现乱码、空行问题 - 加了响应状态码判断,只有请求成功才会进入解析逻辑,减少爬取过程中的报错概率
- 如果你的
data不是pandas的DataFrame格式,只需要保证循环时能同步拿到每行的URL和对应postcode即可,核心传参逻辑不变
内容的提问来源于stack exchange,提问作者Mensa 23
相关产品推荐
相关产品推荐

