You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫仅返回最后一项及DataFrame拆分邮编列求助

爬虫问题解决方案

问题1:循环爬取仅返回最后一项

你当前代码里,company_info在循环中会被每次迭代的内容覆盖,最后data = {company_info}只保留了最后一次循环的结果,而且用集合{}会导致数据去重丢失,正确的做法是用列表收集所有经销商信息:

import requests
from bs4 import BeautifulSoup
import pandas as pd

URL = "https://www.matki.co.uk/matki-dealers/"
page = requests.get(URL)
soup = BeautifulSoup(page.content, "html.parser")
results = soup.find(class_="dealer-overview") 
company_elements = results.find_all("article")

# 初始化列表存储所有经销商信息
company_list = []
for company_element in company_elements:
    company_info = company_element.getText(separator=u', ').replace('Find out more »', '')
    company_list.append(company_info)
    print(company_info)

# 用列表创建DataFrame,指定列名
df = pd.DataFrame(company_list, columns=['full_dealer_info'])
print(df.shape)
print(df)

问题2:拆分邮编到单独列

英国邮编格式通常为「字母+数字组合 数字+字母组合」(如SW1A 1AA),可以用正则表达式提取。在生成DataFrame后,添加新列存储提取出的邮编:

# 提取邮编,正则匹配英国邮编格式
df['postcode'] = df['full_dealer_info'].str.extract(r'([A-Z]{1,2}\d[A-Z\d]? \d[A-Z]{2})', expand=False)

# 查看结果
print(df[['full_dealer_info', 'postcode']])

完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

URL = "https://www.matki.co.uk/matki-dealers/"
page = requests.get(URL)
soup = BeautifulSoup(page.content, "html.parser")
results = soup.find(class_="dealer-overview") 
company_elements = results.find_all("article")

company_list = []
for company_element in company_elements:
    company_info = company_element.getText(separator=u', ').replace('Find out more »', '')
    company_list.append(company_info)

df = pd.DataFrame(company_list, columns=['full_dealer_info'])
# 提取邮编
df['postcode'] = df['full_dealer_info'].str.extract(r'([A-Z]{1,2}\d[A-Z\d]? \d[A-Z]{2})', expand=False)

print(df)

内容的提问来源于stack exchange,提问作者HRol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 11:55:24