如何提取列表存储的链接对应企业信息并汇总存储为表格?
问题根因
- 未模拟浏览器请求头,网站拦截批量非浏览器请求
- 请求频率过高,触发网站反爬限流机制
- 页面信息提取规则适配性差,无法兼容不同页面的结构差异
- 未做异常捕获,单个请求出错后直接中断整个循环
可运行实现代码
import requests import re import time import pandas as pd from bs4 import BeautifulSoup # 配置请求头模拟浏览器 HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'de-DE,de;q=0.9,en-US;q=0.8,en;q=0.7' } # 请求间隔(秒),避免触发限流 REQUEST_DELAY = 2 # 待爬取链接列表 URL_LIST = [ 'https://allianz-entwicklung-klima.de/kompensationspartner/aera-group/', 'https://allianz-entwicklung-klima.de/kompensationspartner/atmosfair-ggmbh/', 'https://allianz-entwicklung-klima.de/kompensationspartner/bischoff-ditze-energy-gmbh-co-kg/', 'https://allianz-entwicklung-klima.de/kompensationspartner/climate-extender-gmbh/', 'https://allianz-entwicklung-klima.de/kompensationspartner/climatepartner-gmbh/', 'https://allianz-entwicklung-klima.de/kompensationspartner/die-klimamanufaktur-gmbh/', 'https://allianz-entwicklung-klima.de/kompensationspartner/die-ofenmacher-e-v/', 'https://allianz-entwicklung-klima.de/kompensationspartner/first-climate/', 'https://allianz-entwicklung-klima.de/kompensationspartner/fokus-zukunft-gmbh-co-kg/' ] # 正则匹配规则 EMAIL_PATTERN = re.compile(r'[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+') PHONE_PATTERN = re.compile(r'\+49\s?\d+[\s\d/-]{5,}') # 德国地址匹配:五位邮编 + 城市名 ADDRESS_PATTERN = re.compile(r'\d{5}\s+[a-zA-ZäöüÄÖÜß\s-]+', re.IGNORECASE) def extract_company_info(url): try: # 发送请求 resp = requests.get(url, headers=HEADERS, timeout=10) resp.raise_for_status() resp.encoding = 'utf-8' soup = BeautifulSoup(resp.text, 'html.parser') # 提取企业名称(默认取h1标题) company_name = soup.find('h1').get_text(strip=True) if soup.find('h1') else '' # 提取页面所有文本用于正则匹配 page_text = soup.get_text(separator=' ', strip=True) # 匹配邮箱 email_match = EMAIL_PATTERN.search(page_text) email = email_match.group() if email_match else '' # 匹配电话 phone_match = PHONE_PATTERN.search(page_text) phone = phone_match.group().strip() if phone_match else '' # 匹配地址 address_match = ADDRESS_PATTERN.search(page_text) address = address_match.group().strip() if address_match else '' return { '企业名称': company_name, '地址': address, '联系电话': phone, '邮箱': email } except Exception as e: print(f"链接{url}爬取失败:{str(e)}") return { '企业名称': '', '地址': '', '联系电话': '', '邮箱': '' } if __name__ == '__main__': result_list = [] for idx, url in enumerate(URL_LIST): print(f"正在爬取第{idx+1}个链接:{url}") info = extract_company_info(url) info['来源链接'] = url result_list.append(info) # 仅非最后一个请求加延迟 if idx != len(URL_LIST)-1: time.sleep(REQUEST_DELAY) # 转成表格输出 df = pd.DataFrame(result_list) # 打印表格 print(df.to_string(index=False)) # 保存为csv文件 df.to_csv('企业信息汇总.csv', index=False, encoding='utf-8-sig')
使用说明
- 运行前先安装依赖:
pip install requests beautifulsoup4 pandas - 若部分字段提取为空,可打印对应页面的html内容,调整正则匹配规则或者增加css选择器定位逻辑,提升提取准确率
- 若触发反爬可适当调大
REQUEST_DELAY参数,或者增加代理IP配置
内容的提问来源于stack exchange,提问作者maltinho
相关产品推荐
相关产品推荐

