You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取列表存储的链接对应企业信息并汇总存储为表格?

问题根因

  • 未模拟浏览器请求头,网站拦截批量非浏览器请求
  • 请求频率过高,触发网站反爬限流机制
  • 页面信息提取规则适配性差,无法兼容不同页面的结构差异
  • 未做异常捕获,单个请求出错后直接中断整个循环

可运行实现代码

import requests
import re
import time
import pandas as pd
from bs4 import BeautifulSoup

# 配置请求头模拟浏览器
HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept-Language': 'de-DE,de;q=0.9,en-US;q=0.8,en;q=0.7'
}
# 请求间隔(秒),避免触发限流
REQUEST_DELAY = 2

# 待爬取链接列表
URL_LIST = [
    'https://allianz-entwicklung-klima.de/kompensationspartner/aera-group/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/atmosfair-ggmbh/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/bischoff-ditze-energy-gmbh-co-kg/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/climate-extender-gmbh/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/climatepartner-gmbh/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/die-klimamanufaktur-gmbh/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/die-ofenmacher-e-v/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/first-climate/',
    'https://allianz-entwicklung-klima.de/kompensationspartner/fokus-zukunft-gmbh-co-kg/'
]

# 正则匹配规则
EMAIL_PATTERN = re.compile(r'[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+')
PHONE_PATTERN = re.compile(r'\+49\s?\d+[\s\d/-]{5,}')
# 德国地址匹配:五位邮编 + 城市名
ADDRESS_PATTERN = re.compile(r'\d{5}\s+[a-zA-ZäöüÄÖÜß\s-]+', re.IGNORECASE)

def extract_company_info(url):
    try:
        # 发送请求
        resp = requests.get(url, headers=HEADERS, timeout=10)
        resp.raise_for_status()
        resp.encoding = 'utf-8'
        soup = BeautifulSoup(resp.text, 'html.parser')
        
        # 提取企业名称(默认取h1标题)
        company_name = soup.find('h1').get_text(strip=True) if soup.find('h1') else ''
        
        # 提取页面所有文本用于正则匹配
        page_text = soup.get_text(separator=' ', strip=True)
        
        # 匹配邮箱
        email_match = EMAIL_PATTERN.search(page_text)
        email = email_match.group() if email_match else ''
        
        # 匹配电话
        phone_match = PHONE_PATTERN.search(page_text)
        phone = phone_match.group().strip() if phone_match else ''
        
        # 匹配地址
        address_match = ADDRESS_PATTERN.search(page_text)
        address = address_match.group().strip() if address_match else ''
        
        return {
            '企业名称': company_name,
            '地址': address,
            '联系电话': phone,
            '邮箱': email
        }
    except Exception as e:
        print(f"链接{url}爬取失败:{str(e)}")
        return {
            '企业名称': '',
            '地址': '',
            '联系电话': '',
            '邮箱': ''
        }

if __name__ == '__main__':
    result_list = []
    for idx, url in enumerate(URL_LIST):
        print(f"正在爬取第{idx+1}个链接:{url}")
        info = extract_company_info(url)
        info['来源链接'] = url
        result_list.append(info)
        # 仅非最后一个请求加延迟
        if idx != len(URL_LIST)-1:
            time.sleep(REQUEST_DELAY)
    
    # 转成表格输出
    df = pd.DataFrame(result_list)
    # 打印表格
    print(df.to_string(index=False))
    # 保存为csv文件
    df.to_csv('企业信息汇总.csv', index=False, encoding='utf-8-sig')

使用说明

  1. 运行前先安装依赖:pip install requests beautifulsoup4 pandas
  2. 若部分字段提取为空,可打印对应页面的html内容,调整正则匹配规则或者增加css选择器定位逻辑,提升提取准确率
  3. 若触发反爬可适当调大REQUEST_DELAY参数,或者增加代理IP配置

内容的提问来源于stack exchange,提问作者maltinho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 18:39:01