You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取求助:仅获取title字段,其余字段返回空值

修复爬虫无法获取除title外其他字段的问题

我帮你看了下你的代码,发现几个导致其他字段为空的核心问题,咱们一步步来修复:

问题分析

  • 没有返回收集的数据:get_detail_data函数只在内部处理了字段,但没有把结果返回出来,所以你根本看不到后续字段的取值情况
  • 元素定位不准确:页面中class="col-xs-12 col-sm-4"的div有多个,直接用find()只会拿到第一个匹配的元素,而这个元素大概率不是你要找的地址或电话所在的容器
  • 字段取值不完整:比如add只找到了strong标签但没提取文本;phone直接取整个div的文本会包含大量无关内容,需要更精准的定位

修复后的完整代码

import requests
from bs4 import BeautifulSoup

def get_page(url):
    response = requests.get(url)
    if not response.ok:
        print('server responded:', response.status_code)
        return None
    else:
        soup = BeautifulSoup(response.text, 'html.parser')
        return soup

def get_detail_data(soup):
    # 初始化结果字典
    data = {}
    
    # 获取title
    try:
        title = soup.find('span', class_="text-info h4", id=False).find('strong').text.strip()
        data['title'] = title
    except Exception as e:
        data['title'] = 'empty'
        print(f"获取title出错: {e}")
    
    # 获取地址:通过后续的文本内容或更精准的父容器定位
    try:
        # 找到包含"Address"的标签,再定位对应的内容
        address_label = soup.find('strong', string="Address:")
        if address_label:
            add = address_label.find_next_sibling(text=True).strip()
            data['add'] = add
        else:
            data['add'] = 'empty add'
    except Exception as e:
        data['add'] = 'empty add'
        print(f"获取地址出错: {e}")
    
    # 获取电话:同样通过标签文本定位
    try:
        phone_label = soup.find('strong', string="Phone:")
        if phone_label:
            phone = phone_label.find_next_sibling(text=True).strip()
            data['phone'] = phone
        else:
            data['phone'] = 'empty phone'
    except Exception as e:
        data['phone'] = 'empty phone'
        print(f"获取电话出错: {e}")
    
    # 返回收集到的数据
    return data

def main():
    url = "https://www.dobsearch.com/people-finder/view.php?searchnum=287404084791&sessid=vusqgp50pm8r38lfe13la8ta1l"
    soup = get_page(url)
    if soup:
        detail_data = get_detail_data(soup)
        print("爬取结果:")
        for key, value in detail_data.items():
            print(f"{key}: {value}")

if __name__ == '__main__':
    main()

关键修复点说明

  1. 返回数据:给get_detail_data添加了返回值,用字典存储所有字段,方便后续使用或保存
  2. 精准定位元素:不再依赖容易重复的class,而是通过标签的文本内容(比如"Address:"、"Phone:")来定位目标字段,这种方式更稳定,不容易因为页面布局小改动失效
  3. 完善异常处理:给每个异常捕获添加了错误信息打印,方便你排查问题
  4. 文本处理:使用strip()去除文本前后的空格和换行符,让结果更整洁

现在运行代码,应该就能正常获取到title、地址和电话字段的内容了。

内容的提问来源于stack exchange,提问作者M.Akram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:57:41