You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取li元素转字典时遇ValueError问题求助

解决抓取
  • 生成指定字典的错误
  • 错误原因

    你遇到的ValueError: dictionary update sequence element #0 has length 23; 2 is required,是因为直接将单个字符串组成的列表传入dict()函数。字典要求每个元素是长度为2的序列(如键值对元组),但你的代码生成的s列表中,每个元素是完整的属性文本(比如Property type: Apartment),并非键值对结构,因此无法直接转换为字典。

    修正后的代码

    import requests
    from bs4 import BeautifulSoup as bs
    import pandas as pd
    
    link = 'https://www.propertyfinder.eg/en/plp/rent/apartment-for-rent-cairo-hay-el-maadi-degla-street-207-3455087.html'
    # 补充headers示例,实际可根据需求调整
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
    r = requests.get(link, headers=headers)
    soup = bs(r.content, 'lxml')
    data = {}
    
    # 精准定位目标属性列表的ul(根据页面实际类名调整,避免抓取无关ul)
    target_ul = soup.find('ul', class_='pf-property-features__list')
    if target_ul:
        for li in target_ul.find_all('li'):
            li_text = li.get_text(strip=True)
            # 按第一个冒号分割键和值,避免值中含冒号导致分割错误
            if ':' in li_text:
                key, value = li_text.split(':', 1)
                key = key.strip()
                value = value.strip()
                # 将数值型值转为整数(如Bedrooms、Bathrooms)
                if value.isdigit():
                    value = int(value)
                data[key] = value
    
    # 转换为DataFrame
    df = pd.DataFrame([data])
    print(data)
    print(df)
    

    关键调整说明

    1. 精准定位目标UL:原代码soup.find('ul')会抓取页面第一个<ul>,可能不是属性列表,通过添加类名(如pf-property-features__list,需根据页面实际结构确认)确保抓取正确的元素。
    2. 分割键值对:对每个<li>的文本按第一个冒号分割,将内容拆分为属性名和属性值,避免值中含冒号时出现错误分割。
    3. 类型转换:将可转为整数的属性值(如卧室、浴室数量)转为数值类型,匹配你期望的字典格式。

    内容的提问来源于stack exchange,提问作者ahmed sayed

    相关产品推荐
    方舟 Agent Plan

    超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

    最近更新时间:2026.08.14 08:05:25