You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas.Series.apply()与BeautifulSoup时触发startswith属性错误

问题:使用BeautifulSoup提取DataFrame列中HTML文本时触发属性错误

DataFrame基本信息

<class 'pandas.core.frame.DataFrame'>
Index: 55778 entries, 0 to html
Data columns (total 1 columns):
 #   Column  Non-Null Count  Dtype 
---  ------  --------------  ----- 
 0   html    55778 non-null  object
dtypes: object(1)
memory usage: 2.9+ MB

实现代码

import bs4
import pandas as pd

def get_location_text(markup: str) -> str:
    selector = ".column-sm-2" # "div > ul > li"
    location = bs4.BeautifulSoup(markup, "lxml").select_one(selector)
    location_text = location.get_text(strip=True)
    return location_text

df_locations["text"] = df_locations["html"].apply(get_location_text)

错误信息

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
/var/folders/bl/b9g4bbkj381ghzzvws3tmcgm0000gn/T/ipykernel_50020/1441714440.py in <module>
----> 1 df_locations["text"] = df_locations["html"].apply(get_location_text)

~/bert2bert/venv/lib/python3.7/site-packages/pandas/core/series.py in apply(self, func, convert_dtype, args, **kwargs)
   4355         dtype: float64
   4356         """
-> 4357         return SeriesApply(self, func, convert_dtype, args, kwargs).apply()
   4358 
   4359     def _reduce(

~/bert2bert/venv/lib/python3.7/site-packages/pandas/core/apply.py in apply(self)
   1041             return self.apply_str()
   1042 
-> 1043         return self.apply_standard()
   1044 
   1045     def agg(self):

~/bert2bert/venv/lib/python3.7/site-packages/pandas/core/apply.py in apply_standard(self)
   1099                     values,
   1100                     f,  # type: ignore[arg-type]
-> 1101                     convert=self.convert_dtype,
   1102                 )
   1103 

~/bert2bert/venv/lib/python3.7/site-packages/pandas/_libs/lib.pyx in pandas._libs.lib.map_infer()

/var/folders/bl/b9g4bbkj381ghzzvws3tmcgm0000gn/T/ipykernel_50020/222721594.py in get_location_text(markup)
      1 def get_location_text(markup: str) -> str:
      2     selector = ".colcount-sm-2" # "div > ul > li"
----> 3     location = bs4.BeautifulSoup(markup, "lxml").select_one(selector)
      4     location_text = location.get_text(strip=True)
      5     return location_text

~/bert2bert/venv/lib/python3.7/site-packages/bs4/__init__.py in __init__(self, markup, features, builder, parse_only, from_encoding, exclude_encodings, element_classes, **kwargs)
    327          self.contains_replacement_characters) in (
    328              self.builder.prepare_markup(
-> 329                  markup, from_encoding, exclude_encodings=exclude_encodings)):
    330             self.reset()
    331             self.builder.initialize_soup(self)

~/bert2bert/venv/lib/python3.7/site-packages/bs4/builder/_lxml.py in prepare_markup(self, markup, user_specified_encoding, exclude_encodings, document_declared_encoding)
    180             # We're in HTML mode, so if we're given XML, that's worth
    181             # noting.
-> 182             DetectsXMLParsedAsHTML.warn_if_markup_looks_like_xml(markup)
    183         else:
    184             self.processing_instruction_class = XMLProcessingInstruction

~/bert2bert/venv/lib/python3.7/site-packages/bs4/builder/__init__.py in warn_if_markup_looks_like_xml(cls, markup)
    533 
    534         if (markup is not None
-> 535             and markup.startswith(prefix)
    536             and not looks_like_html.search(markup[:500])
    537         ):

~/bert2bert/venv/lib/python3.7/site-packages/pandas/core/generic.py in __getattr__(self, name)
   5485         ):
   5486             return self[name]
-> 5487         return object.__getattribute__(self, name)
   5488 
   5489     def __setattr__(self, name: str, value) -> None:

AttributeError: 'Series' object has no attribute 'startswith'

调试输出

执行df_locations["html"].describe():

count                                                 55778
unique                                                17077
top       <section class="mfp-hide modal-dialog modal-co...
freq                                                   3709
Name: html, dtype: object

执行df_locations["html"].head(2):

0    <section class="mfp-hide modal-dialog modal-co...
4    <section class="mfp-hide modal-dialog modal-co...
Name: html, dtype: object

执行type(df_locations["html"].iloc[0]):

<class 'str'>

解决方案

1. 根源分析

错误提示'Series' object has no attribute 'startswith',说明传入BeautifulSoup的markup参数不是字符串,而是Series对象。可能原因:

  • 数据列中混入了非字符串类型元素(虽首行是字符串,但不排除其他行存在异常)
  • 索引名html与列名html冲突,导致df_locations["html"]未正确取到目标列

2. 修复步骤

步骤1:排查数据类型异常

先检查列中所有元素的类型,找出非字符串行:

# 查看所有元素的类型
print(df_locations["html"].apply(type).unique())

# 筛选出非字符串的行
invalid_rows = df_locations[~df_locations["html"].apply(lambda x: isinstance(x, str))]
print(invalid_rows)

步骤2:修改提取函数,增加异常处理

修改函数确保输入为字符串,同时处理找不到目标元素的情况:

import bs4
import pandas as pd

def get_location_text(markup) -> str:
    # 过滤非字符串输入
    if not isinstance(markup, str):
        return ""  # 或根据需求返回pd.NA
    
    # 注意:使用正确的选择器,需匹配真实HTML结构(之前代码注释与实际写的选择器不一致)
    selector = ".column-sm-2"
    soup = bs4.BeautifulSoup(markup, "lxml")
    location = soup.select_one(selector)
    
    # 处理找不到元素的情况
    if location:
        return location.get_text(strip=True)
    else:
        return ""  # 或返回pd.NA

# 明确指定取列,避免索引与列名冲突
df_locations["text"] = df_locations.loc[:, "html"].apply(get_location_text)

步骤3:验证修复结果

执行后查看新生成的text列:

print(df_locations["text"].head())

内容的提问来源于stack exchange,提问作者node_env

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 23:31:00