You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何提取Yelp爬虫所得a标签中>与<之间的用户名

问题原因

你获取到的name是BeautifulSoup的Tag类型对象,不是原始HTML字符串:

  1. 你尝试查找的&gt;是HTML转义后的显示内容,实际Tag对象内不存在该转义字符
  2. name.find[start+1:end]属于语法错误,find是方法而非可切片序列,不能用方括号取值

解决方案

直接调用Tag对象内置的get_text()方法即可提取标签包裹的文本,同时还可以优化冗余逻辑,修正后代码如下:

from selenium import webdriver
from bs4 import BeautifulSoup
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service as ChromeService

options = Options()
options.headless = True
options.add_experimental_option('excludeSwitches', ['enable-logging'])

CHROMEDRIVER_PATH ='../Selenium/chromedriver.exe'
service = ChromeService(executable_path=CHROMEDRIVER_PATH)
driver = webdriver.Chrome(service=service, options=options)
driver.get('https://www.yelp.com/biz/taste-of-texas-houston')

User = []
content = driver.page_source
# 显式指定解析器避免运行警告
soup = BeautifulSoup(content, 'lxml')
data = soup.findAll("li", attrs={'class':'margin-b5__09f24__pTvws border-color--default__09f24__NPAKY'})

for each in data:
    # 无需二次创建BeautifulSoup对象,直接在父标签上查询即可
    names = each.find_all("a", attrs={'class':'css-1422juy' , "href": lambda L: L and L.startswith("/user_details?userid=")})
    for name in names:
        # strip=True自动清理首尾多余的空格、换行符
        user_name = name.get_text(strip=True)
        User.append(user_name)
        # 若需要同时提取用户ID可取消下一行注释
        # user_id = name['href'].split('userid=')[-1]

# 所有用户名已存入User列表,可直接使用
print(User)

# 运行结束关闭驱动,避免后台残留进程
driver.quit()

内容的提问来源于stack exchange,提问作者Slavisha84

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 14:24:00