You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取网站特定子类?提取球员姓名求助

Hey there! Let's tackle your two questions step by step.

1. Extracting Specific Child Elements with BeautifulSoup

BeautifulSoup gives you several straightforward ways to target specific child elements:

  • .find() / .find_all(): Use these to search for elements by tag name, class, ID, or other attributes. For example, to get all <p> tags inside a <div> with class content:
    content_div = soup.find('div', class_='content')
    specific_paragraphs = content_div.find_all('p', class_='highlight')
    
  • CSS Selectors with .select(): This is great for more complex hierarchies. For instance, to get the second <li> child of a <ul> with ID menu:
    target_li = soup.select_one('ul#menu li:nth-child(2)')
    
  • Direct Child Traversal: Use .children or .findChildren(recursive=False) (like you did in your code) to get only immediate children of an element, ignoring nested ones.
2. Fixing Your Code to Extract Only Player Names

Looking at your current code, you're skipping the first column entirely with [1:], but we need to dig into that first column to pull out just the player's name. Here's how to adjust it:

First, let's modify the loop to handle the first column separately, since it likely contains extra elements (like images or links) alongside the name. Assuming the player name is inside an <a> tag (common on soccer market value sites), here's the updated code:

import csv
import requests
from bs4 import BeautifulSoup

def few(urls, file):
    f = open(file, 'a', newline='', encoding="utf-8")
    writer = csv.writer(f)
    url = urls
    page = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'})
    soup = BeautifulSoup(page.content, 'lxml')
    tbody = soup('table', {"class": "items"})[0].find_all('tr')
    
    for row in tbody:
        # Skip header rows if needed (uncomment if your table has a header tr)
        # if row.find('th'):
        #     continue
        
        # Extract player name from the first column
        first_col = row.findChildren(recursive=False)[0]
        # Adjust the selector below based on the actual HTML structure
        # If the name is in a span with class 'player-name', use:
        # player_name = first_col.find('span', class_='player-name').text.strip()
        player_name = first_col.find('a').text.strip()
        
        # Process remaining columns as before, excluding unwanted indices
        other_cols = row.findChildren(recursive=False)[1:]
        exclude = [0, 3, 4, 6, 7, 8]
        filtered_cols = [ele.text.strip() for idx, ele in enumerate(other_cols) if idx not in exclude]
        
        # Combine name and filtered columns, then write to CSV
        writer.writerow([player_name] + filtered_cols)
    f.close()  # Don't forget to close the file!

Key Notes:

  • Adjust the selector for player name: If the name isn't in an <a> tag, right-click the player name on the website → "Inspect" to find its exact HTML element (e.g., a <span> with a specific class). Then update the first_col.find(...) line to match.
  • Skip header rows: If your table has a header row (with <th> tags), uncomment the check to skip it so you don't write header text into your CSV.
  • Close the file: Added f.close() to ensure your data is properly saved.

内容的提问来源于stack exchange,提问作者Gofos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 07:57:39