You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于BeautifulSoup4实现高效编码?——parse_table_data函数性能优化问询

Optimizing Your BeautifulSoup4 Code for Speed

First, let's look at how to speed up your existing BeautifulSoup implementation without switching libraries—these changes can cut down on unnecessary DOM traversals and redundant checks:

  1. Use a Faster Parser
    The built-in html.parser is reliable but slow. Switching to lxml (a C-based parser) will immediately speed up the initial parsing step. You'll need to install it first with pip install lxml.

  2. Combine Selectors to Reduce Traversals
    Instead of finding all scaledRoad--7fdfb divs first and then looping through each to find SVGs, use a combined CSS selector to target only the relevant SVGs directly. This minimizes the number of times you traverse the DOM.

  3. Optimize Condition Checks
    Replacing substring checks ("Banker" in name) with set membership (first_part in {'Banker', 'Player'}) is faster because set lookups are O(1). Also, splitting the name once and checking the first part avoids redundant string operations.

Here's the optimized BeautifulSoup code:

import typing
from bs4 import BeautifulSoup

def parse_table_data(self) -> typing.Union[dict, None]:
    page_source = self.driver.page_source
    # Use lxml for faster parsing
    soup = BeautifulSoup(page_source, "lxml")
    
    road_result_container = {"A": [], "B": [], "C": [], "D": [], "E": [], "F": []}
    column_letters = ["A", "B", "C", "D", "E", "F"]
    
    # Get all column divs in order (assuming they map to A-F)
    column_divs = soup.select("div.scaledRoad--7fdfb")
    
    for idx, div in enumerate(column_divs):
        if idx >= len(column_letters):
            break  # Handle unexpected extra divs gracefully
        current_column = column_letters[idx]
        
        # Target only relevant SVGs within this column
        for svg in div.select('svg.svg--34293[name][data-type="roadItem"]'):
            name = svg["name"]
            first_part = name.split()[0]
            if first_part in {"Banker", "Player"}:
                road_result_container[current_column].append(first_part)
    
    return road_result_container

Alternative Libraries for Maximum Speed

If you still need more performance after optimizing BeautifulSoup, consider using lxml directly (it's the engine behind BeautifulSoup's fast mode) or pyquery (which uses jQuery-style selectors). Both avoid the abstraction layer of BeautifulSoup, leading to faster DOM operations.

Example with lxml Directly:

import typing
from lxml import html

def parse_table_data(self) -> typing.Union[dict, None]:
    page_source = self.driver.page_source
    tree = html.fromstring(page_source)
    
    road_result_container = {"A": [], "B": [], "C": [], "D": [], "E": [], "F": []}
    column_letters = ["A", "B", "C", "D", "E", "F"]
    
    # XPath to get all scaledRoad divs
    column_divs = tree.xpath('//div[contains(@class, "scaledRoad--7fdfb")]')
    
    for idx, div in enumerate(column_divs):
        if idx >= len(column_letters):
            break
        current_column = column_letters[idx]
        
        # XPath to get relevant SVGs in this column
        svgs = div.xpath('.//svg[contains(@class, "svg--34293") and @name and @data-type="roadItem"]')
        for svg in svgs:
            name = svg.get("name")
            first_part = name.split()[0]
            if first_part in {"Banker", "Player"}:
                road_result_container[current_column].append(first_part)
    
    return road_result_container

Key Takeaways

  • Quick Win: Switch to the lxml parser with BeautifulSoup—this alone can reduce parsing time significantly.
  • Reduce DOM Hits: Use combined selectors to avoid unnecessary traversals.
  • Skip Abstraction: For maximum speed, use lxml directly instead of BeautifulSoup, as it cuts out the middle layer.

Content of the question originates from Stack Exchange, asked by Kate shim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 23:12:36