You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:如何用Python(BeautifulSoup)在bb farma批量搜产品并提取ATC编码

Got it, let's walk through how to implement this with Python and BeautifulSoup—this should cover your needs of searching each product, extracting ATC codes from the table, handling exceptions, and saving the results.

Step 1: Set Up Dependencies

First, make sure you have the required libraries installed. If you haven't already:

pip install beautifulsoup4 requests  # requests is optional if you already have all HTML files saved locally

Step 2: Core Implementation

Here's a complete, commented script that handles your workflow. I've included options for both loading pre-saved HTML files and fetching pages in real-time (in case you need that later):

from bs4 import BeautifulSoup
import csv
import requests

# ----------------------
# Configuration
# ----------------------
product_list = ["Product A", "Product B", "Product C"]  # Replace with your actual product list
output_file = "atc_code_results.csv"

# Optional: Function to fetch search page HTML in real-time (replace with actual BB Farma search URL)
def fetch_search_page(product):
    search_url = f"https://www.bbfarma.com/search?query={product}"  # Update this to the real search endpoint
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    try:
        response = requests.get(search_url, headers=headers)
        response.raise_for_status()  # Raise error for HTTP status codes >=400
        return response.text
    except requests.exceptions.RequestException as e:
        print(f"Failed to fetch page for {product}: {str(e)}")
        return None

# ----------------------
# Main Processing Logic
# ----------------------
results = []

for product in product_list:
    try:
        # Option 1: Fetch HTML in real-time
        html_content = fetch_search_page(product)
        
        # Option 2: Load from local file (uncomment if using pre-saved HTML)
        # with open(f"{product}_search.html", "r", encoding="utf-8") as f:
        #     html_content = f.read()

        if not html_content:
            results.append({"product": product, "ATC_code": "Failed to retrieve page"})
            continue

        # Parse HTML with BeautifulSoup
        soup = BeautifulSoup(html_content, "html.parser")

        # Locate the drug list table (adjust selector to match actual page structure)
        # Use browser dev tools to find the table's class/id, e.g., <table class="drug-results-table">
        drug_table = soup.find("table", class_="drug-results-table")
        if not drug_table:
            results.append({"product": product, "ATC_code": "No results table found"})
            print(f"No table detected for {product}")
            continue

        # Find which column contains ATC codes (using table headers)
        headers = [th.text.strip() for th in drug_table.find("tr").find_all("th")]
        try:
            atc_col_index = headers.index("ATC")  # Replace with actual header text (e.g., localized labels like "ATC-Kod")
        except ValueError:
            results.append({"product": product, "ATC_code": "ATC column not found in table"})
            print(f"ATC column missing for {product}")
            continue

        # Extract ATC codes from all rows (skip header row)
        atc_codes = []
        rows = drug_table.find_all("tr")[1:]  # Skip first row (headers)
        for row in rows:
            cells = row.find_all("td")
            if len(cells) > atc_col_index:
                atc_code = cells[atc_col_index].text.strip()
                if atc_code:  # Ignore empty values
                    atc_codes.append(atc_code)

        # Save results (join multiple codes if present, or note none found)
        results.append({
            "product": product,
            "ATC_code": ", ".join(atc_codes) if atc_codes else "No ATC codes found"
        })
        print(f"Successfully processed: {product}")

    except Exception as e:
        # Catch-all for unexpected errors
        error_msg = f"Processing failed: {str(e)}"
        results.append({"product": product, "ATC_code": error_msg})
        print(f"Error with {product}: {error_msg}")

# ----------------------
# Save Results to CSV
# ----------------------
with open(output_file, "w", newline="", encoding="utf-8") as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=["product", "ATC_code"])
    writer.writeheader()
    writer.writerows(results)

print(f"\nAll done! Results saved to {output_file}")

Key Notes to Adjust for Your Use Case:

  • Table Selector: Use your browser's dev tools (right-click → Inspect) to find the exact class or id of the results table. Replace "drug-results-table" with the actual value from BB Farma's page.
  • ATC Header Text: Check the table's header row—if the ATC column is labeled something else (like "ATC Code" or a localized version), update the headers.index("ATC") line to match.
  • HTML Loading: Choose between fetching pages in real-time (Option 1) or loading local files (Option 2) based on how you're storing the HTML.

Exception Handling Details:

The script uses multiple layers of error handling:

  • Catches HTTP errors if fetching pages live
  • Checks for missing tables or ATC columns
  • Uses a catch-all for unexpected issues (like malformed HTML)
  • Logs errors clearly so you can debug specific products

内容的提问来源于stack exchange,提问作者Lusian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:53:12