技术问询:如何用Python(BeautifulSoup)在bb farma批量搜产品并提取ATC编码
Got it, let's walk through how to implement this with Python and BeautifulSoup—this should cover your needs of searching each product, extracting ATC codes from the table, handling exceptions, and saving the results.
Step 1: Set Up Dependencies
First, make sure you have the required libraries installed. If you haven't already:
pip install beautifulsoup4 requests # requests is optional if you already have all HTML files saved locally
Step 2: Core Implementation
Here's a complete, commented script that handles your workflow. I've included options for both loading pre-saved HTML files and fetching pages in real-time (in case you need that later):
from bs4 import BeautifulSoup import csv import requests # ---------------------- # Configuration # ---------------------- product_list = ["Product A", "Product B", "Product C"] # Replace with your actual product list output_file = "atc_code_results.csv" # Optional: Function to fetch search page HTML in real-time (replace with actual BB Farma search URL) def fetch_search_page(product): search_url = f"https://www.bbfarma.com/search?query={product}" # Update this to the real search endpoint headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: response = requests.get(search_url, headers=headers) response.raise_for_status() # Raise error for HTTP status codes >=400 return response.text except requests.exceptions.RequestException as e: print(f"Failed to fetch page for {product}: {str(e)}") return None # ---------------------- # Main Processing Logic # ---------------------- results = [] for product in product_list: try: # Option 1: Fetch HTML in real-time html_content = fetch_search_page(product) # Option 2: Load from local file (uncomment if using pre-saved HTML) # with open(f"{product}_search.html", "r", encoding="utf-8") as f: # html_content = f.read() if not html_content: results.append({"product": product, "ATC_code": "Failed to retrieve page"}) continue # Parse HTML with BeautifulSoup soup = BeautifulSoup(html_content, "html.parser") # Locate the drug list table (adjust selector to match actual page structure) # Use browser dev tools to find the table's class/id, e.g., <table class="drug-results-table"> drug_table = soup.find("table", class_="drug-results-table") if not drug_table: results.append({"product": product, "ATC_code": "No results table found"}) print(f"No table detected for {product}") continue # Find which column contains ATC codes (using table headers) headers = [th.text.strip() for th in drug_table.find("tr").find_all("th")] try: atc_col_index = headers.index("ATC") # Replace with actual header text (e.g., localized labels like "ATC-Kod") except ValueError: results.append({"product": product, "ATC_code": "ATC column not found in table"}) print(f"ATC column missing for {product}") continue # Extract ATC codes from all rows (skip header row) atc_codes = [] rows = drug_table.find_all("tr")[1:] # Skip first row (headers) for row in rows: cells = row.find_all("td") if len(cells) > atc_col_index: atc_code = cells[atc_col_index].text.strip() if atc_code: # Ignore empty values atc_codes.append(atc_code) # Save results (join multiple codes if present, or note none found) results.append({ "product": product, "ATC_code": ", ".join(atc_codes) if atc_codes else "No ATC codes found" }) print(f"Successfully processed: {product}") except Exception as e: # Catch-all for unexpected errors error_msg = f"Processing failed: {str(e)}" results.append({"product": product, "ATC_code": error_msg}) print(f"Error with {product}: {error_msg}") # ---------------------- # Save Results to CSV # ---------------------- with open(output_file, "w", newline="", encoding="utf-8") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["product", "ATC_code"]) writer.writeheader() writer.writerows(results) print(f"\nAll done! Results saved to {output_file}")
Key Notes to Adjust for Your Use Case:
- Table Selector: Use your browser's dev tools (right-click → Inspect) to find the exact
classoridof the results table. Replace"drug-results-table"with the actual value from BB Farma's page. - ATC Header Text: Check the table's header row—if the ATC column is labeled something else (like "ATC Code" or a localized version), update the
headers.index("ATC")line to match. - HTML Loading: Choose between fetching pages in real-time (Option 1) or loading local files (Option 2) based on how you're storing the HTML.
Exception Handling Details:
The script uses multiple layers of error handling:
- Catches HTTP errors if fetching pages live
- Checks for missing tables or ATC columns
- Uses a catch-all for unexpected issues (like malformed HTML)
- Logs errors clearly so you can debug specific products
内容的提问来源于stack exchange,提问作者Lusian
相关产品推荐
相关产品推荐

