Python开发:统计表格指定caseref行上方'Unallocated'次数并截断表格
Solution for Scraping Internal Table & Counting 'Unallocated' Entries
I’ll walk you through a Python-based solution tailored to your needs—since it’s flexible for internal page scraping and easy to tweak to match your specific table structure.
Step 1: Set Up Dependencies
First, install the tools we’ll need if you haven’t already:
pip install requests beautifulsoup4
Step 2: Full Code Implementation
This script fetches your internal page, parses the table, stops at the row containing the user-input caseref, and counts how many times 'Unallocated' appears in the second column above that row:
import requests from bs4 import BeautifulSoup def scrape_table_and_count_unallocated(internal_url, target_caseref): # Handle internal site authentication (adjust based on your site's requirements) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", # Add session cookies or auth tokens here (grab these from your browser's dev tools if needed) # "Cookie": "your_session_cookie_here" } try: # Fetch the internal page content response = requests.get(internal_url, headers=headers) response.raise_for_status() # Throw an error if the request fails (e.g., 404, 500) # Parse the HTML with BeautifulSoup soup = BeautifulSoup(response.text, 'html.parser') # Locate the target table (modify selector if your table has an ID/class) table = soup.find('table') if not table: raise ValueError("Could not find the table on the internal page") unallocated_count = 0 target_found = False # Loop through each row in the table for row in table.find_all('tr'): cells = row.find_all(['td', 'th']) # Handle both data and header cells if not cells: continue # Skip empty rows # Check if this row contains the target caseref (adjust column index as needed) # Example: assuming caseref is in the first column (index 0) caseref_text = cells[0].text.strip() if caseref_text == target_caseref.strip(): target_found = True break # Stop processing once we hit the target row # Count 'Unallocated' in the second column (index 1) if len(cells) >= 2: second_col_text = cells[1].text.strip() if second_col_text == 'Unallocated': unallocated_count += 1 if not target_found: print(f"Warning: Caseref '{target_caseref}' was not found in the table") return unallocated_count except Exception as e: print(f"Scraping failed: {str(e)}") return -1 # Example usage internal_page_url = "https://your-internal-site.com/table-page" user_caseref = input("Enter the caseref to stop at: ") count = scrape_table_and_count_unallocated(internal_page_url, user_caseref) if count != -1: print(f"Number of 'Unallocated' entries above the target caseref: {count}")
Key Adjustments for Your Environment:
- Authentication: Internal sites often require session cookies or auth tokens. Update the
headersdictionary with the values from your browser’s dev tools (look in the "Network" tab for request headers). - Table & Column Indices: If your table has a specific ID or class, replace
soup.find('table')withsoup.find('table', id='your-table-id'). Adjust the column indices (e.g.,cells[0]for caseref,cells[1]for the second column) to match your table’s structure. - Strict Matching: The script uses exact matches for caseref—if you need partial matches, change the check to
target_caseref.strip() in caseref_text.
Edge Cases to Handle:
- If the target caseref isn’t present, the script will warn you and return 0 counts.
- Rows with fewer than 2 cells are skipped to avoid index errors.
- If your table has header rows you want to ignore, add a check to skip rows where cells are
<th>elements.
内容的提问来源于stack exchange,提问作者Matthew Palmer
相关产品推荐
相关产品推荐

