使用grequests从API及SharePoint批量获取CSV文件的技术咨询
Got it, let's tackle this SharePoint bulk CSV download issue alongside your grequests workflow—here are a few practical approaches you can implement right away:
Approach 1: Fetch File List First, Then Concurrent Downloads
Since you already know how to pull individual CSVs, the simplest way to batch is to first get all CSV files in your target folder, then use grequests to fetch them in parallel.
Step-by-Step Breakdown:
- Retrieve the folder's file list
Call SharePoint's API to get metadata for all files in theData Sourcesfolder. This gives you the details needed to build download URLs for each CSV. - Filter for CSV files
From the file list, extract only entries where the filename ends with.csv. - Concurrent download with grequests
Build a download URL for each CSV, then use grequests to send all requests at once (with proper rate limiting to avoid SharePoint throttling).
Example Code:
import grequests # Reuse your existing auth headers headers = { 'Authorization': 'Bearer eyJ...', 'Accept': 'application/json;odata=verbose' } # 1. Get all files in the target folder folder_url = "https://mycompany.sharepoint.com/teams/a/g/_api/web/GetFolderByServerRelativeUrl('Data%20Sources')/Files" folder_resp = grequests.get(folder_url, headers=headers).send().response files_metadata = folder_resp.json() # 2. Filter CSV files and build download URLs csv_download_urls = [] for file in files_metadata['d']['results']: if file['Name'].lower().endswith('.csv'): # Use ServerRelativeUrl for more reliable path handling (avoids special character issues) download_url = f"https://mycompany.sharepoint.com/teams/a/g/_api/web/GetFileByServerRelativeUrl('{file['ServerRelativeUrl']}')/$value" csv_download_urls.append(download_url) # 3. Fetch all CSVs concurrently (limit to 5 concurrent requests to avoid throttling) requests = [grequests.get(url, headers=headers) for url in csv_download_urls] responses = grequests.map(requests, size=5) # Process each CSV response for idx, resp in enumerate(responses): if resp and resp.status_code == 200: csv_content = resp.text print(f"Successfully fetched {csv_download_urls[idx]}") # Save to local file or process directly # with open(f"sharepoint_csv_{idx}.csv", 'w', encoding='utf-8') as f: # f.write(csv_content) else: status_code = resp.status_code if resp else "Request failed" print(f"Failed to fetch {csv_download_urls[idx]}: {status_code}")
Approach 2: Use SharePoint OData Batch Requests (For Large File Sets)
If you have dozens/hundreds of CSVs, using OData batch requests reduces the number of HTTP connections by packing multiple download requests into one. This is more efficient but requires constructing a multipart request body.
Example Code Snippet:
import grequests import uuid # Generate a unique boundary for the batch request batch_boundary = f"batch_{uuid.uuid4().hex}" headers = { 'Authorization': 'Bearer eyJ...', 'Content-Type': f'multipart/mixed; boundary={batch_boundary}' } # First, get the CSV file list (same as Approach 1) folder_url = "https://mycompany.sharepoint.com/teams/a/g/_api/web/GetFolderByServerRelativeUrl('Data%20Sources')/Files" folder_resp = grequests.get(folder_url, headers={'Authorization': 'Bearer eyJ...', 'Accept': 'application/json;odata=verbose'}).send().response csv_files = [f for f in folder_resp.json()['d']['results'] if f['Name'].lower().endswith('.csv')] # Build the batch request body batch_body_parts = [] for file in csv_files: req_path = f"/teams/a/g/_api/web/GetFileByServerRelativeUrl('{file['ServerRelativeUrl']}')/$value" batch_body_parts.extend([ f"--{batch_boundary}", "Content-Type: application/http", "Content-Transfer-Encoding: binary", "", f"GET {req_path} HTTP/1.1", "Accept: text/csv", "", "" ]) batch_body_parts.append(f"--{batch_boundary}--") batch_body = "\r\n".join(batch_body_parts) # Send the batch request batch_url = "https://mycompany.sharepoint.com/teams/a/g/_api/$batch" batch_resp = grequests.post(batch_url, headers=headers, data=batch_body).send().response # Parse the multipart response response_parts = batch_resp.text.split(f"--{batch_boundary}") for part in response_parts[1:-1]: # Skip empty first part and closing boundary if "HTTP/1.1 200 OK" in part: # Extract CSV content (skip response headers) content_start = part.find("\r\n\r\n") + 4 csv_content = part[content_start:].strip() print(f"CSV content:\n{csv_content[:100]}...") # Preview first 100 chars
Key Notes to Avoid Issues:
- Throttling: SharePoint enforces rate limits—use the
sizeparameter ingrequests.map()to cap concurrent requests (5-10 is safe for most tenants). - Permissions: Ensure your Bearer token has read access to all files in the folder (check for 403 errors if some files fail to download).
- Encoding: Some CSVs might use non-UTF-8 encoding (like GBK). If you see garbled text, try specifying the encoding when reading/writing files (e.g.,
encoding='gbk'). - Error Handling: Add an
error_callbacktogrequests.map()to catch timeouts or network failures gracefully.
内容的提问来源于stack exchange,提问作者user3871
相关产品推荐
相关产品推荐

