Scrapy结合MySQL存储数据报错:TypeError: connect()参数应为字符串而非None
Hey there! Let's break down why you're hitting this error and fix it step by step. That TypeError means one of the arguments passed to MySQLdb.connect() is None instead of a required string/value—almost certainly because your Scrapy settings aren't properly configured for MySQL, or there are small bugs in your pipeline code.
1. First: Check Your settings.py MySQL Config
This is the most common culprit. You need to add these MySQL connection details to your project's settings.py file, and make sure the names match exactly what you're using in your pipeline:
# settings.py MYSQL_HOST = 'localhost' # Or your remote database host MYSQL_DBNAME = 'your_database_name' # Name of your MySQL database MYSQL_PORT = 3306 # Keep this as an integer, not a string MYSQL_USER = 'your_db_username' MYSQL_PASSWD = 'your_db_password'
Double-check for typos (like using MYSQL_PASSWORD instead of MYSQL_PASSWD—your pipeline references MYSQL_PASSWD, so the setting name must match perfectly!).
2. Fix Pipeline Code Issues
Your pipeline has a couple of small issues that could cause problems (or break entirely in Python 3):
a. Add Safety Checks for Missing Settings
Modify the from_settings method to catch missing config values before they trigger a connection error:
@classmethod def from_settings(cls, settings): dbargs = dict( host=settings.get('MYSQL_HOST', 'localhost'), db=settings.get('MYSQL_DBNAME'), port=settings.get('MYSQL_PORT', 3306), user=settings.get('MYSQL_USER'), passwd=settings.get('MYSQL_PASSWD'), charset='utf8', cursorclass=MySQLdb.cursors.DictCursor, use_unicode=True, ) # Ensure required settings aren't missing required_settings = ['db', 'user', 'passwd'] for setting in required_settings: if dbargs[setting] is None: raise ValueError(f"Missing required MySQL setting: {setting}") dbpool = adbapi.ConnectionPool('MySQLdb', **dbargs) return cls(dbpool)
b. Fix Python 3 Print Syntax
Your current print statements use Python 2 syntax (no parentheses), which will throw errors in Python 3. Update them:
if result: print("added a model into db") else: print("failed insert into pricing")
c. Correct Item Validation Logic
Your current item check loops over field names instead of their actual values. Update it to properly verify that field values aren't empty:
def _do_upinsert(self, conn, item, spider): valid = True for field_name in item.fields: field_value = item.get(field_name) if field_value is None or not str(field_value).strip(): valid = False spider.logger.warning(f"Missing or empty value for field: {field_name}") break if valid: result = conn.execute(""" insert into pricing(model, price) values(%s, %s) """, (item['model'], item['price'])) if result: print("added a model into db") else: print("failed insert into pricing")
3. Make Sure Your Pipeline Is Enabled
Don't forget to enable your pipeline in settings.py! Add this if it's missing:
ITEM_PIPELINES = { 'scraper.pipelines.ScraperPipeline': 300, }
The number 300 is the priority (lower numbers run first)—adjust as needed if you have other pipelines in your project.
4. Test MySQL Connection Separately
To rule out database-side issues, run a quick test script to confirm you can connect to MySQL with your credentials:
import MySQLdb try: conn = MySQLdb.connect( host='localhost', db='your_database_name', port=3306, user='your_db_username', passwd='your_db_password', charset='utf8' ) print("MySQL connection successful!") conn.close() except Exception as e: print(f"Connection failed: {str(e)}")
If this fails, fix your MySQL credentials/host configuration first before troubleshooting Scrapy further.
内容的提问来源于stack exchange,提问作者mark o'reilly

