SeaweedFS新增Volume Server后无法上传文件问题求助
问题描述
- 部署架构:三台独立机器分别部署master(x.x.x.61)、volume(x.x.x.63)、filer+s3API(x.x.x.62)
- 核心问题:原volume服务器(x.x.x.63)存储空间耗尽,新增volume服务器(x.x.x.64)后,仍无法通过filer UI(http://x.x.x.62:8888)上传新文件
- filer报错日志(系统仍尝试写入已耗尽空间的原volume服务器):
E1221 11:09:48.027930 upload_content.go:351 unmarshal http://x.x.x.63:8080/7,2bafadaa4666: {"error":"failed to write to local disk: write data/chrisDir_7.dat: no space left on device"}{"name":"app_progress4.apk","size":2353734,"eTag":"92b10892"} W1221 11:09:48.027950 upload_content.go:168 uploading 2 to http://x.x.x.63:8080/7,2bafadaa4666: unmarshal http://x.x.x.63:8080/7,2bafadaa4666: invalid character '{' after top-level value E1221 11:09:48.027965 filer_server_handlers_write_upload.go:209 upload error: unmarshal http://x.x.x.63:8080/7,2bafadaa4666: invalid character '{' after top-level value I1221 11:09:48.028022 common.go:70 response method:POST URL:/buckets/chrisDir/ with httpStatus:500 and JSON:{"error":"unmarshal http://x.x.x.63:8080/2,2ba84b2894a7: invalid character '{' after top-level value"}
- master日志确认新volume节点已添加,rebalance脚本已执行:
I1221 11:36:09.522690 node.go:225 topo:DefaultDataCenter:DefaultRack adds child x.x.x.64:8080 I1221 11:36:09.522716 node.go:225 topo:DefaultDataCenter:DefaultRack:x.x.x.64:8080 adds child I1221 11:36:09.522724 master_grpc_server.go:138 added volume server 0: x.x.x.64:8080 [3caad049-38a6-43f6-8192-d1082c5e838b] I1221 11:36:09.522744 master_grpc_server.go:49 found new uuid:x.x.x.64:8080 [3caad049-38a6-43f6-8192-d1082c5e838b] , map[x.x.x.63:8080:[5005b287-c812-4dba-ba41-9b5a6a022f12] x.x.x.64:8080:[3caad049-38a6-43f6-8192-d1082c5e838b]] I1221 11:36:09.522866 volume_layout.go:393 Volume 11 becomes writable I1221 11:36:09.522880 master_grpc_server.go:199 master see new volume 11 from x.x.x.64:8080 I1221 11:38:33.481721 master_server.go:323 executing: lock [] I1221 11:38:33.482821 master_server.go:323 executing: ec.encode [-fullPercent=95 -quietFor=1h] I1221 11:38:33.483925 master_server.go:323 executing: ec.rebuild [-force] I1221 11:38:33.484372 master_server.go:323 executing: ec.balance [-force] I1221 11:38:33.484777 master_server.go:323 executing: volume.balance [-force] 2022/12/21 11:38:48 copying volume 21 from x.x.x.63:8080 to x.x.x.64:8080 I1221 11:38:48.486778 volume_layout.go:407 Volume 21 has 0 replica, less than required 1 I1221 11:38:48.486798 volume_layout.go:380 Volume 21 becomes unwritable I1221 11:38:48.494998 volume_layout.go:393 Volume 21 becomes writable 2022/12/21 11:38:48 tailing volume 21 from x.x.x.63:8080 to x.x.x.64:8080 2022/12/21 11:38:58 deleting volume 21 from x.x.x.63:8080 ....
- 各组件启动命令:
- Master:
./weed master -mdir='.' - Volume:
./weed volume -max=100 -mserver="x.x.x.61:9333" -dir="$dataDir" - Filer和S3:
./weed filer -master="x.x.x.61:9333" -s3
- Master:
- $HOME/.seaweedfs目录内容:
drwxrwxr-x 2 seaweedfs seaweedfs 4096 Dec 20 16:01 . drwxr-xr-x 20 seaweedfs seaweedfs 4096 Dec 20 16:01 .. -rw-r--r-- 1 seaweedfs seaweedfs 2234 Dec 20 15:57 master.toml
- master.toml配置:
# Put this file to one of the location, with descending priority # ./master.toml # $HOME/.seaweedfs/master.toml # /etc/seaweedfs/master.toml # this file is read by master [master.maintenance] # periodically run these scripts are the same as running them from 'weed shell' scripts = """ lock ec.encode -fullPercent=95 -quietFor=1h ec.rebuild -force ec.balance -force volume.deleteEmpty -quietFor=24h -force volume.balance -force volume.fix.replication s3.clean.uploads -timeAgo=24h unlock """ sleep_minutes = 7 # sleep minutes between each script execution [master.sequencer] type = "raft" # Choose [raft|snowflake] type for storing the file id sequence # when sequencer.type = snowflake, the snowflake id must be different from other masters sequencer_snowflake_id = 0 # any number between 1~1023 # configurations for tiered cloud storage # old volumes are transparently moved to cloud for cost efficiency [storage.backend] [storage.backend.s3.default] enabled = false aws_access_key_id = "" # if empty, loads from the shared credentials file (~/.aws/credentials). aws_secret_access_key = "" # if empty, loads from the shared credentials file (~/.aws/credentials). region = "us-east-2" bucket = "your_bucket_name" # an existing bucket endpoint = "" storage_class = "STANDARD_IA" # create this number of logical volumes if no more writable volumes # count_x means how many copies of data. # e.g.: # 000 has only one copy, copy_1 # 010 and 001 has two copies, copy_2 # 011 has only 3 copies, copy_3 [master.volume_growth] copy_1 = 7 # create 1 x 7 = 7 actual volumes copy_2 = 6 # create 2 x 6 = 12 actual volumes copy_3 = 3 # create 3 x 3 = 9 actual volumes copy_other = 1 # create n x 1 = n actual volumes # configuration flags for replication [master.replication] # any replication counts should be considered minimums. If you specify 010 and # have 3 different racks, that's still considered writable. Writes will still # try to replicate to all available volumes. You should only use this option # if you are doing your own replication or periodic sync of volumes. treat_replication_as_minimums = false
- 系统状态信息:
curl http://localhost:9333/dir/assign?pretty=y { "fid": "9,2bb2fd75d706", "url": "x.x.x.63:8080", "publicUrl": "x.x.x.63:8080", "count": 1 } curl http://x.x.x.61:9333/cluster/status?pretty=y { "IsLeader": true, "Leader": "x.x.x.61:9333", "MaxVolumeId": 21 } curl "http://x.x.x.61:9333/dir/status?pretty=y" { "Topology": { "Max": 200, "Free": 179, "DataCenters": [ { "Id": "DefaultDataCenter", "Racks": [ { "Id": "DefaultRack", "DataNodes": [ { "Url": "x.x.x.63:8080", "PublicUrl": "x.x.x.63:8080", "Volumes": 20, "EcShards": 0, "Max": 100, "VolumeIds": " 1-10 12-21" }, { "Url": "x.x.x.64:8080", "PublicUrl": "x.x.x.64:8080", "Volumes": 1, "EcShards": 0, "Max": 100, "VolumeIds": " 11" } ] } ] } ], "Layouts": [ { "replication": "000", "ttl": "", "writables": [ 6, 1, 2, 7, 3, 4, 5 ], "collection": "chrisDir" }, { "replication": "000", "ttl": "", "writables": [ 16, 19, 17, 21, 15, 18, 20 ], "collection": "chrisDir2" }, { "replication": "000", "ttl": "", "writables": [ 8, 12, 13, 9, 14, 10, 11 ], "collection": "" } ] }, "Version": "30GB 3.37 438146249f50bf36b4c46ece02a430f44152777f" }
疑问:是否缺少相关配置,导致系统无法使用新增的Volume Server?
分析与解决方案
问题根源
从系统状态和日志可明确:
- 新增的x.x.x.64节点仅拥有无集合标识(
collection="")的可写卷11,而你上传文件使用的chrisDir集合对应的可写卷全部在已耗尽空间的x.x.x.63节点上。 - 自动执行的
volume.balance仅迁移了chrisDir2集合的卷21,未处理chrisDir集合的卷,也未在新节点为该集合创建新的可写卷。
解决步骤
1. 为chrisDir集合在新节点创建可写卷
通过weed shell手动为目标集合分配新卷:
./weed shell # 为chrisDir集合创建单副本可写卷 > volume.create -collection=chrisDir -replication=000
执行后再次调用curl http://x.x.x.61:9333/dir/status?pretty=y,确认chrisDir集合的writables列表中出现新卷(归属x.x.x.64节点)。
2. 标记原节点的chrisDir卷为只读
原节点的chrisDir卷已无空间,需手动标记为只读,避免master继续分配这些卷:
./weed shell # 列出chrisDir集合的所有卷 > volume.list -collection=chrisDir # 逐个标记为只读(替换为实际卷ID) > volume.markReadOnly -volumeId=1 > volume.markReadOnly -volumeId=2 ... > volume.markReadOnly -volumeId=7
3. 手动触发balance脚本
强制执行卷均衡,确保新节点的卷被正常分配使用:
./weed shell > lock > volume.balance -force > unlock
4. 验证分配结果
执行以下命令,确认master会将chrisDir集合的文件分配到新节点:
curl http://x.x.x.61:9333/dir/assign?collection=chrisDir&pretty=y
若返回的url为x.x.x.64:8080,即可通过filer UI正常上传文件。
补充说明
你的master配置中[master.volume_growth]的copy_1=7是全局默认设置,但针对特定集合的卷增长需要确保新节点有足够容量。后续可针对集合单独配置卷增长策略,避免类似问题重复出现。
内容的提问来源于stack exchange,提问作者chrizonline
相关产品推荐
相关产品推荐

