Neo4j Docker容器重启后挂起并退出的问题排查求助
问题描述
我创建了一个自定义Docker镜像,通过wrapper脚本启动Neo4j并加载初始数据。首次启动容器时偶尔失败,推测是缓存问题或等待Neo4j启动的时间不足。
核心问题:停止容器后重启时,容器会下载插件,但随后挂起,报错./wrapper.sh: line 57: fg: job has terminated,且重启时/logs/debug.log无日志输出,无法定位是否为权限问题。
我的wrapper脚本
#!/bin/bash # THANK YOU! Special shout-out to @marcellodesales on GitHub # https://github.com/marcellodesales/neo4j-with-cypher-seed-docker/blob/master/wrapper.sh for such a great example script # Log the info with the same format as NEO4J outputs log_info() { # https://www.howtogeek.com/410442/how-to-display-the-date-and-time-in-the-linux-terminal-and-use-it-in-bash-scripts/ # printf '%s %s\n' "$(date -u +"%Y-%m-%d %H:%M:%S:%3N%z") INFO Wrapper: $1" # Display UTC time printf '%s %s\n' "$(date +"%Y-%m-%d %H:%M:%S:%3N%z") INFO Wrapper: $1" # Display local time (PST/PDT) return } # Adapted from https://github.com/neo4j/docker-neo4j/issues/166#issuecomment-486890785 # Alpine is not supported anymore, so this is newer # Refactoring: Marcello.deSales+github@gmail.com # turn on bash's job control # https://stackoverflow.com/questions/11821378/what-does-bashno-job-control-in-this-shell-mean/46829294#46829294 set -m # Start the primary process and put it in the background /docker-entrypoint.sh neo4j & # Wait for Neo4j log_info "Checking to see if Neo4j has started at http://${DB_HOST}:${DB_PORT}..." wget --quiet --tries=20 --waitretry=10 -O /dev/null http://${DB_HOST}:${DB_PORT} log_info "Neo4j has started 🤓" log_info "Importing data with auth ${NEO4J_AUTH}" # Import data log_info "Loading and importing Cypher file(s)..." for cypherFile in /var/lib/neo4j/import/*.data.cypher; do [ -f "$cypherFile" ] || break log_info "Running cypher ${cypherFile}" cat ${cypherFile} | bin/cypher-shell -u ${NEO4J_USER} -p ${NEO4J_PASSWORD} --fail-fast --format plain log_info "Renaming import file ${cypherFile}" mv ${cypherFile} ${cypherFile}.applied done log_info "Finished loading data" log_info "Running startup cypher script..." for cypherFile in /var/lib/neo4j/import/*.startup.cypher; do [ -f "$cypherFile" ] || break log_info "Running cypher ${cypherFile}" cat ${cypherFile} | bin/cypher-shell -u ${NEO4J_USER} -p ${NEO4J_PASSWORD} --fail-fast --format plain done log_info "Finished running startup script" # now we bring the primary process back into the foreground # and leave it there fg %1
我的Dockerfile
FROM neo4j ENV NEO4J_USER=neo4j ENV NEO4J_PASSWORD=s3cr3t ENV NEO4J_AUTH=${NEO4J_USER}/${NEO4J_PASSWORD} ENV NEO4JLABS_PLUGINS='["apoc", "graph-data-science"]' ENV NEO4J_HOME='/var/lib/neo4j' ENV DB_HOST='localhost' ENV DB_PORT=7474 ENV NEO4J_dbms_logs_debug_level='DEBUG' ENV NEO4J_dbms_logs_user_stdout__enabled='true' EXPOSE 7474 7473 7687 COPY initial-data/ /var/lib/neo4j/import/ COPY ./docker-scripts/wrapper.sh wrapper.sh ENTRYPOINT ["./wrapper.sh"]
解决方案与定位方法
一、紧急修复:跳过重复初始化逻辑
重启时初始数据已导入(文件已重命名为.applied),但脚本仍会等待Neo4j启动,若此时Neo4j进程提前退出就会触发fg报错。添加判断跳过已完成的初始化:
修改wrapper脚本,在启动Neo4j后加入以下逻辑:
# 新增:检查是否已完成数据导入 if ls /var/lib/neo4j/import/*.data.cypher.applied 1> /dev/null 2>&1; then log_info "Initial data already imported, skipping data loading step" else # Wait for Neo4j log_info "Checking to see if Neo4j has started at http://${DB_HOST}:${DB_PORT}..." wget --quiet --tries=20 --waitretry=10 -O /dev/null http://${DB_HOST}:${DB_PORT} log_info "Neo4j has started 🤓" log_info "Importing data with auth ${NEO4J_AUTH}" # Import data log_info "Loading and importing Cypher file(s)..." for cypherFile in /var/lib/neo4j/import/*.data.cypher; do [ -f "$cypherFile" ] || break log_info "Running cypher ${cypherFile}" cat ${cypherFile} | bin/cypher-shell -u ${NEO4J_USER} -p ${NEO4J_PASSWORD} --fail-fast --format plain log_info "Renaming import file ${cypherFile}" mv ${cypherFile} ${cypherFile}.applied done log_info "Finished loading data" fi # 启动脚本仍需每次执行 log_info "Running startup cypher script..." for cypherFile in /var/lib/neo4j/import/*.startup.cypher; do [ -f "$cypherFile" ] || break log_info "Running cypher ${cypherFile}" cat ${cypherFile} | bin/cypher-shell -u ${NEO4J_USER} -p ${NEO4J_PASSWORD} --fail-fast --format plain done log_info "Finished running startup script"
二、定位Neo4j进程提前退出原因
1. 捕获后台进程状态
在启动Neo4j后记录PID,定期检查进程是否存活:
# Start the primary process and put it in the background /docker-entrypoint.sh neo4j & NEO4J_PID=$! # 新增:检查Neo4j进程是否存活 check_neo4j_alive() { if ! kill -0 $NEO4J_PID 2>/dev/null; then log_info "Neo4j process (PID $NEO4J_PID) has exited unexpectedly" # 打印Neo4j日志以便排查 cat /var/lib/neo4j/logs/neo4j.log 2>/dev/null || log_info "No neo4j.log found" exit 1 fi } # 在wget等待前后都检查进程状态 check_neo4j_alive wget --quiet --tries=20 --waitretry=10 -O /dev/null http://${DB_HOST}:${DB_PORT} check_neo4j_alive
2. 修复日志输出问题
重启时debug.log无输出,大概率是权限或配置问题:
- 在Dockerfile中添加权限配置:
RUN chmod +x wrapper.sh # 确保日志目录权限正确 RUN chown -R neo4j:neo4j /var/lib/neo4j/logs - 修改
log_info函数,同时输出到文件和stdout:log_info() { local msg="$1" local timestamp=$(date +"%Y-%m-%d %H:%M:%S:%3N%z") local log_line="${timestamp} INFO Wrapper: ${msg}" echo "$log_line" echo "$log_line" >> /var/lib/neo4j/logs/wrapper.log return }
3. 优化启动等待逻辑
wget仅检查HTTP端口,改用cypher-shell测试Bolt端口(真正就绪状态):
# 替换原wget等待逻辑 log_info "Waiting for Neo4j Bolt port to be ready..." until bin/cypher-shell -u ${NEO4J_USER} -p ${NEO4J_PASSWORD} -d system "RETURN 1" --format plain >/dev/null 2>&1; do log_info "Neo4j not ready yet, waiting 5 seconds..." sleep 5 done log_info "Neo4j has started 🤓"
三、权限问题排查
- 检查容器内文件权限:执行
docker exec -it <container_id> ls -l /var/lib/neo4j/,确保import、logs目录所属用户为neo4j - 在Dockerfile中设置文件权限:
COPY initial-data/ /var/lib/neo4j/import/ RUN chown -R neo4j:neo4j /var/lib/neo4j/import/ COPY ./docker-scripts/wrapper.sh wrapper.sh RUN chmod +x wrapper.sh && chown neo4j:neo4j wrapper.sh - 避免以root用户运行脚本,在wrapper开头添加:
if [ "$(id -u)" = "0" ]; then exec su - neo4j -c "$0 $@" fi
内容的提问来源于stack exchange,提问作者gdantimi
相关产品推荐
相关产品推荐

