一、架构规划
1.1 集群节点信息
哨兵节点(Sentinel)
- 10.20.2.36:27799
- 10.20.2.37:27799
- 10.20.2.38:27799
Redis 主从节点
- 主节点(Master):10.20.2.122:7799
- 从节点(Slave):
- 10.20.2.118:7799
- 10.20.2.121:7799
二、环境准备
2.1 创建工作目录
在每台服务器上创建 Redis 和 Sentinel 的工作目录:
# Redis 数据目录
mkdir ml-produce-caster-7799
# Sentinel 工作目录
mkdir ml-produce-caster-sentinel-27799
三、Redis 主从配置
3.1 主节点配置(10.20.2.122)
创建 redis.conf 配置文件:
bind 0.0.0.0
protected-mode no
port 7799
daemonize yes
logfile "redis.log"
pidfile "redis.pid"
dir "/home/ant/ml-produce-caster-7799"
appendonly no
maxmemory 4gb
maxmemory-policy allkeys-lru
save 86400 1
3.2 从节点配置(10.20.2.118 和 10.20.2.121)
创建 redis.conf 配置文件,在主节点配置基础上添加:
bind 0.0.0.0
protected-mode no
port 7799
daemonize yes
logfile "redis.log"
pidfile "redis.pid"
dir "/home/ant/ml-produce-caster-7799"
appendonly no
maxmemory 4gb
maxmemory-policy allkeys-lru
save 86400 1
# 指定主节点信息
replicaof 10.20.2.122 7799
3.3 启动 Redis 服务
创建 redis.sh 启动脚本:
#!/bin/sh
# Simple Redis init.d script
REDISPORT=7799
EXEC=redis-server
CLIEXEC=redis-cli
PIDFILE=redis.pid
CONF="redis.conf"
case "$1" in
start)
if [ -f $PIDFILE ]
then
echo "$PIDFILE exists, process is already running or crashed"
else
echo "Starting Redis server..."
$EXEC $CONF
fi
;;
stop)
if [ ! -f $PIDFILE ]
then
echo "$PIDFILE does not exist, process is not running"
else
PID=$(cat $PIDFILE)
echo "Stopping ..."
$CLIEXEC -h localhost -p $REDISPORT shutdown
while [ -x /proc/${PID} ]
do
echo "Waiting for Redis to shutdown ..."
sleep 1
done
echo "Redis stopped"
fi
;;
*)
echo "Please use start or stop as first argument"
;;
esac
添加执行权限:
chmod +x redis.sh
在三台服务器上分别启动 Redis:
./redis.sh start
验证 Redis 进程:
netstat -tunlp | grep 7799
四、Sentinel 哨兵配置
4.1 Sentinel 配置文件
在三台哨兵服务器(10.20.2.36、10.20.2.37、10.20.2.38)上创建 sentinel.conf:
protected-mode no
port 27799
daemonize yes
logfile "redis.log"
pidfile "redis.pid"
dir "/home/ant/ml-produce-caster-sentinel-27799"
# 监控主节点配置
# 至少需要 2 个哨兵节点同意,才能判定主节点故障并进行故障转移
sentinel monitor mymaster 10.20.2.122 7799 2
# 判定服务器 down 掉的时间周期,默认 3000 毫秒(3秒)
sentinel down-after-milliseconds mymaster 3000
# 故障节点的最大超时时间为 1800 毫秒(1.8秒)
sentinel failover-timeout mymaster 1800
4.2 Sentinel 启动脚本
创建 redis.sh 启动脚本:
#!/bin/sh
# Simple Redis Sentinel init.d script
REDISPORT=27799
EXEC=redis-sentinel
CLIEXEC=redis-cli
PIDFILE=redis.pid
CONF="sentinel.conf"
case "$1" in
start)
if [ -f $PIDFILE ]
then
echo "$PIDFILE exists, process is already running or crashed"
else
echo "Starting Redis sentinel..."
$EXEC $CONF
fi
;;
stop)
if [ ! -f $PIDFILE ]
then
echo "$PIDFILE does not exist, process is not running"
else
PID=$(cat $PIDFILE)
echo "Stopping ..."
$CLIEXEC -h localhost -p $REDISPORT shutdown
while [ -x /proc/${PID} ]
do
echo "Waiting for Redis to shutdown ..."
sleep 1
done
echo "Redis stopped"
fi
;;
*)
echo "Please use start or stop as first argument"
;;
esac
添加执行权限:
chmod +x redis.sh
4.3 启动 Sentinel
在三台服务器上分别启动 Sentinel:
./redis.sh start
或直接使用命令:
redis-sentinel sentinel.conf
五、健康检查与验证
5.1 基础检查命令
# 1. 检查 Sentinel 进程是否运行
ps -ef | grep redis-sentinel | grep -v grep
# 2. Sentinel PING 测试
redis-cli -h 10.20.2.37 -p 27799 ping
# 3. 检查 Sentinel 端口监听
netstat -tunlp | grep 27799
5.2 监控信息查询
# 查看所有被 Sentinel 监控的 master
redis-cli -h 10.20.2.37 -p 27799 sentinel masters
# 查看指定 master(mymaster)状态
redis-cli -h 10.20.2.37 -p 27799 sentinel master mymaster
# 查看指定 master(mymaster)的 slave 列表
redis-cli -h 10.20.2.37 -p 27799 sentinel slaves mymaster
# 查看 Sentinel 自身运行状态
redis-cli -h 10.20.2.37 -p 27799 info sentinel
5.3 故障转移测试
# 1. 手动停止主节点,模拟故障
redis-cli -h 10.20.2.122 -p 7799 shutdown
# 2. 等待 Sentinel 自动执行故障转移(约 3-6 秒)
# 3. 查看新的主节点信息
redis-cli -h 10.20.2.118 -p 7799 info replication
redis-cli -h 10.20.2.121 -p 7799 info replication
redis-cli -h 10.20.2.122 -p 7799 info replication
# 4. 通过 Sentinel 查看最新的主节点
redis-cli -h 10.20.2.37 -p 27799 sentinel get-master-addr-by-name mymaster
六、配置说明
6.1 Redis 配置参数说明
| 参数 | 说明 |
|---|---|
| bind | 绑定地址,0.0.0.0 表示允许所有 IP 访问 |
| protected-mode | 保护模式,设置为 no |
| port | Redis 服务端口 |
| daemonize | 是否后台运行 |
| maxmemory | 最大内存限制 |
| maxmemory-policy | 内存淘汰策略 |
| save | RDB 持久化策略 |
| replicaof | 指定主节点地址(仅从节点配置) |
6.2 Sentinel 配置参数说明
| 参数 | 说明 |
|---|---|
| sentinel monitor | 监控的主节点及仲裁数量 |
| sentinel down-after-milliseconds | 判定节点下线的时间阈值 |
| sentinel failover-timeout | 故障转移超时时间 |
七、常见问题
7.1 Sentinel 无法连接主节点
检查防火墙规则和网络连通性:
telnet 10.20.2.122 7799
7.2 故障转移未自动执行
检查 Sentinel 配置中的仲裁数量(quorum),确保至少有该数量的 Sentinel 节点在线。
7.3 查看 Sentinel 日志
tail -f /home/ant/ml-produce-caster-sentinel-27799/redis.log
八、运维建议
- 监控告警:建议接入监控系统,监控 Sentinel 和 Redis 的运行状态
- 日志管理:定期清理和归档日志文件,避免磁盘占用过高
- 定期演练:定期进行故障转移演练,确保高可用机制正常工作
- 备份策略:根据业务需求配置合理的持久化和备份策略