全部笔记All notes

Redis 哨兵集群搭建教程

阅读 4m 19s4m 19s read

一、架构规划

1.1 集群节点信息

哨兵节点(Sentinel)

  • 10.20.2.36:27799
  • 10.20.2.37:27799
  • 10.20.2.38:27799

Redis 主从节点

  • 主节点(Master):10.20.2.122:7799
  • 从节点(Slave):
    • 10.20.2.118:7799
    • 10.20.2.121:7799

二、环境准备

2.1 创建工作目录

在每台服务器上创建 Redis 和 Sentinel 的工作目录:

# Redis 数据目录
mkdir ml-produce-caster-7799

# Sentinel 工作目录
mkdir ml-produce-caster-sentinel-27799

三、Redis 主从配置

3.1 主节点配置(10.20.2.122)

创建 redis.conf 配置文件:

bind 0.0.0.0
protected-mode no
port 7799
daemonize yes
logfile "redis.log"
pidfile "redis.pid"
dir "/home/ant/ml-produce-caster-7799"
appendonly no
maxmemory 4gb
maxmemory-policy allkeys-lru
save 86400 1

3.2 从节点配置(10.20.2.118 和 10.20.2.121)

创建 redis.conf 配置文件,在主节点配置基础上添加:

bind 0.0.0.0
protected-mode no
port 7799
daemonize yes
logfile "redis.log"
pidfile "redis.pid"
dir "/home/ant/ml-produce-caster-7799"
appendonly no
maxmemory 4gb
maxmemory-policy allkeys-lru
save 86400 1

# 指定主节点信息
replicaof 10.20.2.122 7799

3.3 启动 Redis 服务

创建 redis.sh 启动脚本:

#!/bin/sh
# Simple Redis init.d script

REDISPORT=7799
EXEC=redis-server
CLIEXEC=redis-cli
PIDFILE=redis.pid
CONF="redis.conf"

case "$1" in
    start)
        if [ -f $PIDFILE ]
        then
                echo "$PIDFILE exists, process is already running or crashed"
        else
                echo "Starting Redis server..."
                $EXEC $CONF
        fi
        ;;
    stop)
        if [ ! -f $PIDFILE ]
        then
                echo "$PIDFILE does not exist, process is not running"
        else
                PID=$(cat $PIDFILE)
                echo "Stopping ..."
                $CLIEXEC -h localhost -p $REDISPORT shutdown
                while [ -x /proc/${PID} ]
                do
                    echo "Waiting for Redis to shutdown ..."
                    sleep 1
                done
                echo "Redis stopped"
        fi
        ;;
    *)
        echo "Please use start or stop as first argument"
        ;;
esac

添加执行权限:

chmod +x redis.sh

在三台服务器上分别启动 Redis:

./redis.sh start

验证 Redis 进程:

netstat -tunlp | grep 7799

四、Sentinel 哨兵配置

4.1 Sentinel 配置文件

在三台哨兵服务器(10.20.2.36、10.20.2.37、10.20.2.38)上创建 sentinel.conf:

protected-mode no
port 27799
daemonize yes
logfile "redis.log"
pidfile "redis.pid"
dir "/home/ant/ml-produce-caster-sentinel-27799"

# 监控主节点配置
# 至少需要 2 个哨兵节点同意,才能判定主节点故障并进行故障转移
sentinel monitor mymaster 10.20.2.122 7799 2

# 判定服务器 down 掉的时间周期,默认 3000 毫秒(3秒)
sentinel down-after-milliseconds mymaster 3000

# 故障节点的最大超时时间为 1800 毫秒(1.8秒)
sentinel failover-timeout mymaster 1800

4.2 Sentinel 启动脚本

创建 redis.sh 启动脚本:

#!/bin/sh
# Simple Redis Sentinel init.d script

REDISPORT=27799
EXEC=redis-sentinel
CLIEXEC=redis-cli
PIDFILE=redis.pid
CONF="sentinel.conf"

case "$1" in
    start)
        if [ -f $PIDFILE ]
        then
                echo "$PIDFILE exists, process is already running or crashed"
        else
                echo "Starting Redis sentinel..."
                $EXEC $CONF
        fi
        ;;
    stop)
        if [ ! -f $PIDFILE ]
        then
                echo "$PIDFILE does not exist, process is not running"
        else
                PID=$(cat $PIDFILE)
                echo "Stopping ..."
                $CLIEXEC -h localhost -p $REDISPORT shutdown
                while [ -x /proc/${PID} ]
                do
                    echo "Waiting for Redis to shutdown ..."
                    sleep 1
                done
                echo "Redis stopped"
        fi
        ;;
    *)
        echo "Please use start or stop as first argument"
        ;;
esac

添加执行权限:

chmod +x redis.sh

4.3 启动 Sentinel

在三台服务器上分别启动 Sentinel:

./redis.sh start

或直接使用命令:

redis-sentinel sentinel.conf

五、健康检查与验证

5.1 基础检查命令

# 1. 检查 Sentinel 进程是否运行
ps -ef | grep redis-sentinel | grep -v grep

# 2. Sentinel PING 测试
redis-cli -h 10.20.2.37 -p 27799 ping

# 3. 检查 Sentinel 端口监听
netstat -tunlp | grep 27799

5.2 监控信息查询

# 查看所有被 Sentinel 监控的 master
redis-cli -h 10.20.2.37 -p 27799 sentinel masters

# 查看指定 master(mymaster)状态
redis-cli -h 10.20.2.37 -p 27799 sentinel master mymaster

# 查看指定 master(mymaster)的 slave 列表
redis-cli -h 10.20.2.37 -p 27799 sentinel slaves mymaster

# 查看 Sentinel 自身运行状态
redis-cli -h 10.20.2.37 -p 27799 info sentinel

5.3 故障转移测试

# 1. 手动停止主节点,模拟故障
redis-cli -h 10.20.2.122 -p 7799 shutdown

# 2. 等待 Sentinel 自动执行故障转移(约 3-6 秒)

# 3. 查看新的主节点信息
redis-cli -h 10.20.2.118 -p 7799 info replication
redis-cli -h 10.20.2.121 -p 7799 info replication
redis-cli -h 10.20.2.122 -p 7799 info replication

# 4. 通过 Sentinel 查看最新的主节点
redis-cli -h 10.20.2.37 -p 27799 sentinel get-master-addr-by-name mymaster

六、配置说明

6.1 Redis 配置参数说明

参数说明
bind绑定地址,0.0.0.0 表示允许所有 IP 访问
protected-mode保护模式,设置为 no
portRedis 服务端口
daemonize是否后台运行
maxmemory最大内存限制
maxmemory-policy内存淘汰策略
saveRDB 持久化策略
replicaof指定主节点地址(仅从节点配置)

6.2 Sentinel 配置参数说明

参数说明
sentinel monitor监控的主节点及仲裁数量
sentinel down-after-milliseconds判定节点下线的时间阈值
sentinel failover-timeout故障转移超时时间

七、常见问题

7.1 Sentinel 无法连接主节点

检查防火墙规则和网络连通性:

telnet 10.20.2.122 7799

7.2 故障转移未自动执行

检查 Sentinel 配置中的仲裁数量(quorum),确保至少有该数量的 Sentinel 节点在线。

7.3 查看 Sentinel 日志

tail -f /home/ant/ml-produce-caster-sentinel-27799/redis.log

八、运维建议

  1. 监控告警:建议接入监控系统,监控 Sentinel 和 Redis 的运行状态
  2. 日志管理:定期清理和归档日志文件,避免磁盘占用过高
  3. 定期演练:定期进行故障转移演练,确保高可用机制正常工作
  4. 备份策略:根据业务需求配置合理的持久化和备份策略